Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Currently submitted to: JMIR AI

Date Submitted: Aug 14, 2026
Open Peer Review Period: Aug 20, 2026 - Oct 15, 2026
(currently open for review)

Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.

Expert Evaluation of Clinical Large Language Models in Intensive Care Medicine: A Multi-rater Assessment Study

  • Julian Klug; 
  • Raphaël Burger; 
  • Jean Bonnemain; 
  • Lionel Carrel; 
  • Matthieu Raboud; 
  • Victor Montaut; 
  • Aurélie Leuenberger; 
  • Alexis Bikfalvi; 
  • Giorgia Carra; 
  • Mary-Anne Hartley; 
  • Jean-Louis Raisaro

ABSTRACT

Background:

Large language models are increasingly being used for clinical decision support. Yet, a systematic evaluation of the performance and trustworthiness of such models in the intensive care unit (ICU) is lacking.

Objective:

This study aims to create an evaluation framework for large language models in the ICU based on clinical scenarios and evaluates the performance of Meditron-3, a 70-billion-parameter open-source medical large language model.

Methods:

Intensive care medicine specialists from a Swiss university hospital were recruited to create 200 clinical questions. The model used in this evaluation, Meditron-3, generated two independent answers per question. The same experts voted on the preferred answer and then rated it across 10 dimensions on a 5-point Likert scale: Alignment with Guidelines, Question Comprehension, Logical Reasoning, Relevance and Completeness, Harmlessness, Fairness, Contextual Awareness, Rater Confidence, Model Confidence, and Communication and Clarity. Interrater agreement was assessed using Fleiss' kappa, Krippendorff's alpha, and Gwet's AC1. Performance was stratified by interrater agreement quartiles, subspecialty, and clinical task type.

Results:

Eight experts produced 658 expert ratings and 788 answer evaluations. Vote agreement was fair (Fleiss' kappa 0.301, 95% confidence interval 0.23 to 0.38; Gwet's AC1 0.356, 95% confidence interval 0.28 to 0.43). Questions with high interrater agreement scored higher across all dimensions compared to questions with low agreement. Alignment with Guidelines correlated negatively with rater disagreement (Spearman rho -0.433, P < 0.001). The evaluated model scored highest on Communication and Clarity (mean 4.10, standard deviation 0.88) and Fairness (4.07, standard deviation 0.97), and lowest on Relevance and Completeness (3.06, standard deviation 1.12). Alignment with Guidelines was moderate (3.33, SD 1.09). Performance was consistent across subspecialties.

Conclusions:

A clinician-based evaluation framework can effectively benchmark large language models in the ICU and flag domain-specific short-comings. However, the fair interrater agreement highlights the inherent difficulty of evaluating large language model outputs in complex scenarios and underscores the need for standardized evaluation frameworks. The Meditron-3 model demonstrates acceptable performance in communication clarity and fairness but falls short in relevance, completeness, and guideline alignment for intensive care applications.


 Citation

Please cite as:

Klug J, Burger R, Bonnemain J, Carrel L, Raboud M, Montaut V, Leuenberger A, Bikfalvi A, Carra G, Hartley MA, Raisaro JL

Expert Evaluation of Clinical Large Language Models in Intensive Care Medicine: A Multi-rater Assessment Study

JMIR Preprints. 14/08/2026:109644

DOI: 10.2196/preprints.109644

URL: https://preprints.jmir.org/preprint/109644

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.