Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: Journal of Medical Internet Research

Date Submitted: Apr 14, 2026
Date Accepted: Aug 7, 2026

The final, peer-reviewed published version of this preprint can be found here:

The Reliability of Human Evaluation of Large Language Models in Health Care Settings: Scoping Review

Yang E, Ko S, Woo H

The Reliability of Human Evaluation of Large Language Models in Health Care Settings: Scoping Review

J Med Internet Res 2026;28:e98184

DOI: 10.2196/98184

PMID: 42618044

The Reliability of Human Evaluation of Large Language Models in Healthcare Settings: A Scoping Review

  • Euijun Yang; 
  • Siyeon Ko; 
  • Hyekyung Woo

ABSTRACT

Background:

Integration of large language models (LLMs) into healthcare has accelerated rapidly, yet concerns about reliability issues pose risks to patient safety. Although human evaluation has emerged as a core approach when assessing LLM reliability, a systematic understanding of how this can be operationalized remains lacking across the various studies.

Objective:

This scoping review analyzed the evaluation indicators, evaluator characteristics, and evaluation approaches used for human evaluation-based LLM reliability assessments in healthcare settings, and compared differences between the clinical and public health domains.

Methods:

In line with the PRISMA-ScR guidelines, PubMed, the Web of Science, and Google Scholar were searched from January 2016 to July 2025. Original studies assessing the reliability of LLM-generated healthcare responses as assessed by humans were included. The extracted data were analyzed across 3 dimensions: what was evaluated, who evaluated, and how evaluation was conducted.

Results:

Of an initial 1,306 records, 76 studies were included (clinical: n = 30; public health: n = 46). 6 core reliability indicators were identified: Accuracy, Relevance, Completeness, Clarity, Safety, and Consistency. The clinical domain emphasized guideline concordance, internal consistency, and structural coherence, whereas the public health domain prioritized understandability and actionability. Clinician-only panels and Likert scale-based approaches predominated in the clinical domain, whereas mixed evaluator compositions and rubric-based approaches were more commonly employed in the public health domain. Common limitations included a limited evaluation scope, evaluator subjectivity, and a lack of standardized metrics.

Conclusions:

LLM reliability in healthcare is a context-dependent construct that cannot be adequately assessed using a single, uniform standard evaluation. Current structural limitations impede both cross-study comparability and evidence accumulation. Standardized, domain-specific evaluation frameworks are needed. As LLMs evolve toward agentic artificial intelligence (AI), the scope of evaluation should be extended to encompass reasoning processes and action selections.


 Citation

Please cite as:

Yang E, Ko S, Woo H

The Reliability of Human Evaluation of Large Language Models in Health Care Settings: Scoping Review

J Med Internet Res 2026;28:e98184

DOI: 10.2196/98184

PMID: 42618044

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.