Accepted for/Published in: Journal of Medical Internet Research
Date Submitted: Apr 14, 2026
Date Accepted: Aug 7, 2026
The Reliability of Human Evaluation of Large Language Models in Healthcare Settings: A Scoping Review
ABSTRACT
Background:
Integration of large language models (LLMs) into healthcare has accelerated rapidly, yet concerns about reliability issues pose risks to patient safety. Although human evaluation has emerged as a core approach when assessing LLM reliability, a systematic understanding of how this can be operationalized remains lacking across the various studies.
Objective:
This scoping review analyzed the evaluation indicators, evaluator characteristics, and evaluation approaches used for human evaluation-based LLM reliability assessments in healthcare settings, and compared differences between the clinical and public health domains.
Methods:
In line with the PRISMA-ScR guidelines, PubMed, the Web of Science, and Google Scholar were searched from January 2016 to July 2025. Original studies assessing the reliability of LLM-generated healthcare responses as assessed by humans were included. The extracted data were analyzed across 3 dimensions: what was evaluated, who evaluated, and how evaluation was conducted.
Results:
Of an initial 1,306 records, 76 studies were included (clinical: n = 30; public health: n = 46). 6 core reliability indicators were identified: Accuracy, Relevance, Completeness, Clarity, Safety, and Consistency. The clinical domain emphasized guideline concordance, internal consistency, and structural coherence, whereas the public health domain prioritized understandability and actionability. Clinician-only panels and Likert scale-based approaches predominated in the clinical domain, whereas mixed evaluator compositions and rubric-based approaches were more commonly employed in the public health domain. Common limitations included a limited evaluation scope, evaluator subjectivity, and a lack of standardized metrics.
Conclusions:
LLM reliability in healthcare is a context-dependent construct that cannot be adequately assessed using a single, uniform standard evaluation. Current structural limitations impede both cross-study comparability and evidence accumulation. Standardized, domain-specific evaluation frameworks are needed. As LLMs evolve toward agentic artificial intelligence (AI), the scope of evaluation should be extended to encompass reasoning processes and action selections.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.