Accepted for/Published in: Online Journal of Public Health Informatics
Date Submitted: Apr 5, 2026
Open Peer Review Period: Apr 5, 2026 - May 31, 2026
Date Accepted: Jul 27, 2026
(closed for review but you can still tweet)
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Safety-Oriented Evaluation of Large Language Models in Healthcare: A Guideline-Informed Systematic Review
ABSTRACT
Background:
Large language models (LLMs) are rapidly emerging in healthcare, offering opportunities in decision support, education, and research, but raising critical concerns about safety, reliability, and ethics. While several guidelines for trustworthy AI exist in business and technology, few systematic reviews have applied them to medical contexts.
Objective:
This study aimed to conduct a systematic review of LLM research in healthcare, applying the AI Guidelines for Business as a framework across twelve domains including safety, reliability, ethics, transparency, fairness, inclusiveness, privacy, security, robustness, data quality, and verifiability.
Methods:
Following PRISMA methodology, databases such as PubMed, Web of Science, Scopus, arXiv, and IEEE Xplore were searched. The literature was systematically searched on January 15, 2025. Eligible studies were classified and independently verified. A total of 247 studies were included. Each article was analyzed to determine the corresponding medical specialty, classification as surgical, internal medicine, or emergency care, the intended purpose of use, and the study population. Performance measures were converted into percentages on a 0–100 scale for comparability.
Results:
Research was concentrated on internal medicine (23.1%), surgery (23.1%), and radiology (14.2%), while many specialties were underrepresented. Evaluation domains emphasized accuracy (57.9%), fairness and inclusiveness (37.7%), data quality (19.0%), and misinformation control (17.8%). Reported performance was moderate to high, with accuracy at 75.0%, fairness at 58.1%, data quality at 75.9%, and misinformation control at 73.8%. Privacy protection (n=2) and security assurance (n=0) were almost entirely absent from the literature reviewed.
Conclusions:
The review highlights strengths but also methodological gaps and disciplinary imbalances, underscoring the need for broader specialty engagement and comprehensive, guideline-based evaluation to ensure safe, reliable, and equitable implementation of LLMs in healthcare.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.