Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: Online Journal of Public Health Informatics

Date Submitted: Apr 5, 2026
Open Peer Review Period: Apr 5, 2026 - May 31, 2026
Date Accepted: Jul 27, 2026
(closed for review but you can still tweet)

The final, peer-reviewed published version of this preprint can be found here:

Safety-Oriented Evaluation of Large Language Models in Health Care: Guideline-Informed Systematic Review

Matsuoka H, Takahashi T, Semitsu T, Matsuoka K, Hayashi R

Safety-Oriented Evaluation of Large Language Models in Health Care: Guideline-Informed Systematic Review

Online J Public Health Inform 2026;18:e97240

DOI: 10.2196/97240

PMID: 42623243

PMCID: 13492481

Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.

Safety-Oriented Evaluation of Large Language Models in Healthcare: A Guideline-Informed Systematic Review

  • Hikaru Matsuoka; 
  • Takayuki Takahashi; 
  • Takayuki Semitsu; 
  • Keiko Matsuoka; 
  • Ryoma Hayashi

ABSTRACT

Background:

Large language models (LLMs) are rapidly emerging in healthcare, offering opportunities in decision support, education, and research, but raising critical concerns about safety, reliability, and ethics. While several guidelines for trustworthy AI exist in business and technology, few systematic reviews have applied them to medical contexts.

Objective:

This study aimed to conduct a systematic review of LLM research in healthcare, applying the AI Guidelines for Business as a framework across twelve domains including safety, reliability, ethics, transparency, fairness, inclusiveness, privacy, security, robustness, data quality, and verifiability.

Methods:

Following PRISMA methodology, databases such as PubMed, Web of Science, Scopus, arXiv, and IEEE Xplore were searched. The literature was systematically searched on January 15, 2025. Eligible studies were classified and independently verified. A total of 247 studies were included. Each article was analyzed to determine the corresponding medical specialty, classification as surgical, internal medicine, or emergency care, the intended purpose of use, and the study population. Performance measures were converted into percentages on a 0–100 scale for comparability.

Results:

Research was concentrated on internal medicine (23.1%), surgery (23.1%), and radiology (14.2%), while many specialties were underrepresented. Evaluation domains emphasized accuracy (57.9%), fairness and inclusiveness (37.7%), data quality (19.0%), and misinformation control (17.8%). Reported performance was moderate to high, with accuracy at 75.0%, fairness at 58.1%, data quality at 75.9%, and misinformation control at 73.8%. Privacy protection (n=2) and security assurance (n=0) were almost entirely absent from the literature reviewed.

Conclusions:

The review highlights strengths but also methodological gaps and disciplinary imbalances, underscoring the need for broader specialty engagement and comprehensive, guideline-based evaluation to ensure safe, reliable, and equitable implementation of LLMs in healthcare.


 Citation

Please cite as:

Matsuoka H, Takahashi T, Semitsu T, Matsuoka K, Hayashi R

Safety-Oriented Evaluation of Large Language Models in Health Care: Guideline-Informed Systematic Review

Online J Public Health Inform 2026;18:e97240

DOI: 10.2196/97240

PMID: 42623243

PMCID: 13492481

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.