Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: Journal of Medical Internet Research

Date Submitted: Apr 2, 2026
Date Accepted: Jul 2, 2026

The final, peer-reviewed published version of this preprint can be found here:

Effectiveness, Safety, and Workflow Burden of Large Language Model–Based Medical Report Generation: Systematic Review

Huang JL, Zhu JQ, Ni X

Effectiveness, Safety, and Workflow Burden of Large Language Model–Based Medical Report Generation: Systematic Review

J Med Internet Res 2026;28:e97007

DOI: 10.2196/97007

PMID: 42715525

Effectiveness, Safety, and Workflow Burden of Large Language Model-Based Medical Report Generation: A Systematic Review

  • Jie-Lin Huang; 
  • Ji-Qing Zhu; 
  • Xiaoguang Ni

ABSTRACT

Background:

LLM-based systems are increasingly used for medical report generation, but clinical readiness remains uncertain because safety and workflow outcomes are sparsely and inconsistently reported.

Objective:

To assess effectiveness, safety, and workflow impact of LLM-based medical report generation.

Methods:

We searched PubMed/MEDLINE, Embase, Web of Science Core Collection, Scopus, and the Cochrane Library for studies published from January 1, 2016, to March 18, 2026. Eligible studies evaluated LLMs, multimodal LLMs, or vision-language models for image-to-report generation, findings-to-impression generation, or report drafting in radiology, pathology, ultrasound, or endoscopy. Structured extraction and risk-of-bias assessment were performed. Main outcomes were clinically significant, omission, and commission error rates, reporting time, edit burden, expert acceptance, and blinded expert preference. Meta-analysis was not performed because no clinically comparable primary outcome had at least 2 analyzable studies.

Results:

Forty-six studies were included. Chest x-ray predominated (24 studies); 9 studies had moderate risk of bias, 29 high, and 8 serious. Clinically informative safety and workflow evidence came from only a small number of single studies. In one chest radiograph study, acceptance of AI-generated reports was similar to that of radiologist reports (70.5% vs 73.3%), but false-negative findings remained slightly higher (18.5% vs 17.8%). In one clinician-collaboration chest x-ray study, AI reports were equivalent or preferred in 77.7% and 56.1% of 2 datasets, yet clinically significant errors persisted. In one brain MRI study, AI assistance reduced reading time from 61 to 53 seconds, whereas in one impression-drafting study AI increased editing time and edit distance.

Conclusions:

LLM-based report generation may have assistive potential as a support tool, but sparse and heterogeneous safety and workflow data preclude pooled meta-analysis and limit confidence in routine clinical deployment. Current evidence is heavily skewed toward templated chest radiography, and generalizability to complex cross-sectional imaging, pathology, or endoscopic reporting remains unproven.


 Citation

Please cite as:

Huang JL, Zhu JQ, Ni X

Effectiveness, Safety, and Workflow Burden of Large Language Model–Based Medical Report Generation: Systematic Review

J Med Internet Res 2026;28:e97007

DOI: 10.2196/97007

PMID: 42715525

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.