Accepted for/Published in: Journal of Medical Internet Research
Date Submitted: Apr 2, 2026
Date Accepted: Jul 2, 2026
Effectiveness, Safety, and Workflow Burden of Large Language Model-Based Medical Report Generation: A Systematic Review
ABSTRACT
Background:
LLM-based systems are increasingly used for medical report generation, but clinical readiness remains uncertain because safety and workflow outcomes are sparsely and inconsistently reported.
Objective:
To assess effectiveness, safety, and workflow impact of LLM-based medical report generation.
Methods:
We searched PubMed/MEDLINE, Embase, Web of Science Core Collection, Scopus, and the Cochrane Library for studies published from January 1, 2016, to March 18, 2026. Eligible studies evaluated LLMs, multimodal LLMs, or vision-language models for image-to-report generation, findings-to-impression generation, or report drafting in radiology, pathology, ultrasound, or endoscopy. Structured extraction and risk-of-bias assessment were performed. Main outcomes were clinically significant, omission, and commission error rates, reporting time, edit burden, expert acceptance, and blinded expert preference. Meta-analysis was not performed because no clinically comparable primary outcome had at least 2 analyzable studies.
Results:
Forty-six studies were included. Chest x-ray predominated (24 studies); 9 studies had moderate risk of bias, 29 high, and 8 serious. Clinically informative safety and workflow evidence came from only a small number of single studies. In one chest radiograph study, acceptance of AI-generated reports was similar to that of radiologist reports (70.5% vs 73.3%), but false-negative findings remained slightly higher (18.5% vs 17.8%). In one clinician-collaboration chest x-ray study, AI reports were equivalent or preferred in 77.7% and 56.1% of 2 datasets, yet clinically significant errors persisted. In one brain MRI study, AI assistance reduced reading time from 61 to 53 seconds, whereas in one impression-drafting study AI increased editing time and edit distance.
Conclusions:
LLM-based report generation may have assistive potential as a support tool, but sparse and heterogeneous safety and workflow data preclude pooled meta-analysis and limit confidence in routine clinical deployment. Current evidence is heavily skewed toward templated chest radiography, and generalizability to complex cross-sectional imaging, pathology, or endoscopic reporting remains unproven.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.