Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: Journal of Medical Internet Research

Date Submitted: Mar 12, 2026
Date Accepted: Jul 23, 2026

The final, peer-reviewed published version of this preprint can be found here:

Incremental Diagnostic Value of Clinical Information for Large Language Models Across Multiple Organs: Retrospective Study

Zhang J, Wang X, Zhao Y, Cui X, Zhao L, Zhao X, Zhang H, Zhu Z

Incremental Diagnostic Value of Clinical Information for Large Language Models Across Multiple Organs: Retrospective Study

J Med Internet Res 2026;28:e94904

DOI: 10.2196/94904

PMID: 42647068

Incremental Diagnostic Value of Clinical Information in Multi-Organ Diagnosis: A Comparison Between Large Language Models and Radiologists

  • Jinqi Zhang; 
  • Xiaoyi Wang; 
  • Yanfeng Zhao; 
  • Xiaolin Cui; 
  • Liyun Zhao; 
  • Xinming Zhao; 
  • Hongmei Zhang; 
  • Zheng Zhu

ABSTRACT

Background:

Although large language models (LLMs) have demonstrated the ability to generate the impression section from radiology findings automatically, the incremental diagnostic value of clinical information for these models remains unclear.

Objective:

To evaluate the incremental diagnostic value of clinical information for LLMs and compare their performance with that of radiologists.

Methods:

This retrospective study included radiology reports from patients with histopathologically confirmed liver, lung, and breast diseases from two institutions between October 2021 and February 2025. We defined three progressive information input scenarios: (a) basic patient information and imaging findings; (b) Scenario a plus chief complaint or clinical history; and (c) Scenario b plus key laboratory results. Scenario-based data were input into three general-purpose LLMs (DeepSeek-R1, Gemini 2.5 Pro, and GPT-4o), generating 2709 entries. Diagnostic accuracy was assessed for both benign-malignant differentiation and disease diagnosis, with histopathology serving as the reference standard. Accuracy was compared among scenarios and against radiologist performance using McNemar tests, and P values were adjusted using the Holm–Bonferroni correction for multiple comparisons.

Results:

A total of 301 pathologically confirmed cases were included (mean age, 53.5 years ± 12.0 [SD]; 208 female). In the liver cohort, A numerical trend toward higher accuracy was observed in Scenario c compared with Scenario a across all three models (Scenario c range, 72.3%–76.2% vs Scenario a range, 64.4%–68.3%; all P > .05); however, these differences did not reach statistical significance after Holm–Bonferroni correction. Notably, the DeepSeek-R1 model in Scenario c achieved the highest diagnostic accuracy (76.2% [77 of 101]), with no evidence of a difference compared with radiologists (81.2% [82 of 101]; P = .18). In contrast, results in the lung and breast cohorts were more heterogeneous. In the lung cohort, GPT-4o achieved its highest accuracy in Scenario a for disease diagnosis (73.9%), which exceeded its performance in Scenario b (69.6%) and Scenario c (71.7%), suggesting that additional clinical information did not confer a consistent benefit. Gemini 2.5 Pro in Scenario b achieved the highest accuracy in this cohort (78.3% [72 of 92]); however, no statistically significant difference was found compared with radiologists (87.0% [80 of 92]; adjusted P = .25). In the breast cohort, the diagnosis accuracy was numerically highest in Scenario a, and it decreased numerically with the addition of laboratory tests, although no significant difference was found between Scenarios a and c (67.6% [73 of 108] vs 65.7% [71 of 108]; P = .77).

Conclusions:

While the addition of clinical information was associated with a numeric trend toward higher diagnostic accuracy overall, this trend was heterogeneous across models and disease types, and no statistically significant improvement was demonstrated after adjustment for multiple comparisons.


 Citation

Please cite as:

Zhang J, Wang X, Zhao Y, Cui X, Zhao L, Zhao X, Zhang H, Zhu Z

Incremental Diagnostic Value of Clinical Information for Large Language Models Across Multiple Organs: Retrospective Study

J Med Internet Res 2026;28:e94904

DOI: 10.2196/94904

PMID: 42647068

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.