Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Currently submitted to: JMIR AI

Date Submitted: Aug 17, 2026
Open Peer Review Period: Aug 24, 2026 - Oct 19, 2026
(currently open for review)

Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.

Multimodal large language models versus physicians for real-world inpatient diagnosis: Retrospective comparative

  • Bruce Bassett; 
  • Amy Rouillard; 
  • Sitwala Mundia; 
  • Michael Cameron Gramanie; 
  • Linda Camara; 
  • Ziyaad Dangor; 
  • Shabir Madhi; 
  • Kajal Morar; 
  • Marlvin Themba Ncube; 
  • Ismail Kalla; 
  • Haroon Saloojee

ABSTRACT

Background:

Large language models (LLMs) are increasingly proposed for diagnostic support. However, validation of their performance in real-world multimodal inpatient care remains limited, particularly in low- and middle-income country (LMIC) hospital settings.

Objective:

This study aimed to evaluate the diagnostic accuracy and clinical reasoning of leading multimodal LLMs compared to routine clinical practice within a South African tertiary hospital.

Methods:

We conducted a retrospective evaluation of 10 leading multimodal LLMs using 539 complete inpatient cases from a South African tertiary public hospital. The dataset included radiology imaging, radiology reports, laboratory results, and clinical narratives. Expert clinician panels adjudicated 300 cases to establish reference diagnoses, differential diagnoses, and clinical reasoning. Model outputs and routine recorded ward diagnoses were scored using a panel-calibrated, three-model LLM Jury.

Results:

Mean LLM scores were tightly clustered despite 50-fold cost differences, and on average all models exceeded routine ward diagnostic performance (>99.9% bootstrap probability). In a senior clinician-adjudicated tie-breaker subset maximised for model–panel disagreement, a predefined Top-3 LLM ensemble exceeded the original expert-panel consensus with 97% bootstrap probability. Residual safety risks and model non-response motivate prospective evaluation with rigorous adjudication and clinical oversight.

Conclusions:

Multimodal LLMs demonstrated high diagnostic performance in a complex, real-world LMIC inpatient setting, surpassing recorded routine ward diagnoses in 60-80% of cases depending on the model. Our finding that even the cheapest models outperform the routine ward diagnoses and are comparable to the best models suggests that affordable AI clinical decision support in resource-constrained environments is now worth serious consideration.


 Citation

Please cite as:

Bassett B, Rouillard A, Mundia S, Gramanie MC, Camara L, Dangor Z, Madhi S, Morar K, Ncube MT, Kalla I, Saloojee H

Multimodal large language models versus physicians for real-world inpatient diagnosis: Retrospective comparative

JMIR Preprints. 17/08/2026:109825

DOI: 10.2196/preprints.109825

URL: https://preprints.jmir.org/preprint/109825

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.