Currently submitted to: JMIR AI
Date Submitted: Aug 17, 2026
Open Peer Review Period: Aug 24, 2026 - Oct 19, 2026
(currently open for review)
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Multimodal large language models versus physicians for real-world inpatient diagnosis: Retrospective comparative
ABSTRACT
Background:
Large language models (LLMs) are increasingly proposed for diagnostic support. However, validation of their performance in real-world multimodal inpatient care remains limited, particularly in low- and middle-income country (LMIC) hospital settings.
Objective:
This study aimed to evaluate the diagnostic accuracy and clinical reasoning of leading multimodal LLMs compared to routine clinical practice within a South African tertiary hospital.
Methods:
We conducted a retrospective evaluation of 10 leading multimodal LLMs using 539 complete inpatient cases from a South African tertiary public hospital. The dataset included radiology imaging, radiology reports, laboratory results, and clinical narratives. Expert clinician panels adjudicated 300 cases to establish reference diagnoses, differential diagnoses, and clinical reasoning. Model outputs and routine recorded ward diagnoses were scored using a panel-calibrated, three-model LLM Jury.
Results:
Mean LLM scores were tightly clustered despite 50-fold cost differences, and on average all models exceeded routine ward diagnostic performance (>99.9% bootstrap probability). In a senior clinician-adjudicated tie-breaker subset maximised for model–panel disagreement, a predefined Top-3 LLM ensemble exceeded the original expert-panel consensus with 97% bootstrap probability. Residual safety risks and model non-response motivate prospective evaluation with rigorous adjudication and clinical oversight.
Conclusions:
Multimodal LLMs demonstrated high diagnostic performance in a complex, real-world LMIC inpatient setting, surpassing recorded routine ward diagnoses in 60-80% of cases depending on the model. Our finding that even the cheapest models outperform the routine ward diagnoses and are comparable to the best models suggests that affordable AI clinical decision support in resource-constrained environments is now worth serious consideration.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.