Accepted for/Published in: Journal of Medical Internet Research
Date Submitted: Dec 22, 2025
Date Accepted: Jul 3, 2026
Stepwise Diagnostic Evaluation of Chinese Large Language Models: A Comparative Study of Common and Rare Diseases
ABSTRACT
Background:
The rapid development of Chinese large language models (LLMs) offers significant potential for clinical decision support; however, their comparative diagnostic performance across common versus rare diseases within the specific context of the Chinese medical system remains underexplored.
Objective:
Our study aimed to evaluate the diagnostic capabilities of LLMs for common diseases and rare diseases using clinical vignettes within a Hypothetico-deductive framework, and to identify their potential and limitations for clinical diagnosis.
Methods:
We evaluated four Chinese LLMs using a clinical scenario approach, with incrementally provided patient information. This study included 28 cases of chronic obstructive pulmonary disease ( COPD) and 28 cases of relapsing polychondritis (RP). Evaluation metrics included the accuracy of the top three differential diagnoses, the first-ranked differential diagnosis, and the final diagnosis.
Results:
Data collection from the China Clinical Case Results Database occurred March 31–April 14, 2025. Overall, LLMs demonstrated significantly higher final diagnostic accuracy for COPD compared to RP (P<.001). While weighted accuracy scores were similar for COPD (P=.53) , they varied significantly for RP (P=.01) , with DouBao (1.18) outperforming Leftdoctor GPT (0.32; P=.05). Furthermore, incremental information significantly improved RP diagnosis for DeepSeek (32.14% to 71.43%; P=.007) and DouBao (35.71% to 78.57%; P=.003) , whereas Kimi and Leftdoctor GPT showed no significant improvement.
Conclusions:
Chinese LLMs show promise in assisting clinical diagnosis, particularly for common diseases. However, their diagnostic ability for rare diseases with limited data remains a concern. Given the performance variations among LLMs, their use should be carefully considered based on disease type and clinical scenario.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.