Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: Journal of Medical Internet Research

Date Submitted: May 18, 2026
Date Accepted: Jun 24, 2026

The final, peer-reviewed published version of this preprint can be found here:

Initial-Visit Specialty Triage in Rare Diseases Using Large Language Models: Retrospective Benchmarking Study

Song J, Xu Z, Xiao M, Bi C, Zhang Y, Zheng X, Li X, Cao Q, Lu Z, Yang H, Shen B

Initial-Visit Specialty Triage in Rare Diseases Using Large Language Models: Retrospective Benchmarking Study

J Med Internet Res 2026;28:e101711

DOI: 10.2196/101711

PMID: 42490550

Initial-Visit Specialty Triage in Rare Diseases Using Large Language Models: A Retrospective Benchmarking Study

  • Jie Song; 
  • Zhichuan Xu; 
  • Meng Xiao; 
  • Cheng Bi; 
  • Yuxin Zhang; 
  • Xin Zheng; 
  • Xiaoran Li; 
  • Qiongfang Cao; 
  • Ziyu Lu; 
  • Hao Yang; 
  • Bairong Shen

ABSTRACT

Background:

Specialty triage at first contact is an overlooked step in early diagnostic pathways for rare diseases. Patients often present with overlapping, multisystem, and atypical manifestations, making first-visit specialty selection challenging and potentially prolonging diagnostic pathways.

Objective:

To evaluate the accuracy, response time, and consistency of large language models (LLMs) for initial-visit specialty triage in rare diseases across multiple datasets, and to compare their performance with registered nurses and non-medical participants.

Methods:

In this retrospective benchmarking study, we used five rare disease datasets: a publication-derived case set, three RareBench-derived datasets, and a Facial phenotype-Gene-Disease Dataset-derived set. Fourteen LLMs were evaluated over five independent runs per case. Performance was assessed using accuracy, response time, and consistency, with subgroup analyses by model accessibility, reasoning mode, parameter scale, and phenotype count. Human comparison was conducted on the publication-derived case set using registered nurses and non-medical participants.

Results:

Across datasets, accuracy varied from 0.4378 to 0.7141. Claude-opus-4-5 achieved the highest accuracy (0.7141) and consistency (0.9653), averaging 10.79 s per case. GPT-5.1 had the shortest response time (3.39 s per case) and high accuracy (0.6948). Proprietary models achieved higher average accuracy than open-weight models (0.6973 vs 0.6365). Standard models achieved higher average accuracy than thinking models (0.6789 vs 0.5826) and had shorter response times. Accuracy varied by phenotype count, with higher performance in cases with 1–2 or more than 14 phenotypes. On the publication-derived case set, LLMs achieved higher average accuracy than registered nurses and non-medical participants (0.5978 vs 0.4914 and 0.4573).

Conclusions:

LLMs showed potential as assistive tools for initial-visit specialty triage in rare diseases. Model choice, reasoning mode, and phenotype information density substantially influenced performance, suggesting that deployment should prioritize validated models with strong accuracy–efficiency balance. Future work should evaluate LLM-based specialty triage in prospective clinical settings and develop clinician-supervised workflows with traceable evidence support. Clinical Trial: Not applicable.


 Citation

Please cite as:

Song J, Xu Z, Xiao M, Bi C, Zhang Y, Zheng X, Li X, Cao Q, Lu Z, Yang H, Shen B

Initial-Visit Specialty Triage in Rare Diseases Using Large Language Models: Retrospective Benchmarking Study

J Med Internet Res 2026;28:e101711

DOI: 10.2196/101711

PMID: 42490550

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.