Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: Journal of Medical Internet Research

Date Submitted: Mar 26, 2026
Date Accepted: Jul 27, 2026

The final, peer-reviewed published version of this preprint can be found here:

Machine Learning, Large Language Models, and Multimodal AI for Diagnosing Pediatric Rare Diseases: Scoping Review

Zhao J, Luo J, Li Q, Chen Y

Machine Learning, Large Language Models, and Multimodal AI for Diagnosing Pediatric Rare Diseases: Scoping Review

J Med Internet Res 2026;28:e96169

DOI: 10.2196/96169

Machine Learning, Large Language Models, and Multimodal AI for Diagnosing Pediatric Rare Diseases: A Scoping Review

  • Jungang Zhao; 
  • Jiawei Luo; 
  • Qiu Li; 
  • Yaolong Chen

ABSTRACT

Background:

Pediatric rare diseases often cause a prolonged diagnostic odyssey. Artificial intelligence (AI)—including machine learning (ML), deep learning (DL), large language models (LLMs), and multimodal systems—may support diagnosis, but the landscape of applications in children has not been mapped systematically.

Objective:

To map and synthesize the published evidence on the use of machine learning, deep learning, large language models, and multimodal artificial intelligence for diagnosing pediatric rare diseases, and to identify key gaps in validation, reporting, and real-world clinical implementation.

Methods:

We conducted a scoping review following JBI methodology and PRISMA-ScR. We searched PubMed, Scopus, and Web of Science from 1 January 2015 to 31 December 2025 for studies using ML/DL, LLMs, or multimodal AI to diagnose rare diseases in children (0–18 years). Eligibility was defined by the PCC framework (Population: pediatric rare disease; Concept: diagnostic AI; Context: any setting). Two reviewers screened titles/abstracts and full texts; data were extracted on disease type, model type, input modality, performance metrics, and reported challenges.

Results:

We retrieved 237 records (1 January 2015 to 31 December 2025) and removed 42 duplicates. After screening 195 unique records, we excluded 140 at title/abstract. Of 55 full-text reports assessed for eligibility, 13 were excluded; 42 studies were included. Studies spanned 2018–2025, with a sharp rise from 2019 and a peak of 16 in 2025. By technology, classical ML accounted for 23 studies (55%), deep learning for 8 (19%), facial AI for 4 (10%), multimodal for 4 (10%), and LLM for 3 (7%). The most represented disease groups were Mendelian/rare genetic (9), metabolic/newborn screening (5), and genetic syndromes (4). Input modalities were dominated by EHR/claims/text (15), facial image (7), and genomic data (5). Reported performance varied widely; LLM studies showed modest top-1 accuracy (8–13% on rare pediatric case reports) and emphasized the need for clinical oversight. Common challenges included single-center or retrospective design, small samples, limited external validation, and ethical issues such as privacy, bias, and false positives/negatives.

Conclusions:

The literature on ML, LLMs, and multimodal AI for diagnosing pediatric rare diseases is growing but fragmented. Classical ML and deep learning dominate the evidence base, while LLM and multimodal applications remain limited and LLM performance in rare pediatric settings is modest. Key gaps include prospective and multi-center validation, LLM studies in Chinese pediatric rare disease practice, and integration of agentic AI with real-world clinical workflows. This scoping map can guide future primary studies and systematic reviews. Clinical Trial: The protocol for this scoping review is registered on PROSPERO (CRD420261326146).


 Citation

Please cite as:

Zhao J, Luo J, Li Q, Chen Y

Machine Learning, Large Language Models, and Multimodal AI for Diagnosing Pediatric Rare Diseases: Scoping Review

J Med Internet Res 2026;28:e96169

DOI: 10.2196/96169

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.