Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: Journal of Medical Internet Research

Date Submitted: May 6, 2026
Open Peer Review Period: May 5, 2026 - Jun 30, 2026
Date Accepted: Jun 30, 2026
(closed for review but you can still tweet)

The final, peer-reviewed published version of this preprint can be found here:

Conversational Large Language Models for Vestibular Diagnosis in Outpatient Clinics: Prospective Multicenter Diagnostic Accuracy Study

Li H, Lu C, Zhang R, Jiang H, Yu Y, Zhang S, Lin Q, Wu P

Conversational Large Language Models for Vestibular Diagnosis in Outpatient Clinics: Prospective Multicenter Diagnostic Accuracy Study

J Med Internet Res 2026;28:e100442

DOI: 10.2196/100442

Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.

Conversational Large Language Models for Vestibular Diagnosis in Outpatient Clinics: Prospective Multicenter Diagnostic Accuracy Study

  • Huawei Li; 
  • Chongkai Lu; 
  • Ruiqi Zhang; 
  • Huaili Jiang; 
  • Yanping Yu; 
  • Sulin Zhang; 
  • Qin Lin; 
  • Peixia Wu

ABSTRACT

Background:

Background:

Vestibular disorders are common, burdensome, and frequently misdiagnosed, particularly in non-specialist settings where history-taking is often incomplete or inconsistently structured. Digital health tools that standardize symptom elicitation could improve diagnostic triage, but most existing systems rely on static questionnaires or rule-based logic. Large language models (LLMs) offer a more flexible alternative through adaptive, natural-language consultations, but prospective evidence from real clinical workflows remains scarce.

Objective:

Objective:

This study aimed to benchmark LLM diagnostic performance using static vestibular histories and to prospectively evaluate a locally deployed conversational LLM agent embedded in outpatient vertigo clinics.

Methods:

Methods:

We conducted a two-phase diagnostic accuracy study. In the history-based evaluation (HBE), 10 LLMs and 5 senior otolaryngologists independently reviewed 227 structured vertigo histories, including 138 real patient questionnaires and 89 expert-simulated cases. In the prospective clinical evaluation (PCE), 176 outpatients with vertigo or dizziness were included in the analytic cohort across five centers in China. A nurse-assisted tablet-based agent powered by DeepSeek-R1, deployed locally within institutional infrastructure, conducted multi-turn symptom-history dialogues. The agent received history information only and did not receive physical examination or ancillary test findings. Attending clinicians were blinded to the agent output and recorded final clinical diagnoses after routine assessment. The primary outcome was Top-1 diagnostic concordance with the reference diagnosis; secondary outcomes included HBE Top-3 accuracy and disorder-specific performance.

Results:

Results:

In the HBE, Gemini-2.5-pro and o1 achieved the highest Top-1 accuracy (69.6%; 95% CI 63.2%-75.5% for each), and no LLM significantly outperformed the specialist panel majority vote (63.9%; 95% CI 57.3%-70.1%; McNemar test, P>.10 for all models). In the PCE, the median age was 51 years (IQR 38-61), and 128 of 176 participants (72.7%) were women. The conversational agent matched the reference diagnosis in 140 of 176 patients (79.55%; 95% CI 72.8%-85.2%). Concordance was highest for benign paroxysmal positional vertigo (43/44, 97.73%; 95% CI 88.0%-99.9%) and Ménière disease (27/30, 90.00%; 95% CI 73.5%-97.9%), and lower for vestibular migraine (32/45, 71.11%; 95% CI 55.7%-83.6%) and persistent postural-perceptual dizziness (19/28, 67.86%; 95% CI 47.6%-84.1%).

Conclusions:

Conclusions:

A locally deployed, history-only conversational LLM agent achieved approximately 80% concordance with specialist final diagnoses in a prospective multicenter outpatient workflow, with particularly high performance for benign paroxysmal positional vertigo. These findings support the development of conversational LLMs as clinician-facing tools for structured history-taking and diagnostic support, especially in settings with limited vestibular expertise. Future studies should test whether such systems improve clinical decisions, reduce unnecessary resource use, and maintain safety across languages and health care settings.


 Citation

Please cite as:

Li H, Lu C, Zhang R, Jiang H, Yu Y, Zhang S, Lin Q, Wu P

Conversational Large Language Models for Vestibular Diagnosis in Outpatient Clinics: Prospective Multicenter Diagnostic Accuracy Study

J Med Internet Res 2026;28:e100442

DOI: 10.2196/100442

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.