Accepted for/Published in: JMIR Medical Informatics
Date Submitted: Oct 19, 2025
Date Accepted: Jun 16, 2026
Large Language Models for Endodontic Symptom Assessment and Treatment Planning Using Image-Free Clinical Records: A Comparative Evaluation Study
ABSTRACT
Background:
Accurate assessment of the pulpal status is essential for achieving successful endodontic outcomes. However, direct evaluation remains inherently challenging because the pulp is surrounded by calcified tissue, necessitating reliance on clinical and radiographic examinations for making diagnostic and prognostic decisions. These procedures demand substantial clinical expertise and time, and less experienced clinicians often face challenges that may lead to errors in diagnosis and treatment planning. Recent advancements in large language models (LLMs) offer promising opportunities to enhance clinical reasoning by facilitating evidence integration and supporting methodical diagnostic decision-making.
Objective:
This study aimed to evaluate the clinical applicability of LLMs by comparing the diagnostic accuracy and clinical validity of their responses with those of human experts.
Methods:
Between January 2011 and December 2022, 100 cases were randomly selected from the clinical records of outpatients who visited the Department of Conservative Dentistry or Advanced General Dentistry (AGD) at Yonsei University Dental Hospital. Four prompt types, combining 2 variables (language and role), were used as input for 4 LLMs to generate diagnostic and treatment plan responses. Diagnostic accuracy was evaluated on a 0–2 scale, whereas the validity and relevance of treatment plans were assessed using a 5-point Likert scale. Evaluators included specialists, residents, and senior dental students.
Results:
ChatGPT achieved the highest diagnostic accuracy among the LLMs, with a score of 0.98 ± 0.82 on Korean-doctor prompts (p < .001). In contrast, Clova X recorded the lowest accuracy at 0.23 ± 0.63 on English-patient prompts (p = .004). Across all disease categories, AGD specialists demonstrated the highest diagnostic accuracy (pulpal: 0.65, periodontal: 0.71, periapical: 0.83), with higher sensitivity but lower specificity than those exhibited by the other groups. ChatGPT also showed favorable performance among LLMs, with accuracies of 0.63 (95% CI: 0.53–0.72) for pulpal, 0.72 (95% CI: 0.62–0.81) for periodontal, and 0.81 (95% CI: 0.72–0.87) for periapical disease, which were comparable to the accuracies of AGD and endodontics residents.
Conclusions:
Conclusions:
ChatGPT 4.0 generated more consistent and reliable diagnostic and treatment planning responses than did other LLMs, showing potential to support clinical decision-making in endodontic practice. However, hallucination issues and interpretation biases influenced by user experience remain. Therefore, continuous clinical supervision and comprehensive user training are necessary for safe and effective clinical application. Clinical Trial: The study protocol was approved by the Institutional Review Board of the Dental Hospital of Yonsei University (IRB No. 2-2024-0075).
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.