Accepted for/Published in: JMIR Medical Education
Date Submitted: Jan 25, 2026
Date Accepted: Jul 17, 2026
Benchmarking Multimodal LLMs on Dental Chart Image Interpretation: A Comparison with Students and Clinicians
ABSTRACT
Background:
During the transition from preclinical-to-clinical training, dental students must learn to accurately read and interpret complex, tooth-centered electronic dental charts. However, limited protected teaching time in patient-centered clinics can result in gaps in chart-reading literacy, increasing cognitive load and the risk of chart misinterpretation. Multimodal large language models (LLMs) capable of processing chart images may offer scalable, interactive educational support, yet their chart-image performance in dental education contexts has not been rigorously benchmarked.
Objective:
To evaluate the performance of multimodal ChatGPT models on dental chart-image interpretation tasks of varying complexity and to compare the best-performing model with dental students and early-career clinician experts.
Methods:
A retrospective, cross-sectional benchmark study was conducted using de-identified dental chart images from 15 patients treated at Seoul National University Dental Hospital between January 1, 2017, and June 30, 2025. Charts containing Korean and English text were captured as sequential screenshots (154 total images; mean 10.27 images per patient). For each patient, 24 Korean-language questions were developed (360 total) across four categories: (A) general factual retrieval, (B) tooth- or procedure-specific retrieval, (C) interpretation requiring multientry synthesis, and (D) absent-information (unanswerable) questions to assess abstention. Nine multimodal ChatGPT models available from August 2 to 9, 2025 (KST) were evaluated under standardized, reset-conversation conditions. Model outputs were scored against a gold standard established by a senior faculty dentist using seven evaluation metrics, with Sentence-BERT (SBERT) similarity prespecified as the primary semantic measure. Human performance baselines included two third-year dental students and two first-year resident dentists. Nonparametric group comparisons were conducted using Kruskal–Wallis tests with Dunn’s post hoc analyses and Holm–Bonferroni adjustment.
Results:
Across 360 items, GPT-5 Thinking achieved the highest overall SBERT median score (0.900, IQR 0.525–1.000), followed by GPT-5 Pro (0.861, IQR 0.501–1.000) and OpenAI o3 (0.831, IQR 0.489–1.000). Kruskal–Wallis testing demonstrated significant overall group differences (P<.001). Relative to GPT-5 Thinking, Student1 differed only on Type D questions (P=.006), whereas Student2 showed no significant differences for Type A (P=.56) or Type C (P=.43) questions; remaining student comparisons were nonsignificant (P>.99). Early-career clinician experts outperformed GPT-5 Thinking on Type B questions (P=.023 and P=.009) and on Type C questions for one expert (P=.008). Type C tasks showed compressed SBERT distributions and low exact-match rates, indicating persistent difficulty in multientry synthesis.
Conclusions:
In this image-based dental chart interpretation benchmark, multimodal LLMs—particularly GPT-5 Thinking—demonstrated performance comparable to dental students across most task types but did not match early-career clinicians on tooth- or procedure-specific retrieval and integrative interpretation tasks. These findings support the use of multimodal LLMs as supervised educational tools for chart-reading practice and verification, rather than as replacements for clinical expertise.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.