Accepted for/Published in: Journal of Medical Internet Research
Date Submitted: Nov 14, 2025
Date Accepted: Jul 23, 2026
Multimodal Large Language Models versus the Periodontist: Periodontitis Risk Assessment and Prevention Planning
ABSTRACT
Background:
Periodontitis is one of the most prevalent yet preventable oral diseases, indicated by multiple clinical and radiographic factors. As these factors are recorded in electronic health records (EHRs), their reuse offers opportunities for personalized risk assessment and targeted prevention. Predictive AI and traditional machine learning models support fragmented detection tasks, but lack the integration of textual and imaging predictors. Emerging multimodal large language models (M-LLMs) show promise in combining these data sources in clinical assessment. Evaluating the capabilities of M-LLMs and comparing them against the current clinical standard, is therefore essential to determine their potential as digital assistants.
Objective:
This study aimed to evaluate the ability of M-LLMs to assess periodontitis risk and suggest prevention strategies, based on EHR data and radiographic findings. Each M-LLM was individually evaluated by periodontal experts, benchmarked against other models, and compared to the periodontist as a reference.
Methods:
A vignette study was conducted following TRIPOD-guidelines for the evaluation of LLMs. Ten periodontal vignettes were created, each including a panoramic radiograph and textual EHR data. Three LLMs capable of reasoning and handling multimodal data, were compared to a periodontist who generated outputs manually, based on the same prompts and input data. Periodontal experts rated all outputs across six predefined criteria on a 5-point Likert scale. Statistical analyses evaluated overall performance per model, and tested whether performance varied per model, scenario complexity, or rater.
Results:
GPT o1 pro and Claude Sonnet 4 showed strong performance, with 86% of ratings deemed acceptable - comparable to the periodontist’s output (88%). Gemini 2.5 Pro was rated significantly lower than both the periodontist and the other models (60% acceptable; p<0.05). Radiographic interpretation consistently received lower scores than other abilities across all models and the periodontist, with Gemini rated below the acceptable threshold. The time required for completion ranged from a few seconds for Claude, to approximately 30 seconds for Gemini, 3 minutes for GPT, and 6 minutes for the periodontist.
Conclusions:
M-LLMs demonstrated strong reasoning abilities in periodontal assessment, yet their current capacity to interpret radiographs remains limited and potentially unreliable for detailed diagnostics. Periodontal evaluation primarily relies on textual EHR data, whereas radiographic parameters such as alveolar bone levels have a greater influence on risk than plaque-retentive factors. Rather than striving for “perfect” AI, which is neither realistic nor feasible, systems should perform at least comparably to the clinical standard and be satisfactory to periodontal experts. While some M-LLMs already perform at a similar level and much faster than the periodontist, clinical implementation should still consider their capabilities, limitations, and potential risks.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.