Accepted for/Published in: JMIR Formative Research
Date Submitted: Mar 22, 2026
Date Accepted: Jul 30, 2026
Performance and hallucination analysis of large language models on European anaesthesiology examinations: cross-sectional comparative study
ABSTRACT
Background:
Large language models (LLMs) have shown promising performance on medical examinations across specialties. However, comparative evaluations of current-generation LLMs across multiple European anaesthesiology examinations, alongside structured assessment of hallucinations versus question-related confusion, remain lacking.
Objective:
This study aimed to compare the performance of four state-of-the-art LLMs on anaesthesiology and intensive medicine examination questions and to assess their hallucination rates.
Methods:
This computational comparative study analysed 437 multiple-choice questions (1,748 queries) from three sources: nurse anaesthetist school examinations (IADE, n = 100), European Diploma in Anaesthesiology and Intensive Care (EDAIC, n = 219), and EDAIC Online Assessment (OLA, n = 118). Each question was submitted to four LLMs (Claude Sonnet 4.5, Gemini Pro 2.5, GPT-5, and Grok 4) using standardized prompts via default web interface settings. Responses were evaluated through structured consensus review by two examiners for accuracy, hallucinations, and question-related confusion. Statistical analysis included Friedman and Wilcoxon tests with Bonferroni correction, Cochran Q test, and logistic regression.
Results:
Success rates ranged from 86% to 94% across LLMs and examination types, approaching or exceeding typical candidate performance, representing substantial improvement over previously reported GPT-3.5. For EDAIC, Claude and Gemini outperformed GPT-5 and Grok 4 (P=.003), though some pairwise differences were attenuated after Bonferroni correction. Hallucination rates ranged from 11% to 20% without significant inter-model differences. All models exceeded the EDAIC passing threshold (70.25%).
Conclusions:
Current-generation LLMs demonstrate high performance across multiple European anaesthesiology examinations; however, consistent hallucination rates of 11–20% highlight a critical limitation for unsupervised educational use. These findings underscore the need for structured integration frameworks and systematic verification when deploying LLMs as learning tools in medical education. Clinical Trial: Not applicable
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.