Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: JMIR Formative Research

Date Submitted: Mar 22, 2026
Date Accepted: Jul 30, 2026

The final, peer-reviewed published version of this preprint can be found here:

Performance and Hallucination Analysis of Large Language Models on European Anesthesiology Examinations: Cross-Sectional Comparative Study

Andrei S, Giet T, Belouard A, Stefan MG, Popescu M, Tanaka S, Montravers P, Gouel A

Performance and Hallucination Analysis of Large Language Models on European Anesthesiology Examinations: Cross-Sectional Comparative Study

JMIR Form Res 2026;10:e95859

DOI: 10.2196/95859

PMID: 42678529

Performance and hallucination analysis of large language models on European anaesthesiology examinations: cross-sectional comparative study

  • Stefan Andrei; 
  • Thibaut Giet; 
  • Alexis Belouard; 
  • Mihai-Gabriel Stefan; 
  • Mihai Popescu; 
  • Sébastien Tanaka; 
  • Philippe Montravers; 
  • Aurélie Gouel

ABSTRACT

Background:

Large language models (LLMs) have shown promising performance on medical examinations across specialties. However, comparative evaluations of current-generation LLMs across multiple European anaesthesiology examinations, alongside structured assessment of hallucinations versus question-related confusion, remain lacking.

Objective:

This study aimed to compare the performance of four state-of-the-art LLMs on anaesthesiology and intensive medicine examination questions and to assess their hallucination rates.

Methods:

This computational comparative study analysed 437 multiple-choice questions (1,748 queries) from three sources: nurse anaesthetist school examinations (IADE, n = 100), European Diploma in Anaesthesiology and Intensive Care (EDAIC, n = 219), and EDAIC Online Assessment (OLA, n = 118). Each question was submitted to four LLMs (Claude Sonnet 4.5, Gemini Pro 2.5, GPT-5, and Grok 4) using standardized prompts via default web interface settings. Responses were evaluated through structured consensus review by two examiners for accuracy, hallucinations, and question-related confusion. Statistical analysis included Friedman and Wilcoxon tests with Bonferroni correction, Cochran Q test, and logistic regression.

Results:

Success rates ranged from 86% to 94% across LLMs and examination types, approaching or exceeding typical candidate performance, representing substantial improvement over previously reported GPT-3.5. For EDAIC, Claude and Gemini outperformed GPT-5 and Grok 4 (P=.003), though some pairwise differences were attenuated after Bonferroni correction. Hallucination rates ranged from 11% to 20% without significant inter-model differences. All models exceeded the EDAIC passing threshold (70.25%).

Conclusions:

Current-generation LLMs demonstrate high performance across multiple European anaesthesiology examinations; however, consistent hallucination rates of 11–20% highlight a critical limitation for unsupervised educational use. These findings underscore the need for structured integration frameworks and systematic verification when deploying LLMs as learning tools in medical education. Clinical Trial: Not applicable


 Citation

Please cite as:

Andrei S, Giet T, Belouard A, Stefan MG, Popescu M, Tanaka S, Montravers P, Gouel A

Performance and Hallucination Analysis of Large Language Models on European Anesthesiology Examinations: Cross-Sectional Comparative Study

JMIR Form Res 2026;10:e95859

DOI: 10.2196/95859

PMID: 42678529

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.