Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Currently submitted to: JMIR Medical Education

Date Submitted: Aug 6, 2026
Open Peer Review Period: Aug 7, 2026 - Oct 2, 2026
(currently open for review)

Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.

Characterizing large language model generative artificial intelligence variability in the production of objective structured clinical examination stations

  • Kristen Joseph-Delaffon; 
  • Maxime Desgrouas; 
  • Sophie Catanese; 
  • Julien Lejeune; 
  • Julien Nait-Kaci; 
  • Eric Piver; 
  • Isaure Breteau; 
  • Sophie Leducq; 
  • Philippe Gatault; 
  • Raoul Kanav Khanna; 
  • Denis Angoulvant; 
  • Nicolas Vallet

ABSTRACT

Background:

Designing high-quality Objective Structured Clinical Examination (OSCE) stations is a time-consuming process. Generative artificial intelligence (AI) represents a promising path to accelerate content creation by automating the generation of scenarios. A growing number of AI tools is now available for this purpose.

Objective:

To assess the variability between generative AI models in their ability to produce OSCE stations in the field of paediatrics.

Methods:

A structured prompt was developed based on the French national OSCE guidelines for medical education. Five distinct AI models were provided with this prompt, alongside the neonatal jaundice chapter from the French pediatric reference textbook, to generate 6 complete OSCE stations.

Results:

Prompt compliance was high for ChatGPT 5.1, ChatGPT 5.2, Gemini 3.0 Pro, and Claude Opus 4.5, while it was lower for Grok 4.1. Expert-rated quality was generally high, with few factual errors or missing information across models. However usability differed significantly between models. This was also true for several quality dimensions such as checklist clarity, embedding of checklist answers within vignettes, and ease of standardized patient formation. ChatGPT 5.1 required the most revisions and Gemini most often rated usable as is. Significant inter-model differences were observed in diagnostics, only with ChatGPT 5.1 sampling all three neonatal jaundice categories. Contextual variables showed systematic narrowing across models. Clinical grid density was consistent (10-12 items per station), but thematic distribution differed markedly. Soft skills coverage varied significantly across models (p=0.002), none of them consistently representing all communication competency domains.

Conclusions:

Large language models can generate structurally compliant OSCE stations, but surface compliance conceals substantive inter-model differences in diagnostic coverage, contextual diversity, and soft skills representation, that compromise content validity. No model currently meets the criteria for unsupervised deployment in a summative assessment bank. The choice of model carries pedagogical implications and expert curation remains essential before integration into high-stakes assessment workflows.


 Citation

Please cite as:

Joseph-Delaffon K, Desgrouas M, Catanese S, Lejeune J, Nait-Kaci J, Piver E, Breteau I, Leducq S, Gatault P, Khanna RK, Angoulvant D, Vallet N

Characterizing large language model generative artificial intelligence variability in the production of objective structured clinical examination stations

JMIR Preprints. 06/08/2026:108828

DOI: 10.2196/preprints.108828

URL: https://preprints.jmir.org/preprint/108828

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.