Currently submitted to: JMIR Medical Education
Date Submitted: Aug 6, 2026
Open Peer Review Period: Aug 7, 2026 - Oct 2, 2026
(currently open for review)
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Characterizing large language model generative artificial intelligence variability in the production of objective structured clinical examination stations
ABSTRACT
Background:
Designing high-quality Objective Structured Clinical Examination (OSCE) stations is a time-consuming process. Generative artificial intelligence (AI) represents a promising path to accelerate content creation by automating the generation of scenarios. A growing number of AI tools is now available for this purpose.
Objective:
To assess the variability between generative AI models in their ability to produce OSCE stations in the field of paediatrics.
Methods:
A structured prompt was developed based on the French national OSCE guidelines for medical education. Five distinct AI models were provided with this prompt, alongside the neonatal jaundice chapter from the French pediatric reference textbook, to generate 6 complete OSCE stations.
Results:
Prompt compliance was high for ChatGPT 5.1, ChatGPT 5.2, Gemini 3.0 Pro, and Claude Opus 4.5, while it was lower for Grok 4.1. Expert-rated quality was generally high, with few factual errors or missing information across models. However usability differed significantly between models. This was also true for several quality dimensions such as checklist clarity, embedding of checklist answers within vignettes, and ease of standardized patient formation. ChatGPT 5.1 required the most revisions and Gemini most often rated usable as is. Significant inter-model differences were observed in diagnostics, only with ChatGPT 5.1 sampling all three neonatal jaundice categories. Contextual variables showed systematic narrowing across models. Clinical grid density was consistent (10-12 items per station), but thematic distribution differed markedly. Soft skills coverage varied significantly across models (p=0.002), none of them consistently representing all communication competency domains.
Conclusions:
Large language models can generate structurally compliant OSCE stations, but surface compliance conceals substantive inter-model differences in diagnostic coverage, contextual diversity, and soft skills representation, that compromise content validity. No model currently meets the criteria for unsupervised deployment in a summative assessment bank. The choice of model carries pedagogical implications and expert curation remains essential before integration into high-stakes assessment workflows.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.