Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Currently submitted to: JMIR AI

Date Submitted: Jan 12, 2026

Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.

Benchmarking Open- and Closed-Source Large Language Models in Spanish Health Specialized Examinations: A Comparative Performance Study

  • Alejandro Calvera-Rayo; 
  • Gonzalo Verdú; 
  • Manuel Morales-Ruiz; 
  • Aleix Fabregat-Bolufer

ABSTRACT

Background:

The artificial intelligence landscape is evolving rapidly, with competition between closed-source large language models (LLMs) and open-source alternatives. Although general-purpose benchmarks are widely used, domain-specific evaluation remains underexplored, which represents an urgent gap for medical education and laboratory medicine, where specialized reasoning, transparency, and cost-effectiveness will shape safe adoption.

Objective:

To evaluate and benchmark five state-of-the-art LLMs—three closed-source (ChatGPT o3, Gemini 2.5 Pro, Grok 4) and two open-source (DeepSeek R1, Qwen3 235B-A22B)—on the 2024 Spanish FIR (Pharmaceutical Internal Resident) and BIR (Biologist Internal Resident) health specialized examinations.

Methods:

The official 2024 FIR and BIR exams (200 multiple-choice questions each) were submitted using two complementary strategies: single-question prompting with image snippets and full-document prompting. Outcomes included global and subject-specific accuracy, multimodality, influence of prompt length, and response time. Analyses used descriptive statistics, McNemar test and Cohen’s kappa for consistency, and Friedman/Wilcoxon tests for paired comparisons.

Results:

With single-question snippets, closed-source models achieved top-tier accuracy and hypothetical first-rank positioning among human candidates in both exams: ChatGPT o3 (96.5% FIR, 98.5% BIR), Gemini 2.5 Pro (96.5% FIR, 99% BIR), and Grok 4 (92.5% FIR, 99.5% BIR). DeepSeek R1 also ranked first (89% FIR, 98% BIR), whereas Qwen3 showed lower accuracy (51% FIR, 87% BIR). Errors predominantly clustered in multimodal and chemistry-related domains, with Gemini 2.5 Pro demonstrating superior performance (75% accuracy on image-based questions). Full-document prompting negatively affected models' accuracy: ChatGPT o3 (80% FIR, 97.5% BIR), Gemini 2.5 Pro (93.5% FIR, 99% BIR), Grok 4 (23% FIR, 51.5% BIR), DeepSeek R1 (88% FIR, 96.5% BIR), while Qwen3 improved (59% FIR, 97.5% BIR). Per-question response times differed significantly across models (P<.001): ChatGPT o3 was fastest at 1.0 seconds per question, followed by DeepSeek R1 and Grok 4. All models showed longer response times for incorrect answers.

Conclusions:

Closed-source models retained slight advantages, particularly for multimodal reasoning, but open-source competitors, especially DeepSeek R1, approached competitive parity while remaining broadly accessible and democratizing. This convergence highlights the need for evaluation frameworks beyond simple adoption rates, integrating specialized performance and healthcare considerations before deploying LLMs in medical education and clinical laboratory decision support.


 Citation

Please cite as:

Calvera-Rayo A, Verdú G, Morales-Ruiz M, Fabregat-Bolufer A

Benchmarking Open- and Closed-Source Large Language Models in Spanish Health Specialized Examinations: A Comparative Performance Study

JMIR Preprints. 12/01/2026:91300

DOI: 10.2196/preprints.91300

URL: https://preprints.jmir.org/preprint/91300

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.