Currently accepted at: JMIR Medical Education
Date Submitted: Sep 17, 2025
Date Accepted: Jul 23, 2026
This paper has been accepted and is currently in production.
It will appear shortly on 10.2196/84266
The final accepted version (not copyedited yet) is in this tab.
Performance of Cloud-Hosted Large Vision-Language Models on the Japanese National Examination for Clinical Laboratory Technicians: Comparative Benchmarking Study
ABSTRACT
Background:
Cloud-hosted large vision–language models (LVLMs) often outperform open-weight models on multimodal benchmarks, but their applicability to healthcare exams that require both text and image reasoning remains unclear. Existing studies on Japan’s National Examination for Clinical Laboratory Technicians (CLT) mainly used text-only settings and a limited set of models.
Objective:
To evaluate the accuracy and applicability of cloud-hosted LVLMs on the CLT exam, quantify the contribution of images, examine alignment with human performance across medical subfields, and test generalizability on post–knowledge cutoff questions.
Methods:
We built Kensagishi QA, a benchmark of 800 CLT questions (2020–2023) with text-only and image-attached items, standardized JSON, and strict scoring (all correct options required). We tested GPT-5, GPT-4o, GPT-4o-mini, Gemini 1.5 Pro, Gemini 2.5 Pro, and Neva-22B using a uniform “numbers-only” response prompt; images were passed as Base64 when available. Accuracy was computed overall, by item type, and by medical subfield. Human reference was drawn from published CLT results. Spearman correlation assessed model–human alignment. Generalization was evaluated on 200 questions from the 71st CLT (2025), released after the models’ knowledge cutoffs.
Results:
Top models exceeded the 60% passing threshold: GPT-5 reached 93.8% with images (88.0% without), Gemini 2.5 Pro 92.4% (86.6%), and GPT-4o 81.5% (79.6%). Gemini 1.5 Pro achieved 69.3% (68.3%), GPT-4o-mini 55.8% (52.9%), and Neva-22B 29.6% (32.1%). Providing images consistently improved accuracy for GPT-5, Gemini 2.5 Pro, and GPT-4o; Neva-22B showed no such benefit and more formatting errors. By subfield, GPT-5 was uniformly strong; Gemini 2.5 Pro and GPT-4o were stable; Gemini 1.5 Pro varied; GPT-4o-mini/Neva-22B lagged. Model–human rank correlations were moderate for GPT-4o (ρ≈0.62 with images) and significant for Gemini 1.5 Pro without images (ρ≈0.67). On the 71st CLT (2025), accuracies were comparable to historical means (e.g., GPT-5 92.7%, Gemini 2.5 Pro 91.3%, GPT-4o 83.5%), indicating no material degradation on unseen questions.
Conclusions:
Cloud-hosted LVLMs show strong performance on the CLT, and visual input is a key driver of accuracy for leading models. Models that better mirror human difficulty patterns (e.g., GPT-4o) may suit learning support, while highest-accuracy models (e.g., GPT-5, Gemini 2.5 Pro) fit accuracy-critical uses. Future work should add multidimensional quality metrics (factuality, instruction following, safety, Japanese medical terminology), evaluate open-weight medical LVLMs, and explore prompt/RAG strategies to support education and clinical training.
Citation
The author of this paper has made a PDF available, but requires the user to login, or create an account.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.