Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Evaluation of artificial intelligence-generated medicine images through comparison of interpretations by students and specialists
ABSTRACT
Background:
Recent breakthroughs in generative artificial intelligence (AI) have enabled the creation of highly realistic images across diverse domains. However, despite this progress, the application and realism of AI-generated medical images remain insufficiently explored. Accurate medical imaging interpretation is critical for diagnosis and treatment, and any ability of AI to produce images indistinguishable from real ones could have significant implications for medical education, research, and clinical practice.
Objective:
This study aimed to evaluate the realism of AI-generated medical images from the perspectives of medical professionals and students. Specifically, we investigated whether current generative AI models can produce images that are realistic enough to confuse trained medical experts, and how their performance compares across different models and image types.
Methods:
We assessed four generative AI models—Gemini, DALL-E 3, Midjourney, and a customized Stable Diffusion model trained on open radiology datasets. Each model generated four axial abdominal CT and four sagittal brain MR images. Four authentic clinical images were added to each set, resulting in 40 images in total. An extended Visual Turing Test survey was conducted with ten participants: five medical students and five specialists in Anesthesiology and Pain Medicine. Participants rated each image on a five-point realism scale, indicating whether they believed it was real or AI-generated. Statistical analyses compared discrimination performance across participant groups, image modalities, and models.
Results:
Medical students struggled to distinguish real from AI-generated abdominal CT images (P = .064), indicating near-chance performance, but significantly differentiated brain MR images (P = .002). In contrast, specialists demonstrated high discrimination accuracy for both abdominal CT and brain MR images (P < .001). Among AI models, Gemini and the custom Stable Diffusion model achieved higher realism scores (mean rank ≈ 5), whereas DALL-E 3 and Midjourney performed poorly (mean rank ≈ 15), reflecting limited suitability of general-purpose models for medical imaging tasks.
Conclusions:
Current generative AI systems can produce visually convincing medical images, but experienced medical professionals can still reliably detect subtle inconsistencies, indicating that these models have not yet reached the level of realism needed to fully mimic genuine clinical imaging. Moreover, the customized model outperformed general-purpose systems, supporting the feasibility of locally engineered AI solutions for medical institutions. These findings highlight both the potential and current limitations of generative AI in medical imaging, underscoring the importance of domain-specific training for clinical applicability.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.