Previously submitted to: JMIR Formative Research (no longer under consideration since May 20, 2026)
Date Submitted: Aug 1, 2025
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Application Research and Comparative Analysis of Large Language Model-Assisted Lesson Plan Development for Rare Breast Diseases
ABSTRACT
Background:
Rare breast diseases are infrequently encountered in clinical training, leading to gaps in medical students’ knowledge. Large Language Models (LLMs) show promise in assisting educators, particularly for complex topics like rare diseases where standardized teaching materials are scarce.
Objective:
This study aims to evaluate and compare the capabilities of prominent LLMs in generating lesson plans for rare breast diseases targeted at clinical medical students.
Methods:
Ten representative rare breast diseases were selected based on prior bibliometric analysis and expert consensus. Four LLMs - ChatGPT-4o, Grok3, Deepseek R1, and Doubao - were prompted in Chinese to produce structured teaching plans. Three experienced breast surgery teaching faculty evaluated the generated lesson plans using a standardized score (0-5 points) across six dimensions: Generation Capability, Medical Accuracy, Completeness, Readability, Applicability, and Interactivity. Quantitative scores were compared statistically.
Results:
All LLMs successfully generated lesson plans, but with significant differences across all evaluated dimensions (P < 0.001). Deepseek consistently achieved the highest mean scores in all 6 dimensions. ChatGPT ranked second, also demonstrating strong performance, particularly in Generation Capability and Completeness. Grok3 and Doubao showed moderate performance, with Doubao scoring relatively higher in Readability and Accuracy, while Grok3 performed better in Applicability and Interactivity compared to Doubao.
Conclusions:
LLMs, particularly advanced models like Deepseek and ChatGPT, demonstrate significant potential in assisting the generation of high-quality lesson plans for rare breast diseases, while variability in quality and occasional inaccuracies necessitate expert review. Integration of LLM-generated materials into medical curricula holds promise for enhancing rare disease education.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.