Previously submitted to: Journal of Medical Internet Research (no longer under consideration since Sep 09, 2025)
Date Submitted: Jan 9, 2025
(closed for review but you can still tweet)
Large Language Models in Dermatology: A Systematic Review of Applications, Challenges, and Future Directions
Background:
The diagnosis and treatment of dermatological conditions rely heavily on physicians’ expertise, with inconsistencies in care quality persisting at the primary level. Recent advancements in large language models (LLMs), including GPT, MPT, PaLM, and LLaMA, show promise in medical knowledge extraction and intelligent consultation. However, challenges such as training data biases, limited model generalization, and ethical concerns impede the development of standardized LLM-based solutions.
Objective:
This systematic review aims to (1) summarize LLMs' state-of-the-art performance in dermatological tasks; (2) compare different LLM architectures’ technical characteristics and applicability; (3) analyze data-related, methodological, ethical, and practical challenges; and (4) propose recommendations for future research to fully harness LLMs’ potential in dermatological practices.
Methods:
A systematic literature search was conducted in PubMed, Web of Science, Embase, IEEE Xplore, and ACM Digital Library, with search strategy optimized by a medical librarian. The screening process followed PRISMA (Preferred Reporting Items for Systematic reviews and Meta-Analyses) guidelines, and methodological quality was assessed using the PROBAST tool. Two reviewers (WZL&ZYT) independently conducted literature screening, data extraction, and quality assessment.
Results:
The systematic search yielded 1247 records, of which 68 studies met the inclusion criteria. LLMs have demonstrated effectiveness in various dermatological applications, including diagnostic support (eg, Falcon-40B achieving 84% sensitivity and 91% specificity in melanoma classification), patient education (eg, GPT-generated materials with a readability level equivalent to the fifth grade), and drug discovery (eg, Christopoulou et al. applied an ensemble of BiLSTM and Transformer-based models to extract medication-related relations from electronic health records, identifying potential adverse drug events with a sensitivity of 91% and specificity of 95%.). The Results section presents a detailed comparison of the advantages, limitations, and applicability of different dermatology-specific LLM architectures. Despite these advancements, LLMs face persistent challenges in managing rare cases, integrating multimodal data, ensuring data privacy, and executing complex medical reasoning. A critical issue is the limited diversity in training datasets, with underrepresentation of certain skin tones, age groups, and rare conditions, which may exacerbate health disparities if unaddressed. Moreover, the ethical implications of LLM-based decision support, such as potential biases and liability concerns, remain largely unexplored.
Conclusions:
LLMs hold significant potential for advancing dermatological diagnosis and treatment; however, key challenges remain. Future research should focus on improving model generalization through multimodal data fusion and comprehensive knowledge integration, while also developing collaborative diagnostic frameworks that combine human expertise with machine learning. Priority areas include (1) curating large-scale, diverse, and representative dermatological datasets; (2) advancing few-shot learning and transfer learning to adapt LLMs to rare conditions; (3) developing privacy-preserving techniques such as federated learning and secure multiparty computation; and (4) establishing guidelines and protocols for responsible and transparent LLM deployment in clinical settings. Multidisciplinary efforts in data governance, ethics, and regulation will be crucial to ensure the safe, effective, and equitable implementation of these technologies.
Citation
Request queued. Please wait while the file is being generated. It may take some time.