Generative Large Language Models in Mental Health Care Settings - Systematic Review and Meta-analysis
ABSTRACT
Background:
Large language models (LLMs) are emerging as digital tools in mental health care, but their performance and safety across clinical tasks remain unclear.
Objective:
This study aimed to systematically review and meta-analyze the effectiveness of general-purpose LLMs for mental health care tasks, including screening and diagnosis, clinical decision support, therapy support, documentation and monitoring, and patient education.
Methods:
Following Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines, we searched PubMed, ACM Digital Library, IEEE Xplore, and Google Scholar for studies published from November 2022 to June 2025 that evaluated general-purpose LLMs for mental health care applications. The protocol was registered in PROSPERO (CRD420251056593). Eligible designs included observational, experimental, implementation, and simulation or vignette-based studies with empirical outcomes. Risk of bias was assessed using the Mixed Methods Appraisal Tool. Where sufficient data were available, we conducted meta-analyses of LLM performance for screening and diagnosis, pooling sensitivity, specificity, precision, and accuracy using the inverse variance heterogeneity model. Other outcomes were synthesized narratively.
Results:
We included 41 studies evaluating LLMs for screening and diagnosis (n=20), clinical decision support (n=10), therapy support (n=9), documentation and monitoring (n=4), and patient education (n=2). Twelve screening and diagnostic studies (25 effect sizes) were pooled. Overall sensitivity was moderate (0.50; 95% CI 0.13-0.86; I²=99%), indicating that on average LLMs missed approximately half of true cases, while specificity was high (0.92; 95% CI 0.80-1.00; I²=98%). Precision was moderate (0.64; 95% CI 0.53-0.75; I²=94%), and overall accuracy was good (0.80; 95% CI 0.75-0.84; I²=66%). Subgroup analyses suggested higher sensitivity for GPT-4 than for GPT-3.5, with overlapping confidence intervals. Narrative synthesis indicated that LLMs can support clinical reasoning, triage, and documentation tasks, and can generate generally acceptable psychoeducational content. However, models frequently lacked clinical nuance, occasionally produced unsafe or misleading advice in complex presentations, and were rarely evaluated in prospective or real-world clinical settings.
Conclusions:
General-purpose LLMs show promising accuracy for mental health screening and diagnostic support, particularly for ruling out cases, and may augment clinical decision making, therapy support, documentation, and education. However, limited sensitivity, inconsistent performance across conditions and models, and a lack of real-world implementation studies mean that LLMs should currently function only as adjuncts to clinicians. Safe integration will require task-specific validation, continuous monitoring, human oversight, and robust governance and reporting frameworks. Clinical Trial: PROSPERO CRD420251056593.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.