Accepted for/Published in: JMIR Mental Health
Date Submitted: Feb 26, 2026
Date Accepted: Jul 28, 2026
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Effectiveness of Chatbots in Mental Health Screening and Assessment: A Systematic Review
ABSTRACT
Background:
Mental health disorders affect 970 million people globally, yet over 50% do not access timely evaluation due to structural barriers and professional shortages. Chatbots and AI-based conversational agents have emerged as promising tools for mental health screening and assessment.
Objective:
To systematically evaluate the effectiveness, accuracy, reliability, and acceptability of chatbots and AI-based conversational agents in mental health screening and assessment in adults.
Methods:
Systematic search conducted in May 2025 across PubMed/MEDLINE, PsycINFO, Scopus, and Web of Science (2019-2025), following PRISMA 2020 guidelines. Eligible studies evaluated chatbots or AI for mental health screening/assessment in adults (≥18 years). Risk of bias was assessed using appropriate tools (RoB 2, QUADAS-2, JBI checklists, MMAT). This systematic review was registered in PROSPERO (CRD420251072392).
Results:
Eighteen studies (2021-2025) were included, with samples ranging from 20 to 3,902 participants. Rule-based chatbots demonstrated high reliability (Cronbach's α > 0.85) and good acceptability (AIM > 19/25). Generative models (LLMs) achieved sensitivities of 0.84-0.93 and specificities of 0.80-0.96 for depression and anxiety, with correlations up to r = 0.96 with expert clinicians in suicide risk assessment. Hybrid approaches combining LLMs with machine learning achieved exceptional performance in cognitive impairment (F1 = 92.1%, specificity = 99.6%). Most studies reported high user satisfaction (≥70%), though barriers existed in older populations. Methodological quality was heterogeneous with moderate risk of bias in critical dimensions.
Conclusions:
Chatbots and AI conversational agents demonstrate clinically relevant performance in mental health screening and assessment. However, safe implementation requires clear clinical protocols, professional supervision, integration with electronic health records, and active mitigation of algorithmic bias. These technologies should complement rather than replace clinical judgment. Clinical Trial: PROSPERO 2025 CRD420251072392
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.