Currently submitted to: Journal of Medical Internet Research
Date Submitted: Aug 27, 2026
Open Peer Review Period: Aug 27, 2026 - Oct 22, 2026
(currently open for review)
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Exploring the Effectiveness of a Psychological Counseling Domain-Specific Chatbot (Nancy) for Improving Adult Mental Health: A Randomized Controlled Trial
ABSTRACT
Background:
Psychological counseling chatbots may expand access to mental health support, but evidence has largely focused on scripted systems or general-purpose large language models (LLMs). Whether domain-specific psychological counseling LLMs provide greater clinical benefit remains unclear.
Objective:
This study evaluated the clinical effectiveness of Nancy, a psychological counseling chatbot built on the domain-specific PsyLLM, relative to ChatGPT, a psychoeducational E-book, and a waitlist control (WLC).
Methods:
In this 4-arm randomized controlled trial (RCT), 501 Mandarin-speaking adults aged 18 years or older with a Patient Health Questionnaire-9 (PHQ-9) score of 5 or higher or a Generalized Anxiety Disorder-7 (GAD-7) score of 5 or higher were randomized 1:1:1:1 to Nancy, ChatGPT, E-book, or the Waitlist Control (WLC) for 28 days. Primary outcomes were PHQ-9 and GAD-7 scores. Secondary outcomes included affect, perceived empathy, and platform-recorded engagement. Intention-to-treat (ITT) analyses used multiple imputation and baseline-adjusted analysis of covariance; per-protocol (PP) analyses were sensitivity analyses.
Results:
Nancy showed greater PHQ-9 reductions than ChatGPT (adjusted difference −2.13, 95% CI −3.66 to −0.59; Holm-adjusted p=.007; Cohen d=0.44), E-book (−5.81, 95% CI −7.23 to −4.38; p<.001; d=1.10), and WLC (−2.95, 95% CI −4.35 to −1.54; p<.001; d=0.61). GAD-7 reductions were also greater with Nancy than with ChatGPT (−1.58, 95% CI −3.08 to −0.08; Holm-adjusted p=.04; d=0.35), E-book (−3.35, 95% CI −4.70 to −2.00; p<.001; d=0.74), and WLC (−1.90, 95% CI −3.24 to −0.56; p=.01; d=0.44). Per-protocol findings were consistent. Positive and negative affect did not differ significantly between groups. Perceived empathy was similar between Nancy and ChatGPT (35.49 vs 34.92; p=.69; Hedges g=0.07), whereas conversational intensity was higher with Nancy (median 17.21 vs 16.39 turns per recorded use day; p<.001; Cliff δ=0.29).
Conclusions:
Nancy produced greater reductions in depressive and anxiety symptoms than ChatGPT, psychoeducation, and waitlist controls over 28 days. These findings suggest that domain-specific therapeutic alignment may provide added clinical value beyond general-purpose conversational support. Clinical Trial: Chinese Clinical Trial Registry (ChiCTR2600121508); https://www.chictr.org.cn/hvshowproject.html?id=298372&v=1.0
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.