Currently submitted to: Journal of Medical Internet Research
Date Submitted: Sep 10, 2026
Open Peer Review Period: Sep 11, 2026 - Nov 6, 2026
(currently open for review)
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Effectiveness and User Experience of a Hybrid LLM-Based Conversational Agent (Chatbot) for Depression and Anxiety in Korean Young Adults: Randomized Controlled Trial and Thematic Analysis
ABSTRACT
Background:
Depression and anxiety are highly prevalent among young adults, yet access to evidence-based psychological care remains limited by cost, stigma, and workforce shortages. Large language model (LLM)–based conversational agents have been proposed as scalable, low-barrier formats for delivering cognitive behavioral therapy (CBT), but rigorous randomized comparisons against established self-help interventions are scarce, particularly in non-English contexts.
Objective:
This study evaluated the effectiveness and user experience of a hybrid LLM-based CBT conversational agent (chatbot) for depressive and anxiety symptoms in Korean young adults, compared with CBT bibliotherapy as an active control.
Methods:
We conducted a 2-week, 2-arm, parallel-group randomized controlled trial. Korean young adults aged 19 to 29 years who self-reported depressive or anxiety symptoms were randomly assigned to a web-based hybrid LLM-based conversational agent (Mango) or to a CBT-based self-help book (bibliotherapy). The intervention was fully automated, with a single in-person onboarding session and no therapist involvement. The Patient Health Questionnaire-9 (PHQ-9) and Generalized Anxiety Disorder-7 (GAD-7) were self-administered online at baseline (T0), week 1 (T1), and week 2 (T2). All outcomes were analyzed using linear mixed models with an intention-to-treat approach; effect sizes were calculated as repeated-measures Cohen d (dRM) and Hedges g (gPPC2). A corresponding completer analysis served as a sensitivity analysis. Open-ended responses were analyzed thematically. The target sample (n=80) was based on feasibility rather than an a priori power calculation.
Results:
A total of 90 participants were randomized (experimental, n=49; control, n=41); 74 completed all assessments and adhered for at least 7 days. Baseline characteristics, including gender, age, occupation, and initial PHQ-9 or GAD-7 distribution, did not differ between groups. In the intention-to-treat analysis, both groups showed significant reductions from baseline to T2. In the experimental group, mean changes were −2.39 for PHQ-9 (95% CI −3.54 to −1.24; P<.001; dRM=−0.52) and −2.16 for GAD-7 (95% CI −3.07 to −1.26; P<.001; dRM=−0.59). Corresponding mean changes in the control group were −3.63 for PHQ-9 (P<.001; dRM=−0.81) and −3.15 for GAD-7 (P<.001; dRM=−0.83). Between-group differences in change were not statistically significant at either follow-up; at T2, the estimated differences were 1.24 for PHQ-9 (95% CI −0.44 to 2.92; P=.15; gPPC2=0.23) and 0.98 for GAD-7 (95% CI −0.34 to 2.30; P=.15; gPPC2=0.20). Thematic analysis indicated that, despite the absence of significant between-group differences, the 2 formats engaged users through different mechanisms: Mango users described contextualized emotional exploration and a perceived therapeutic alliance alongside the burden of sustained disclosure, whereas bibliotherapy users described declarative CBT learning without personalization.
Conclusions:
A hybrid LLM-based CBT conversational agent did not outperform CBT bibliotherapy over 2 weeks; significant within-group improvement was observed in both conditions in the completer analysis. Because the trial was underpowered and not designed for equivalence testing, these findings do not establish that the CBT conversational agent Mango is as effective as bibliotherapy. The study provides a transparent benchmark of a hybrid LLM architecture against an active comparator and qualitative evidence of distinct engagement mechanisms, informing adequately powered evaluation of future LLM-based interventions. Clinical Trial: Clinical Research Information Service (CRIS) KCT0012309; https://cris.nih.go.kr/cris/search/detailSearch.do?seq=33449
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.