This paper has been accepted and is currently in production.
It will appear shortly on 10.2196/92470
The final accepted version (not copyedited yet) is in this tab.
Optimizing small local language models for culturally competent mental health counseling: comparative evaluation with GPT-4o by psychiatrists
ABSTRACT
Background:
Large language models (LLMs) show promise for mental health counseling applications. However, their deployment faces significant challenges, including data privacy concerns requiring on-premise solutions, high operational costs, and Western-centric biases that limit cultural appropriateness for non-English speaking populations.
Objective:
This study aimed to develop and evaluate culturally competent small language models (SLMs) with approximately 7 to 8 billion parameters for Korean mental health counseling and to validate a 2-stage hybrid evaluation framework combining LLM-as-a-Judge screening with expert clinical validation.
Methods:
We constructed a dataset of 1558 diary-based counseling interactions from a Korean mental health mobile app (Maeumpyeonuijeom). Three Korean-optimized base models (SKT A.X 4.0 Light [7B], LG EXAONE 3.5 [7.8B], and Meta Llama 3.1 [8B]) were fine-tuned using an aggressive fine-tuning strategy with high-rank low-rank adaptation (r=256) over 20 epochs. Inference parameters were optimized using GPT-5 (gpt-5-2025-08-07) as an LLM-as-a-Judge. The optimized SLMs and GPT-4o (gpt-4o-2024-08-06) were evaluated by 10 board-certified psychiatrists (mean clinical experience approximately 10 years) on 50 counseling scenarios across 4 dimensions (individualization, supportiveness, cultural appropriateness, and safety) using 5-point Likert scales. Wilcoxon signed-rank tests with Holm correction were used for statistical comparison, and intraclass correlation coefficients (ICCs) assessed inter-rater reliability.
Results:
The SKT A.X 4.0 Light achieved statistical equivalence with GPT-4o in supportiveness (mean 3.39, SD 0.74 vs mean 3.45, SD 0.90; P=.26) and cultural appropriateness (mean 3.30, SD 0.77 vs mean 3.29, SD 0.79; P=.58). The SKT A.X 4.0 Light ranked first in cultural appropriateness among all models. Inter-rater reliability analysis revealed higher agreement among psychiatrists for local SLMs (ICC=0.575 for SKT A.X 4.0 Light in cultural appropriateness) compared with GPT-4o (ICC=0.211). GPT-4o maintained significantly higher safety scores than all SLMs (P<.001). The LLM-as-a-Judge demonstrated moderate-to-good correlation with human evaluation but exhibited a ceiling effect in safety scoring.
Conclusions:
A 7-8 billion parameter SLM fine-tuned on culturally specific data can achieve performance comparable to GPT-4o in supportiveness and cultural appropriateness for mental health counseling. These findings support the feasibility of locally deployable, privacy-preserving mental health artificial intelligence (AI) systems based on culturally adapted SLMs. However, the safety gap necessitates human expert oversight, and success depends on the availability of high-quality language-specific base models.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.