Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Currently accepted at: JMIR AI

Date Submitted: Jan 29, 2026
Date Accepted: Sep 17, 2026

This paper has been accepted and is currently in production.

It will appear shortly on 10.2196/92470

The final accepted version (not copyedited yet) is in this tab.

Optimizing small local language models for culturally competent mental health counseling: comparative evaluation with GPT-4o by psychiatrists

  • Jaewook Han; 
  • Beomgi So; 
  • Whanbo Jung; 
  • Minwoo Kim; 
  • Hoon Kim; 
  • Daun Shin

ABSTRACT

Background:

Large language models (LLMs) show promise for mental health counseling applications. However, their deployment faces significant challenges, including data privacy concerns requiring on-premise solutions, high operational costs, and Western-centric biases that limit cultural appropriateness for non-English speaking populations.

Objective:

This study aimed to develop and evaluate culturally competent small language models (SLMs) with approximately 7 to 8 billion parameters for Korean mental health counseling and to validate a 2-stage hybrid evaluation framework combining LLM-as-a-Judge screening with expert clinical validation.

Methods:

We constructed a dataset of 1558 diary-based counseling interactions from a Korean mental health mobile app (Maeumpyeonuijeom). Three Korean-optimized base models (SKT A.X 4.0 Light [7B], LG EXAONE 3.5 [7.8B], and Meta Llama 3.1 [8B]) were fine-tuned using an aggressive fine-tuning strategy with high-rank low-rank adaptation (r=256) over 20 epochs. Inference parameters were optimized using GPT-5 (gpt-5-2025-08-07) as an LLM-as-a-Judge. The optimized SLMs and GPT-4o (gpt-4o-2024-08-06) were evaluated by 10 board-certified psychiatrists (mean clinical experience approximately 10 years) on 50 counseling scenarios across 4 dimensions (individualization, supportiveness, cultural appropriateness, and safety) using 5-point Likert scales. Wilcoxon signed-rank tests with Holm correction were used for statistical comparison, and intraclass correlation coefficients (ICCs) assessed inter-rater reliability.

Results:

The SKT A.X 4.0 Light achieved statistical equivalence with GPT-4o in supportiveness (mean 3.39, SD 0.74 vs mean 3.45, SD 0.90; P=.26) and cultural appropriateness (mean 3.30, SD 0.77 vs mean 3.29, SD 0.79; P=.58). The SKT A.X 4.0 Light ranked first in cultural appropriateness among all models. Inter-rater reliability analysis revealed higher agreement among psychiatrists for local SLMs (ICC=0.575 for SKT A.X 4.0 Light in cultural appropriateness) compared with GPT-4o (ICC=0.211). GPT-4o maintained significantly higher safety scores than all SLMs (P<.001). The LLM-as-a-Judge demonstrated moderate-to-good correlation with human evaluation but exhibited a ceiling effect in safety scoring.

Conclusions:

A 7-8 billion parameter SLM fine-tuned on culturally specific data can achieve performance comparable to GPT-4o in supportiveness and cultural appropriateness for mental health counseling. These findings support the feasibility of locally deployable, privacy-preserving mental health artificial intelligence (AI) systems based on culturally adapted SLMs. However, the safety gap necessitates human expert oversight, and success depends on the availability of high-quality language-specific base models.


 Citation

Please cite as:

Han J, So B, Jung W, Kim M, Kim H, Shin D

Optimizing small local language models for culturally competent mental health counseling: comparative evaluation with GPT-4o by psychiatrists

JMIR AI. 17/09/2026:92470 (forthcoming/in press)

DOI: 10.2196/92470

URL: https://preprints.jmir.org/preprint/92470

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.