Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Currently accepted at: JMIR AI

Date Submitted: Jan 29, 2026
Date Accepted: Sep 17, 2026

This paper has been accepted and is currently in production.

It will appear shortly on 10.2196/92470

The final accepted version (not copyedited yet) is in this tab.

Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.

Optimizing small local language models for culturally competent mental health counseling: comparative evaluation with GPT-4o by psychiatrists

  • Jaewook Han; 
  • Beomgi So; 
  • Whanbo Jung; 
  • Minwoo Kim; 
  • Hoon Kim; 
  • Daun Shin

ABSTRACT

Background:

Large language models (LLMs) show promise for mental health counseling applications. However, their deployment faces significant challenges, including data privacy concerns requiring on-premise solutions, high operational costs, and Western-centric biases that limit cultural appropriateness for non-English speaking populations.

Objective:

This study aimed to develop and evaluate culturally competent small language models (SLMs) with approximately 7 to 8 billion parameters for Korean mental health counseling and to validate a 2-stage hybrid evaluation framework combining LLM-as-a-Judge screening with expert clinical validation.

Methods:

We constructed a dataset of 1558 diary-based counseling interactions from a Korean mental health mobile app (Maeumpyeonuijeom). Three Korean-optimized base models (SKT A.X 4.0 Light[7B], LG EXAONE 3.5 [7.8B], and Meta Llama 3.1[8B]) were fine-tuned using a hyperfitting strategy with high-rank low-rank adaptation (r=256) over 20 epochs. Inference parameters were optimized using GPT-5(gpt-5-2025-08-07) as an LLM-as-a-Judge. The optimized SLMs and GPT-4o(gpt-4o-2024-08-06) were evaluated by 10 board-certified psychiatrists (mean clinical experience approximately 10 years) on 50 counseling scenarios across 4 dimensions (individualization, supportiveness, cultural appropriateness, and safety) using 5-point Likert scales. Wilcoxon signed-rank tests with Holm correction were used for statistical comparison, and intraclass correlation coefficients (ICCs) assessed inter-rater reliability.

Results:

The SKT A,X 4.0 Light model achieved statistical equivalence with GPT-4o in supportiveness (mean 3.39, SD 0.74 vs mean 3.45, SD 0.90; P=.26) and cultural appropriateness (mean 3.30, SD 0.77 vs mean 3.29, SD 0.79; P=.58). The SKT A,X 4.0 Light model ranked first in cultural appropriateness among all models. Inter-rater reliability analysis revealed higher agreement among psychiatrists for local SLMs (ICC=0.575 for SKT in cultural appropriateness) compared with GPT-4o (ICC=0.211). GPT-4o maintained significantly higher safety scores than all SLMs (P<.001). The LLM-as-a-Judge demonstrated moderate-to-good correlation with human evaluation but showed a tendency to overestimate safety scores.

Conclusions:

A 7-8 billion parameter SLM fine-tuned on culturally specific data can achieve performance comparable to GPT-4o (estimated 1.76 trillion parameters) in supportiveness and cultural appropriateness for mental health counseling. This represents a parameter efficiency gain of more than 99%, enabling on-premise deployment for enhanced privacy protection. However, the safety gap necessitates human expert oversight, and success depends on the availability of high-quality language-specific base models.


 Citation

Please cite as:

Han J, So B, Jung W, Kim M, Kim H, Shin D

Optimizing small local language models for culturally competent mental health counseling: comparative evaluation with GPT-4o by psychiatrists

JMIR Preprints. 29/01/2026:92470

DOI: 10.2196/preprints.92470

URL: https://preprints.jmir.org/preprint/92470

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.