Accepted for/Published in: Journal of Medical Internet Research
Date Submitted: Jan 14, 2026
Date Accepted: Aug 5, 2026
Large Language Models for Distress Rating in Korean Psycho-Oncology Interviews: Exploratory Clinician-Benchmarked Evaluation Study
ABSTRACT
Background:
Psychiatric distress is common among patients with cancer, yet systematic interview-based screening remains difficult to scale in routine clinical care. Large language models (LLMs) have shown promise as scalable tools for mental health assessment, but most existing evidence is derived from clinician-authored records, translated text, or proxy data. The performance characteristics, error patterns, and explanatory behaviors of contemporary LLMs when applied to authentic, non-English psychiatric interviews remain insufficiently characterized.
Objective:
This study aimed to evaluate the reliability, directional bias, and error mechanisms of contemporary LLMs when reproducing psycho-oncologists’ item-level distress symptom ratings from real-world Korean psycho-oncology interviews, using clinician benchmarks and adjudicated error taxonomy.
Methods:
Between April 2024 and May 2025, 101 adults receiving active cancer treatment in South Korea underwent semi-structured interviews. Board-certified psycho-oncologists provided real-time, time-stamped ratings of 39 items. Generative pre-trained transformer 4o (GPT-4o), Claude 3.5, and Gemini 2.5 generated item scores and brief rationales using identical Korean zero-shot rubrics. Concordance between clinician ratings was evaluated using ordinal and binary screening metrics. A paired Wilcoxon test compared patient-level total symptom burden (binary flag sum). Binary mismatches were clinically adjudicated as ambiguous or definite overestimation/underestimation with an eight-etiology taxonomy. Model-generated rationales were further meta-evaluated by GPT-5 across four dimensions: citation (use of quoted supporting statements), structure (logical organization of the rationale), mapping (consistency between the rationale and the assigned item rating), and expansion (degree of interpretive elaboration beyond the explicit transcript content). The associations between meta-evaluation results and absolute error were examined using cross-classified mixed-effects models.
Results:
Across 3931-item ratings, all models showed good agreement with clinicians (intraclass correlation coefficient: 0.816–0.872), while GPT performing the highest. Screening performance was strong (F1: 0.775–0.837). Claude and Gemini yielded significantly higher patient-level symptom burden (p <0.001 and p =0.002, respectively), whereas GPT did not (p=0.145). Clinician adjudication attributed 29%–39% of mismatches to intrinsic ambiguity in patient speech. Among definite errors, misapplication of severity thresholds was the predominant mechanism across models. Meta-evaluation revealed that greater interpretive expansion, and to a lesser extent higher citation density, were consistently associated with larger absolute error.
Conclusions:
Contemporary LLMs can approximate psycho-oncologists’ item-level ratings in authentic Korean psycho-oncology interviews, but exhibit differences in error profiles. These findings highlight the limitations of relying solely on aggregate performance metrics and emphasize the importance of clinician-benchmarked, error-aware evaluation when designing and deploying LLM-based tools for psychiatric assessment in real-world clinical settings.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.