Accepted for/Published in: JMIR Formative Research
Date Submitted: Nov 2, 2025
Open Peer Review Period: Nov 3, 2025 - Dec 29, 2025
Date Accepted: Jun 27, 2026
(closed for review but you can still tweet)
Measuring Depression Severity With CGI-S Scores From Clinical Notes Using Large Language Models: Validation Study
ABSTRACT
Background:
Real-world psychiatric care is marked by wide heterogeneity in clinical presentations and outcomes, underscoring the need for systematic approaches to outcome measurement. The Clinical Global Impression–Severity (CGI-S) scale is a brief, clinician-rated measure of overall illness severity widely used in psychiatric research but rarely documented in routine care. Large language models (LLMs) may enable automated extraction of CGI-S scores from narrative clinical notes providing scalable outcome measures for real-world clinical care and research.
Objective:
This study aimed to evaluate whether LLMs can estimate CGI-S scores from psychiatric clinical notes in patients with major depressive disorder (MDD) by first generating a clinician consensus gold standard dataset, and then comparing model-generated scores for validation.
Methods:
We used data from the Johns Hopkins electronic health record. Three psychiatrists independently rated 77 clinical notes using a validated depression-specific CGI rubric. Weighted Cohen’s kappa (κ) coefficients were calculated to assess interrater reliability and model–human agreement. Two prompting strategies, zero-shot and few-shot, were tested using GPT-4o, and agreement was compared against average human ratings. Exploratory analyses evaluated whether agreement varied by patient demographics, care setting, or note length.
Results:
Interrater reliability among psychiatrists was high (κ = 0.77–0.78). Agreement between model-generated and average human ratings was similarly strong (κ = 0.85) and was even higher for notes on which all three raters were in complete agreement (κ = 0.88). Weighted κ values remained consistently high across all subgroups (0.82–0.89), with no significant differences by age, sex, race, treatment location, or note length.
Conclusions:
LLMs can accurately estimate clinician-rated CGI-S scores from psychiatric clinical notes, achieving reliability comparable to expert raters. This approach may enable scalable outcome measurement and support the implementation of measurement-based care in real-world psychiatric practice.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.