Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: JMIR Formative Research

Date Submitted: Nov 2, 2025
Open Peer Review Period: Nov 3, 2025 - Dec 29, 2025
Date Accepted: Jun 27, 2026
(closed for review but you can still tweet)

The final, peer-reviewed published version of this preprint can be found here:

Measuring Depression Severity With Clinical Global Impression–Severity Scale Scores From Clinical Notes Using Large Language Models: Validation Study

Li K, Zirikly A, Collica SC, Goes FS, Zhao C, Nguyen T, Gagliardi JP, Goldstein BA, Hong H, Stuart EA, Zandi PP

Measuring Depression Severity With Clinical Global Impression–Severity Scale Scores From Clinical Notes Using Large Language Models: Validation Study

JMIR Form Res 2026;10:e86906

DOI: 10.2196/86906

PMID: 42573558

Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.

Large Language Models to Estimate CGI-S From Clinical Notes as a Measure of Depression Severity

  • Kevin Li; 
  • Ayah Zirikly; 
  • Sarah C. Collica; 
  • Fernando S. Goes; 
  • Congwen Zhao; 
  • Trang Nguyen; 
  • Jane P. Gagliardi; 
  • Benjamin A. Goldstein; 
  • Hwanhee Hong; 
  • Elizabeth A. Stuart; 
  • Peter P. Zandi

ABSTRACT

Background:

Real-world psychiatric care is marked by wide heterogeneity in clinical presentations and outcomes, underscoring the need for systematic approaches to outcome measurement. The Clinical Global Impression–Severity (CGI-S) scale is a brief, clinician-rated measure of overall illness severity widely used in psychiatric research but rarely documented in routine care. Large language models (LLMs) may enable automated extraction of CGI-S scores from narrative clinical notes providing scalable outcome measures for real-world clinical care and research.

Objective:

This study aimed to evaluate whether LLMs can estimate CGI-S scores from psychiatric clinical notes in patients with major depressive disorder (MDD) by first generating a clinician consensus gold standard dataset, and then comparing model-generated scores for validation.

Methods:

We used data from the Johns Hopkins electronic health record. Three psychiatrists independently rated 77 clinical notes using a validated depression-specific CGI rubric. Weighted Cohen’s kappa (κ) coefficients were calculated to assess interrater reliability and model–human agreement. Two prompting strategies, zero-shot and few-shot, were tested using GPT-4o, and agreement was compared against average human ratings. Exploratory analyses evaluated whether agreement varied by patient demographics, care setting, or note length.

Results:

Interrater reliability among psychiatrists was high (κ = 0.77–0.78). Agreement between model-generated and average human ratings was similarly strong (κ = 0.85) and was even higher for notes on which all three raters were in complete agreement (κ = 0.88). Weighted κ values remained consistently high across all subgroups (0.82–0.89), with no significant differences by age, sex, race, treatment location, or note length.

Conclusions:

LLMs can accurately estimate clinician-rated CGI-S scores from psychiatric clinical notes, achieving reliability comparable to expert raters. This approach may enable scalable outcome measurement and support the implementation of measurement-based care in real-world psychiatric practice.


 Citation

Please cite as:

Li K, Zirikly A, Collica SC, Goes FS, Zhao C, Nguyen T, Gagliardi JP, Goldstein BA, Hong H, Stuart EA, Zandi PP

Measuring Depression Severity With Clinical Global Impression–Severity Scale Scores From Clinical Notes Using Large Language Models: Validation Study

JMIR Form Res 2026;10:e86906

DOI: 10.2196/86906

PMID: 42573558

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.