Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: JMIR Research Protocols

Date Submitted: Apr 1, 2026
Open Peer Review Period: Apr 2, 2026 - May 28, 2026
Date Accepted: Jul 22, 2026
(closed for review but you can still tweet)

The final, peer-reviewed published version of this preprint can be found here:

Extraction of Pain Severity and Functional Interference From Clinical Narratives Using Domain-Informed Large Language Models: Protocol for a Development and Validation Study

Zhang X, Wilson C, Eyre H, Reed DE II, Hee-Wai TY, Kloehn A, Rosser EW, Luo G, Zeliadt SB

Extraction of Pain Severity and Functional Interference From Clinical Narratives Using Domain-Informed Large Language Models: Protocol for a Development and Validation Study

JMIR Res Protoc 2026;15:e96346

DOI: 10.2196/96346

PMID: 42842900

Extracting Pain Severity and Functional Interference from Clinical Narratives Using Domain-Informed Large Language Models: Study Protocol

  • Xiaoyi Zhang; 
  • Chris Wilson; 
  • Hannah Eyre; 
  • David E. Reed II; 
  • Travis Y. Hee-Wai; 
  • Alexander Kloehn; 
  • Ethan Wan Rosser; 
  • Gang Luo; 
  • Steven Bacchus Zeliadt

ABSTRACT

Background:

Chronic pain is a leading cause of disability and requires multidimensional assessment of pain intensity and functioning, yet electronic health records (EHRs) rarely capture these measures systematically. By contrast, surveys collecting patient-reported outcomes can assess pain over multiple dimensions but remain resource-intensive and difficult to scale for continuous population-level monitoring.

Objective:

The objective of this study is to develop and validate a domain-informed natural language processing (NLP) framework to derive pain severity and functional interference outcomes from unstructured clinical narratives. We aim to demonstrate that NLP-derived outcomes can serve as a reliable, scalable surrogate for resource-intensive patient-reported surveys.

Methods:

This study utilizes a retrospective cohort of 3,726 Veterans with chronic musculoskeletal pain initiating Complementary and Integrative Health (CIH) therapies across 18 Veterans Health Administration (VA) Whole Health Flagship sites (2021–2023). The dataset encompasses longitudinal patient-reported outcome surveys serving as the benchmark, linked with unstructured clinical narratives from the VA EHR. Guided by established psychometric instruments and subject matter expert (SME) input, we developed a seed lexicon and annotation guidelines to identify language distinguishing three pain domains: pain severity, interference with enjoyment of life, and interference with general activities. Preliminary LLM prompting was used to identify 600 candidate encounters (200 per domain) from 6,747 notes across 260 patients for SME annotation, forming a ground truth validation sample. Two candidate LLMs will be evaluated on this sample; the best-performing LLM will generate a large library of span-level annotations to train a scalable, lightweight language model. The study employs a three-stage validation process: (1) documentation completeness of pain interference in clinical narratives against SME-annotated references; (2) inference accuracy of the LLM-as-annotator and the fine-tuned lightweight model against SME annotations across note-level classification and span-level localization; and (3) concordance between the lightweight model's output and patient-reported PEG scores across a range of temporal windows.

Results:

To date, the cohort of 3,726 Veterans has been identified and linked to clinical notes. The seed lexicon and annotation guidelines have been developed. Applying a developmental LLM to screen 6,747 text notes in 6,642 unique encounters over a 7-month period for 260 patients, at least one of the three pain domains was identified in 75% of notes and 99% of patients. SME validation at the encounter level is in progress.

Conclusions:

This protocol outlines a framework for identifying severe pain intensity and interference from clinical narratives, addressing a critical gap in health care system surveillance. To our knowledge, this is the first study to validate clinical text-based pain outcome extraction against patient-reported outcomes in a nationwide longitudinal cohort. If successful, this approach will enable health care systems to continuously monitor reports of pain-related functional interference and support more holistic, patient-centered pain management at scale.


 Citation

Please cite as:

Zhang X, Wilson C, Eyre H, Reed DE II, Hee-Wai TY, Kloehn A, Rosser EW, Luo G, Zeliadt SB

Extraction of Pain Severity and Functional Interference From Clinical Narratives Using Domain-Informed Large Language Models: Protocol for a Development and Validation Study

JMIR Res Protoc 2026;15:e96346

DOI: 10.2196/96346

PMID: 42842900

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.