Currently submitted to: JMIR Formative Research
Date Submitted: Sep 11, 2026
Open Peer Review Period: Oct 7, 2026 - Dec 2, 2026
(currently open for review)
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Extracting Family Health History of Alcohol Use Disorder From Clinical Notes: Development and Evaluation of a Rule-Based NLP Pipeline
ABSTRACT
Background:
Utilizing family health history (FHH) to identify at-risk patients before they develop alcohol use disorder (AUD) enables early intervention, improving overall health. Clinical documentation of family members’ alcohol use behaviors in electronic health records frequently occurs in the narrative text in clinical notes.
Objective:
This study assesses the feasibility of utilizing natural language processing (NLP) methods to extract FHH of AUD from clinical notes, aiming to establish a pathway for early AUD risk identification.
Methods:
Notes were obtained retrospectively from a large medical center for patients ages 18-45, with either FHH of alcohol or substance use disorder diagnosis codes (FHH_AOSUD) or AUD diagnosis codes or medication prescription (N=1,388 patients with 575,156 notes). We developed a rule-based NLP algorithm using Python’s MedSpaCy NLP library to identify family member AUD in clinical notes using keywords and terms from previous studies, combined with MedSpaCy’s section identifier. This pipeline pre-filtered the notes for manual annotation and analysis. Two cohorts were annotated: 419 notes from patients that had a FHH_AOSUD diagnosis code and NLP-identified mentions of family members and alcohol use in the FHH section, and 200 randomly selected notes from patients that had an AUD diagnosis code without a FHH_AOSUD diagnosis. Precision, recall, F1 and accuracy were used for evaluation.
Results:
The MedSpaCy section identifier successfully returned the portion of text with a FHH section. Rule-based NLP performed well on the FHH_AOSUD cohort (F1=0.91, P=0.84, R=1.0, Accuracy=0.99), but worse on the random AUD cohort (F1=0.25, P=0.16, R=0.63, Accuracy=0.85). Error arose from false positives where content adjacent to the FHH section, such as social history, was included, and false negatives where alcohol terms were missing from the rule-base. An analysis of note format and content showed only 191 of 419 FHH segments contained unique annotatable content, and 31% of those were formatted as lists. Male family members appeared significantly more often in FHH_AUD mentions than in the surrounding context compared to female family members (OR=2.90, 95% CI:2.04-4.20).
Conclusions:
Our work shows MedSpaCy captures the required FHH information, but a more sophisticated method is needed for classification. Future work will require pipeline components that better utilizes context to improve classification accuracy, such as large language models.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.