Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Currently accepted at: Journal of Medical Internet Research

Date Submitted: Jan 8, 2026
Open Peer Review Period: Feb 16, 2026 - Apr 13, 2026
Date Accepted: Jul 22, 2026
(closed for review but you can still tweet)

This paper has been accepted and is currently in production.

It will appear shortly on 10.2196/91089

The final accepted version (not copyedited yet) is in this tab.

Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.

Stigmatizing Language in Gender-Expansive Patient Records: Corpus, Disparity Analysis, and NLP-based Detection

  • Liyang Xue; 
  • Mary Chayko; 
  • Vivek Kumar Singh

ABSTRACT

Background:

Stigmatizing language (SL) in electronic health records (EHRs) can influence clinical decision-making, propagate bias across care encounters, and undermine patient trust. Gender-expansive patients (GEPs) may be particularly vulnerable to documentation-based stigma, yet large-scale quantitative evidence and fairness-aware evaluation of automated SL detection methods remain limited.

Objective:

To quantify stigmatizing language in clinical documentation for gender-expansive patients by introducing a labeled corpus, analyzing demographic disparities, and evaluating fairness-aware natural language processing (NLP) methods for SL detection.

Methods:

We developed a corpus of 780 discharge summaries from a large academic health system, annotated for SL and its subtypes. Notes were categorized by GEP versus non-GEP status. We conducted logistic regression to assess associations between GEP status and SL presence, adjusting for demographics. Multiple NLP models, including transfer learning approaches, were benchmarked for SL detection. We implemented fairness-aware thresholding to reduce subgroup performance gaps.

Results:

SL appeared in 61.9% of GEP notes compared to 26.1% of non-GEP notes. After adjustment, GEP status remained a significant predictor of SL (odds ratio > 4). Baseline NLP models exhibited subgroup disparities, with high performance gaps in accuracy, true positive rate, and false positive rates between GEP and non-GEP patients. Our fairness-aware thresholding approach reduced error disparities (ΔFPR from 21.16% to 6.65% and ΔTPR from 21.11% to 0.00% with minimal accuracy loss), while maintaining overall accuracy.

Conclusions:

Stigmatizing language is common in EHR documentation and disproportionately affects gender-expansive patients, with automated detection models showing persistent subgroup performance gaps. This study introduces the first annotated corpus focused on SL in GEP documentation, quantifies demographic disparities, and demonstrates practical fairness-aware NLP strategies that can reduce error-rate inequities while preserving accuracy. These findings support equity-focused interventions and inform digital health workflows aimed at reducing stigmatizing language in EHRs.


 Citation

Please cite as:

Xue L, Chayko M, Singh VK

Stigmatizing Language in Gender-Expansive Patient Records: Corpus, Disparity Analysis, and NLP-based Detection

JMIR Preprints. 08/01/2026:91089

DOI: 10.2196/preprints.91089

URL: https://preprints.jmir.org/preprint/91089

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.