Currently submitted to: JMIR Medical Education
Date Submitted: Sep 30, 2026
Open Peer Review Period: Oct 5, 2026 - Nov 30, 2026
(currently open for review)
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Identification of Stigmatizing and Biased Language in Mental Health Clinical Notes: A Classification Analysis using LLMs
ABSTRACT
Background:
Stigmatizing language in clinical documentation can negatively affect patient-centered care and contribute to bias in health care. Large language models are increasingly used for processing clinical notes, but their ability to identify stigmatizing language and their behavior across different linguistic and demographic contexts remain insufficiently understood.
Objective:
This study aimed to evaluate the ability of Large Language Models to identify stigmatizing language in clinical notes by analyzing classification of predefined stigmatizing instances against a gold standard annotation and independently identifying and extracting representative spans of stigmatizing language. The study also examined patterns in false negatives and assessed whether model classification performance varied across patient demographic groups.
Methods:
Three LLMs (Gemini-2.5-Flash, Phi-4, Qwen2.5-7B) were evaluated using 750 discharge summaries from MIMIC-IV. In contextual classification, models were provided with a candidate stigmatized keyword or phrase along with surrounding context and were asked to classify the candidate instance as either Stigmatizing or Non Stigmatizing. Classification performance was assessed using accuracy, precision, recall and F1 score. Pairwise model differences were evaluated using McNemar tests and classification performance was also examined across gender, age and race groups. False negative cases were further examined to identify recurring linguistic and contextual patterns. In LLM span detection and extraction, models were provided with the entire clinical note and instructed to identify a single, smallest contiguous span that could independently be classified as Stigmatizing Language. Model generated spans were independently assessed by a human annotator, and the agreement rate, disagreement rate, span hallucinations count, and span length were examined.
Results:
In contextual classification, Gemini-2.5-Flash achieved the highest performance, with an accuracy of 71.86% and F1-score of 71.74%, compared with 55.60% and 54.90% for Qwen2.5-7B and 49.86% and 48.10% for Phi-4, respectively. Pairwise comparisons showed statistically significant differences between all 3 models. Across demographic groups, all models generally demonstrated substantially higher accuracy for Not Stigmatizing Language than Stigmatizing Language. False-negative analysis identified recurring difficulties involving contextual distinctions, word forms, prefixes, and tense. In LLM span detection and extraction, human annotator agreement with model-generated spans was 18.36% for Gemini-2.5-Flash, 9.09% for Phi-4, and 5.57% for Qwen2.5-7B post removal of hallucinated spans. Models frequently generated spans that the human annotator did not consider stigmatizing.
Conclusions:
LLMs demonstrated an ability to identify some stigmatizing language in clinical documentation, but their sensitivity to stigmatizing language was limited and varied across models and demographic groups. Analysis of false negatives revealed recurring linguistic and contextual patterns associated with missed instances. Independent span generation and human assessment further indicated that models may identify language as stigmatizing when it does not meet human annotation criteria. These findings highlight the need for careful evaluation of LLMs before their use in automated detection of stigmatizing language in clinical documentation.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.