Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: JMIR Medical Informatics

Date Submitted: Sep 30, 2025
Date Accepted: Aug 13, 2026

The final, peer-reviewed published version of this preprint can be found here:

Semantic Similarity Search Approach to Extract Exemplars of Stigmatizing and Positive Language in Obstetric Clinical Notes: Exploratory Study

Scroggins JK, Hulchafo II, Barcelona V, Davoudi A, Moen H, Harkins S, Scharp D, Cato K, Tadiello M, Topaz M

Semantic Similarity Search Approach to Extract Exemplars of Stigmatizing and Positive Language in Obstetric Clinical Notes: Exploratory Study

JMIR Med Inform 2026;14:e85088

DOI: 10.2196/85088

PMID: 42771877

Semantic Similarity Search Approach to Extract Exemplars of Stigmatizing and Positive Language in Obstetric Clinical Notes: An Exploratory Study

  • Jihye Kim Scroggins; 
  • Ismael Ibrahim Hulchafo; 
  • Veronica Barcelona; 
  • Anahita Davoudi; 
  • Hans Moen; 
  • Sarah Harkins; 
  • Danielle Scharp; 
  • Kenrick Cato; 
  • Michele Tadiello; 
  • Maxim Topaz

ABSTRACT

Background:

Natural language processing (NLP) can extract meaningful information from clinical notes. However, human annotation is time-consuming, costly, and scarce data poses a challenge.

Objective:

To explore a semantic similarity search approach to extract exemplars of stigmatizing and positive language in obstetric clinical notes.

Methods:

We used electronic health records data from labor and birth admissions at two hospitals in the United States in 2017. We employed a semantic similarity search approach, which used human-annotated exemplars as queries to search for similar exemplars. We extracted the top five candidates with the highest cosine similarities for 200 stratified and randomly selected true exemplars, which were assessed for accuracy.

Results:

We retrieved 1,000 candidates. Average precision was acceptable (≥ 0.7) when candidates with cosine similarity thresholds of 0.75 or higher were included, at which point 69% of candidates accurately represented true cases. Precision was highest (100%) in the following categories: autonomy for birth, power/privilege, questioning patient credibility, and disapproval. Precision was lowest for difficult patient (29.57%) and marginalized identities (33.33%) categories.

Conclusions:

The semantic similarity search approach shows promise in efficiently extracting exemplars while reducing the annotation burden, laying the groundwork for future applications in other domains. Clinical Trial: This is an observational study and does not require trial registration.


 Citation

Please cite as:

Scroggins JK, Hulchafo II, Barcelona V, Davoudi A, Moen H, Harkins S, Scharp D, Cato K, Tadiello M, Topaz M

Semantic Similarity Search Approach to Extract Exemplars of Stigmatizing and Positive Language in Obstetric Clinical Notes: Exploratory Study

JMIR Med Inform 2026;14:e85088

DOI: 10.2196/85088

PMID: 42771877

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.