Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: JMIR Medical Informatics

Date Submitted: Nov 15, 2025
Date Accepted: Jun 29, 2026

The final, peer-reviewed published version of this preprint can be found here:

Persona-Driven Data Augmentation for Disease Name Recognition Across Rare and General Disease Corpora: Comparative Evaluation Study

Pierre JCJ, Nishiyama T, Peng S, Wakamiya S, Aramaki E

Persona-Driven Data Augmentation for Disease Name Recognition Across Rare and General Disease Corpora: Comparative Evaluation Study

JMIR Med Inform 2026;14:e87831

DOI: 10.2196/87831

PMID: 42497410

Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.

Persona-Driven Data Augmentation for Disease Name Recognition

  • Jude Crener Junior Pierre; 
  • Tomohiro Nishiyama; 
  • Shaowen Peng; 
  • Shoko Wakamiya; 
  • Eiji Aramaki

ABSTRACT

Background:

Medical information extraction requires automatically identifying disease names and their associated terms in text. This task, known as Named Entity Recognition (NER), relies on expert-annotated data that are costly to produce and only available in limited quantities. The method of data augmentation (DA) aims to increase available training data, however standard techniques, including synonym substitution and back-translation, sometimes use inappropriate substitutions and do not maintain proper entity-label alignments, which are essential for sequence labeling tasks. Although current Large language models (LLMs) can generate text effectively, they often struggle to maintain factual accuracy.

Objective:

The research investigated how persona-driven, document-level DA using LLM can enhance disease NER performance under low-resources scenarios by rephrasing medical documents into diverse synthetic versions that introduce linguistic bias without altering meaning or entity factuality.

Methods:

We designed a DA framework using multiple personas that vary in medical expertise, personality, tone, and narrative style. Each persona rephrased documents from a medical disease corpus while preserving annotated entity spans through XML-constrained prompting. We measured semantic fidelity and lexical diversity to categorize personas into high-, balanced-, and low-fidelity groups. Models based on bert-base-cased were fine-tuned under four conditions: using (1) gold-standard (GS) data only, (2) single-persona augmentation, (3) persona subsets, and (4) all-persona augmentation. Performance were evaluated by micro-averaged F1-scores, with additional breakdowns by entity type (Disease, RareDisease, Sign, and Symptom).

Results:

Combining GS data with persona-generated texts substantially improved performance. The curated Top-3 mix (P2, P7, P8) raised F1-scores from 68.13 to 71.18, while augmentation with all ten personas achieved 72.29. Entity-level scores improved across categories, particularly Symptoms (from 59.64 to 67.26). Under low-resource settings (60% of GS), persona-augmented models surpassed the baseline trained on full GS data. Error analysis showed reduced confusion among closely related entities.

Conclusions:

Persona-driven augmentation enhances biomedical NER by improving accuracy, reducing entity-type confusion, and requiring less GS data. By preserving span integrity while introducing lexical diversity, the method offers a scalable, robust approach for low-resource biomedical text settings.


 Citation

Please cite as:

Pierre JCJ, Nishiyama T, Peng S, Wakamiya S, Aramaki E

Persona-Driven Data Augmentation for Disease Name Recognition Across Rare and General Disease Corpora: Comparative Evaluation Study

JMIR Med Inform 2026;14:e87831

DOI: 10.2196/87831

PMID: 42497410

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.