Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: JMIR Formative Research

Date Submitted: Sep 28, 2025
Date Accepted: Jul 21, 2026

The final, peer-reviewed published version of this preprint can be found here:

Fine-Tuning Large Language Models for Structured Extraction of Infectious Disease–Related Information From Clinical Notes in Japanese Primary Care: Development and Internal Validation Study

Yoshihara H, Maeda H, Hagiwara Y, Sato D, Kitajima K, Iwata A, Van de Velde N, Nakamura Y, Yamagishi Y, Igarashi A

Fine-Tuning Large Language Models for Structured Extraction of Infectious Disease–Related Information From Clinical Notes in Japanese Primary Care: Development and Internal Validation Study

JMIR Form Res 2026;10:e84974

DOI: 10.2196/84974

PMID: 42721099

Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.

Fine-tuning Large Language Models for Structured Extraction of Infectious Disease-related Information from Clinical Notes in Japanese Primary Care: Development and Validation Study

  • Hiroshi Yoshihara; 
  • Haruka Maeda; 
  • Yuriko Hagiwara; 
  • Daichi Sato; 
  • Kei Kitajima; 
  • Akihiro Iwata; 
  • Nicolas Van de Velde; 
  • Yuta Nakamura; 
  • Yosuke Yamagishi; 
  • Ataru Igarashi

ABSTRACT

Background:

The Coronavirus Disease 2019 (COVID-19) pandemic highlighted the importance of timely infectious disease surveillance.

Objective:

We aimed to develop a natural language processing algorithm to extract structured information on symptoms and vaccination history from free-text clinical notes in Japanese primary care. This approach enables low-latency monitoring using real-world electronic health record data.

Methods:

A total of 773 clinical notes, originating from 526 unique patients, were provided by M3 Inc. through the Japan Medical Data Survey and used for analysis. Three clinicians annotated information related to infectious disease symptoms and vaccination history. The data were divided into 622 training cases and 151 evaluation cases, ensuring no patient overlap between the two sets. We developed a rule-based extraction algorithm using clinician-designed regular expressions and conducted few-shot learning and fine-tuning with both open-source and commercial large language models (LLMs). We compared the extraction accuracy (F1 score) of these approaches.

Results:

Generally, the format of clinical notes tends to be similar across clinicians, allowing rule-based algorithms to achieve reasonable extraction accuracy. LLMs using few-shot prompts could extract unstructured information, such as vaccination history, with high accuracy, achieving high performance across all evaluation metrics (Anthropic Claude 3.5 Sonnet, macro-average F1 score 0.902). We also demonstrated that a fine-tuned small open-source LLM can achieve remarkably higher accuracy (Google Gemma 2 2B, macro-average F1 score 0.932) than commercial models.

Conclusions:

We showed that a fine-tuned LLM can accurately extract and structure infectious disease-related information from free-text clinical notes, which implies the feasibility of digital surveillance using electronic medical record information.


 Citation

Please cite as:

Yoshihara H, Maeda H, Hagiwara Y, Sato D, Kitajima K, Iwata A, Van de Velde N, Nakamura Y, Yamagishi Y, Igarashi A

Fine-Tuning Large Language Models for Structured Extraction of Infectious Disease–Related Information From Clinical Notes in Japanese Primary Care: Development and Internal Validation Study

JMIR Form Res 2026;10:e84974

DOI: 10.2196/84974

PMID: 42721099

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.