Currently submitted to: JMIR Medical Informatics
Date Submitted: Aug 14, 2026
Open Peer Review Period: Aug 21, 2026 - Oct 16, 2026
(currently open for review)
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Large language and rule-based models for diabetes diagnosis date extraction from electronic medical records : comparative study
ABSTRACT
Background:
Diabetes duration is a major determinant of complications and an essential variable for clinical research, yet the diagnosis date is frequently unavailable in structured electronic health records (EHRs). Although large language models (LLMs) have shown promise for clinical information extraction, their added value over established rule-based approaches for this task remains unclear.
Objective:
This study aimed to compare the performance and computational efficiency of prompting-based LLMs and a rule-based extraction method for identifying diabetes diagnosis dates from French clinical notes.
Methods:
Clinical notes of patients with diabetes (600 type 1 and 100 type 2) were extracted from the Assistance Publique-Hôpitaux de Paris Clinical Data Warehouse and manually annotated. Notes of type 1 diabetes were divided into training (n=200), validation (n=200) and test (n=200) sets; notes of type 2 were only used as test set. We compared a rule-based ContextualMatcher with five open-weight LLMs (Qwen3-8B, LLaMA-3.1-8B-Instruct, Ministral-8B-Reasoning, Ministral-14B-Reasoning, and MedGemma-27B) using zero-shot and few-shot prompting, with reasoning enabled when supported. Predictions within ±1 year of the reference annotation were considered correct. Performance was evaluated using F1-score, precision, recall, balanced accuracy, specificity, hallucination rate, no-prediction rate, mean absolute error (MAE), and inference time.
Results:
On the type 1 diabetes test set, MedGemma-27B achieved the highest F1-score (0.94, 95% CI 0.91–0.96), followed by Qwen3-8B with reasoning (0.93, 95% CI 0.90–0.96), compared with 0.86 (95% CI 0.82–0.90) for the best ContextualMatcher. LLMs showed the greatest advantage for relative diagnosis dates (best F1 1.00 vs 0.82). Performance generalized to type 2 diabetes despite model development being conducted exclusively on type 1 notes, with MedGemma-27B achieving an F1-score of 0.96 and the best ContextualMatcher 0.89. Few-shot prompting did not consistently improve performance, and dedicated reasoning models performed inconsistently. The rule-based approach processed notes in 0.01–0.02 seconds using CPUs only, whereas LLM inference required 0.2–2.0 seconds per note and two NVIDIA A100 GPUs. Qwen3-8B with reasoning required approximately 50-fold longer inference time than the best ContextualMatcher.
Conclusions:
Prompting-based LLMs achieved the highest accuracy for diabetes diagnosis date extraction from French clinical notes but provided only modest performance gains over a well-tuned rule-based approach while requiring substantially greater computational resources. For large-scale EHR research, rule-based methods remain an efficient and competitive option, whereas LLMs may be most valuable when maximizing extraction accuracy for linguistically heterogeneous expressions justifies the additional computational cost.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.