Accepted for/Published in: Journal of Medical Internet Research
Date Submitted: Mar 13, 2026
Date Accepted: Aug 10, 2026
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Leveraging generative Large Language Models for temporal relation extraction from French clinical narratives
ABSTRACT
Background:
Temporal relation extraction in clinical narratives is crucial for understanding patient history, disease progression, and treatment pathways. However, it remains challenging due to limited annotated data and to the complexity of clinical text, including domain-specific terminology, inconsistent information, and implicit temporal reasoning.
Objective:
In this work, we explore the potential of open-weight and on-premises Large Language Models (LLMs) on clinical temporal relation extraction and normalization using zero- and few-shot prompting, aiming to reduce the time-consuming and costly annotation process. Based on recent promising capabilities of LLMs in understanding and reasoning over text, we evaluate whether these models can efficiently perform temporal relation extraction and normalization in real-world clinical settings.
Methods:
We cast the temporal relation task as a question-answering task, in which the models extract the temporal expressions associated with a given clinical event. We propose a prompt-chaining strategy that sequentially performs a normalization on the extracted temporal expressions. Our LLM-based approach is evaluated on constructed and annotated French real-word clinical narratives, covering temporal relations involving two types of clinical events: phenotypes and rare disease diagnoses. Four open-weight LLMs were evaluated across multiple prompt configurations and compared with rule-based and neural baselines using exact-match metrics, an LLM-as-a-judge evaluation, and human validation. We further evaluate our approach by applying it on the English 2012 i2b2 corpus, involving other types of clinical events.
Results:
Results show that the proposed approach achieves stable extraction performance with only a small number of in-context examples across both clinical event types, reaching a maximum F-measure of 0.72 for rare disease diagnoses and 0.61 for phenotype events. LLM-as-a-judge evaluation supports these findings by capturing minor temporal variations penalized by exact-match metrics, while human validation further confirms the clinical validity of the extracted relations. On the 2012 i2b2 corpus, the approach remains robust despite increased complexity and minimal supervision. Adding normalization using the prompt-chaining strategy maintains stable overall performance, indicating minimal error propagation and an efficient temporal normalization.
Conclusions:
Our results suggest that prompting LLMs with a minimal set of in-context examples can achieve strong performance for temporal relation extraction and normalization in low-resource settings with limited domain-specific annotations. The proposed approach could be extended to extract key clinical information, such as age at diagnosis or age at onset, particularly in the context of rare diseases.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.