Currently submitted to: JMIR Medical Informatics
Date Submitted: Sep 8, 2026
Open Peer Review Period: Sep 15, 2026 - Nov 10, 2026
(currently open for review)
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Entity Recognition of the Age-Friendly Health System 4Ms in Nursing Home Text Messages Using Large Language Model Revision of Bio-ClinicalBERT Candidate Entities: Development and Evaluation Study
ABSTRACT
Background:
Text messages exchanged among care teams in nursing homes contain information aligned with the Age-Friendly Health Systems 4Ms (What Matters, Medication, Mentation, and Mobility) that need to be extracted for downstream uses. A very limited number of evaluated approaches exist in this domain for extracting 4M information from nursing home text messages, with a most recent and major approach using LLM fine-tuning. An approach that constrains locally deployable LLMs to revision only of encoder-detected candidate spans remains unexplored.
Objective:
This study aimed to develop and evaluate a 4M Entity Recognition (4M-ER) pipeline that combines encoder-based token classification with LLM revision by locally deployable, open-weight instruction-tuned LLMs, to identify 4M entities in nursing home text messages without additional LLM fine-tuning.
Methods:
An expert-annotated dataset of 1,169 messages from 16 Midwest nursing homes was used for development and evaluation. The pipeline uses a fine-tuned Bio-ClinicalBERT token classifier to identify candidate spans, then an LLM revision step (guided by semantically retrieved in-context exemplars) corrects boundaries, evaluates labels, and accepts or rejects candidates. Four 4M-ER variants (Gemma, Phi, Qwen and Mistral-Nemo) were compared against zero-shot LLM extraction, standalone fine-tuned Bio-ClinicalBERT, and a previously published fine-tuned Gemma LLM.
Results:
Evaluating via entity-type F1 score, to ensure semantically accurate assessment rather than strict boundary rigidity and partial-matching leniency, the 4M-ER variants performed better compared to the fine-tuned Gemma with significant improvement for What Matters: Gemma, Qwen, and Mistral-Nemo variants increased What Matters F1 by 0.114, 0.108, and 0.094 respectively (95% bootstrap CIs excluding). Mentation and Mobility F1 scores were generally higher than those of the fine-tuned Gemma, but not significantly. The fine-tuned Gemma performed better for Medication, significantly outperforming the Phi variant of the pipeline by 0.062 and also had higher secondary strict and partial F1 scores.
Conclusions:
The 4M-ER pipeline provides a locally deployable approach to 4M entity recognition in text messages through LLM revision of fine-tuned Bio-ClinicalBERT candidate entities without LLM fine-tuning with significant advantage for What Matters entity recognition.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.