Currently submitted to: Journal of Medical Internet Research
Date Submitted: Sep 17, 2026
Open Peer Review Period: Sep 17, 2026 - Nov 12, 2026
(currently open for review)
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Uncoded Clinical Features from Multilingual Electronic Health Records in Catalonia: Development and Validation Study
ABSTRACT
Background:
Unstructured free‑text narratives in electronic health records (EHRs) contain critical clinical information that is not captured by structured standard codes. In Catalonia, primary care notes follow a semi‑structured format known as MEAP (in Catalan). However, extracting structured variables using cloud‑hosted commercial large language models (LLMs) raises substantial data privacy concerns and incurs high operational costs.
Objective:
To develop and validate a privacy-preserving local hybrid pipeline combining a lightweight small language model (SLM) with deterministic regular expressions for extracting six uncoded urinary tract infection (UTI) clinical features (fever, nitrites, leukocytes, lumbar pain, abdominal pain, and haematuria) from primary care EHR narratives
Methods:
We deployed the open‑weight model (Phi4-mini, 3.8B parameters) within the institutional firewall using Ollama. Prompts were iteratively refined in collaboration with clinicians, and regular expressions were applied post‑inference. Validation was conducted across two independent arms: (1) a double‑blind clinician gold standard review of 60 real patient records (consensus‑resolved cases), and (2) an adversarial synthetic dataset of 720 notes enriched with linguistic noise, generated using GPT‑4.1 and Grok‑4.1. Point estimates and 95% confidence intervals (CIs) were computed for key performance metrics.
Results:
The pipeline processed 15,498 MEAP narratives from a matched cohort of 2,962 primary care patients and identified 4,663 clinical feature occurrences. Patients progressing to acute pyelonephritis (cases) presented a higher features burden than non-progressing controls (71.3% vs. 55.8%; standardized mean difference [SMD] = 0.326), particularly for fever (33.0% vs. 9.0%; SMD = 0.616) and lumbar pain (29.0% vs. 9.8%; SMD = 0.502). In real-world clinician validation, overall performance yielded 93.8% accuracy (95% CI 90.6%-96.1%), 83.6% sensitivity (95% CI 73.0%-91.2%), 96.6% specificity (95% CI 93.6%-98.4%), 87.1% positive predictive value (PPV; 95% CI 77.0%-93.9%), and 95.5% negative predictive value (NPV; 95% CI 92.3%-97.7%). In synthetic stress-testing, the pipeline demonstrated 87.2% accuracy (95% CI 84.6%-89.6%), 73.9% sensitivity (95% CI 68.9%-78.5%), near-perfect specificity (99.5%, 95% CI 98.1%-99.9%), and 99.2% PPV (95% CI 97.2%-99.9%).
Conclusions:
A privacy-preserving local hybrid framework combining a compact open-weight SLM with regular expression rules achieves high specificity and precision for extracting six uncoded clinical features from multilingual primary care narratives. Running entirely within institutional servers, this approach keeps patient data secure, meets privacy standards, and avoids Application Programming Interface (API) costs, making it a practical tool for EHR research, though with lower sensitivity for narratively complex descriptions.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.