Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Currently submitted to: JMIR Formative Research

Date Submitted: Jul 18, 2026
Open Peer Review Period: Jul 18, 2026 - Sep 12, 2026
(currently open for review)

Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.

De-identification of Turkish Electronic Health Records Using Local Large Language Models: A Privacy-Focused and Open-Source Prototype Development Study

  • Murat KOÇAK; 
  • Rashad ISMAYILOV; 
  • Zafer AKÇALI

ABSTRACT

Background:

Electronic health records (EHRs) contain Protected Health Information (PHI) that poses significant privacy and regulatory risks. Cloud-based large language model (LLM) solutions require data to leave institutional boundaries, raising legal and ethical concerns.

Objective:

This study aimed to develop and validate a fully local de-identification pipeline based on HIPAA Safe Harbor criteria, comparatively evaluating 11 open-source local LLMs on Turkish clinical text.

Methods:

Clinical documents from Turkish healthcare institutions (n=221, 488 pages) were digitized from PDF to Markdown via a vision-language model OCR pipeline and processed through a zero-shot de-identification prompt. A manually annotated gold standard of 11,768 PHI spans across 18 HIPAA entity categories was constructed and used for evaluation. Model performance was assessed using precision, recall, F1 score, and Matthews Correlation Coefficient (MCC).

Results:

Binary PHI F1 scores ranged from 0.56 to 0.96. Qwen3.5-27B-NonThinking achieved the highest performance (Binary F1=0.96, Macro F1=0.917, MCC=0.876), establishing itself as the Pareto-dominant model without extended chain-of-thought reasoning. Four models met the Excellent MCC threshold (>0.75) required for HIPAA-regulated deployment. Structurally regular entities (EMAIL, IP, SSN, URL) were detected consistently across models, while contextually dependent categories (DEVICE, HEALTHPLANID, OTHERID) showed high inter-model variance. The domain-specific MedGemma-27B ranked last (F1=0.56, MCC=0.12), indicating that clinical pre-training does not inherently confer advantages in structured PHI detection. A "Performance Cliff" below rank 8 revealed that MCC collapsed while F1 remained superficially acceptable, underscoring F1's insufficiency as a sole evaluation criterion.

Conclusions:

The developed prototype demonstrates the feasibility of privacy-preserving de-identification of Turkish clinical records using local large language models, offering a practical solution for institutions seeking to comply with data protection requirements while leveraging AI-assisted clinical research. The comparative evaluation of 11 models revealed significant differences in performance across identifier categories, highlighting the importance of model selection and language-specific post-processing in non-English clinical contexts. Future work should focus on fine-tuning models on annotated Turkish EHRs and expanding validation across diverse hospital settings. Clinical Trial: N/A


 Citation

Please cite as:

KOÇAK M, ISMAYILOV R, AKÇALI Z

De-identification of Turkish Electronic Health Records Using Local Large Language Models: A Privacy-Focused and Open-Source Prototype Development Study

JMIR Preprints. 18/07/2026:107369

DOI: 10.2196/preprints.107369

URL: https://preprints.jmir.org/preprint/107369

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.