Accepted for/Published in: Journal of Medical Internet Research
Date Submitted: Jan 15, 2026
Date Accepted: Jul 17, 2026
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Comparative Evaluation of Large Language Models and Quantum-Enhanced Hybrid Architectures for Medical Text Anonymization
ABSTRACT
Background:
The widespread adoption of Electronic Health Records (EHRs) has generated large-scale repositories of highly sensitive clinical information, emphasizing the need for robust anonymization strategies to enable secondary use for research while safeguarding patient privacy. Conventional rule-based and machine learning approaches for de-identifying medical text face limitations with the linguistic complexity, variability, and context dependence inherent to clinical documentation. Recent advances in Large Language Models (LLMs), combined with emerging quantum computing paradigms, present novel opportunities to enhance the accuracy, scalability, and resilience of healthcare data anonymization.
Objective:
This study aims to evaluate the efficacy of LLM-based and quantum-enhanced hybrid architectures for medical text anonymization, assessing the effectiveness and computational efficiency across multiple entity types in Portuguese clinical notes.
Methods:
We constructed a gold-standard corpus of 1,000 Portuguese outpatient clinical notes, manually annotated by five domain experts for five protected-entity categories: patient names, dates, identifiers, organizations, and geographic locations. Four anonymization strategies were evaluated: two standalone LLMs (Llama-3.1-8B-instruct and Llama-3.3-70B-instruct) and two quantum-enhanced hybrid models (Dynex-QML with 8B and 70B base models) incorporating quantum optimization via Quadratic Unconstrained Binary Optimization (QUBO) formulations. The quantum-enhanced approach transforms the final attention layer of the LLM into a global constraint satisfaction problem solved via neuromorphic quantum annealing. Model performance was measured on held-out test set of 500 notes using precision, recall, and F1-score metrics. Computational efficiency was quantified through end-to-end processing time.
Results:
The quantum-enhanced Dynex-QML-70B model achieved the highest overall performance with a macro-F1 score of 0.855, outperforming the standalone Llama-3.3-70B (0.726), Dynex-QML-8B (0.733), and Llama-3.1-8B (0.602). Across all entity categories, Dynex-QML-70B demonstrated consistently superior precision, with gains for patient names (0.971 vs 0.624 for Llama-70B) and dates (0.960 vs 0.919). In terms of computational efficiency, the quantum-enhanced 70B model reduced inference time compared to the standalone Llama-70B (7.97±2.87s vs 8.52±3.32s per text).
Conclusions:
Quantum-enhanced hybrid architectures provide substantial improvements in medical text anonymization accuracy compared to standalone LLMs, particularly reductions in false positive rates while preserving high sensitivity. The Dynex-QML-70B model achieved the best balance between performance and efficiency, suggesting that quantum-enhanced optimization offers a strategy for high-fidelity, scalable, and regulation-compliant in de-identification of clinical text. These findings highlight the potential of emerging quantum–AI paradigms to advance the secondary secure use of healthcare data. Clinical Trial: Independent Institutional Review Board (CAPPesq – IRB No. 68, Brazilian National Research Ethics System), under registration number CAAE 89528325.8.0000.0068.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.