Accepted for/Published in: JMIR Medical Informatics
Date Submitted: Apr 8, 2026
Open Peer Review Period: Apr 13, 2026 - May 11, 2026
Date Accepted: Jul 8, 2026
Date Submitted to PubMed: Aug 13, 2026
(closed for review but you can still tweet)
Quantifying the Impact of Anonymization-Induced Clinical Data Quality Loss: A Methodological Case Study Using Primary Diagnosis Codes and Hospital Length of Stay
ABSTRACT
Background:
Secondary use of electronic health record data requires robust privacy protection. k-Anonymity is widely used to enable data sharing by ensuring that each quasi-identifier combination occurs in at least k records, yet its analytical impact on clinically meaningful structures remains insufficiently characterized, particularly for the combination of record suppression and microaggregation that arises when a numeric attribute lacks a natural generalization hierarchy. A further gap is that anonymization tools report internal information-loss values but do not signal the downstream distributional and inferential distortions these transformations introduce.
Objective:
This study evaluated the analytical footprint of k-anonymity at k = 5, 10, and 15 on two core data elements in retrospective hospital research: primary ICD-10-GM diagnosis codes and hospital length of stay (LOS). It aimed to determine and quantify whether anonymization introduces meaningful distortions not captured by the anonymization tool itself, and whether diagnosis-specific LOS patterns remain reproducible after anonymization.
Methods:
We analyzed 719,387 inpatient encounters from University Hospital Mannheim from 2010 to 2024. Anonymization was performed with the ARX tool. It used record suppression and microaggregation. Distributional distortion was assessed with the Kolmogorov-Smirnov (KS) D statistic, quantile shifts, interquartile range changes, and tail changes. Categorical fidelity was assessed with the Jaccard coefficient and Cramer's V. Inferential reproducibility was assessed with a three-level linear mixed model. The model included random intercepts for ICD-3 and patient. We compared the intraclass correlation coefficient (ICC) and diagnosis-level effect concordance. Concordance was quantified using Spearman rho and Lin's concordance correlation coefficient (CCC), both with 95% confidence intervals. A composite traffic-light verdict summarized the results.
Results:
ARX masked quasi-identifier cells rather than deleting rows; the proportion of encounters with a masked cell rose from 0.77% (k=5) to 2.62% (k=15), distributed almost uniformly across admission years. KS D was stable at 0.147. Median LOS shifted by one day and the standard deviation declined by about 6.5 days, while the diagnosis-level mean changed modestly. Jaccard overlap fell from 0.624 to 0.421. The diagnosis ICC rose from 0.294 to 0.837, reflecting variance compression rather than improved signal. Best linear unbiased prediction (BLUP) rank concordance (Spearman rho 0.964-0.970) and aggregate magnitude agreement (Lin CCC 0.959-0.966) were high, yet about 7% of low-signal diagnoses showed sign reversals. Most distributional change occurred at k=5.
Conclusions:
k-Anonymity preserved the ranking of diagnosis-specific LOS effects but altered distributional shape, individual effect magnitudes, and diagnostic vocabulary, none of which was flagged by the internal loss metric. Anonymized data of this type may support ordinal analyses but can mislead analyses requiring faithful variance structure, accurate absolute effects, or complete rare-diagnosis representation. A reporting checklist is provided to document these effects.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.