Currently submitted to: JMIR Medical Informatics
Date Submitted: Aug 27, 2026
Open Peer Review Period: Sep 4, 2026 - Oct 30, 2026
(currently open for review)
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Methods to validate structured electronic health record data from emergency departments
ABSTRACT
Background:
Electronic health records (EHRs) can help emergency departments conduct disease surveillance, evaluate and improve patient care, as well as accelerate health services research and system learning. However, emergency providers face unique workflow and time constraints that make documentation prone to data errors, inconsistencies, and omissions. Comprehensive validation is required to ensure resulting data is suitable for research use.
Objective:
We applied six data quality dimensions to evaluate and refine electronically extracted emergency department EHR data.
Methods:
We extracted EHR data from twelve British Columbia hospitals from July 2025 to March 2026. We used computational and manual assessments to evaluate the conformance, completeness, plausibility, correctness, concordance, and currency of 37 key variables. Computational assessments measured variable standardization, unexpected missingness, distribution, and relationships between related variables. Research assistants completed independent and blinded manual record reviews to create the criterion standard in which electronically extracted variables were evaluated using kappa scores, correlation coefficients, and mean similarity rates. Physicians provided consultations throughout.
Results:
We verified the data of 8,323 patients who made 34,035 emergency department visits. We identified 35 data issues total from all assessments; 18 related to the source data and 17 related to data extraction. Three variables had poor coverage (missingness >80%) and were omitted from the dataset. Timestamps did not always reflect the sequence of care. Obvious data entry errors were rare (<1% of most variables). Most electronically extracted variables had very strong (>95%) agreement with manually extracted variables.
Conclusions:
Iterative data quality assessments in six data quality dimensions allowed us to improve the extraction process and provided an understanding of data accuracy and limitations. Manual chart reviews and physician consultations were critical for this process and identified important limitations of the extracted data. Future standardized assessment frameworks should account for medical specialty, workflow, and documentation practices. Clinical Trial: N/A
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.