Accepted for/Published in: JMIR Formative Research
Date Submitted: Jul 13, 2025
Open Peer Review Period: Aug 25, 2025 - Oct 20, 2025
Date Accepted: May 14, 2026
(closed for review but you can still tweet)
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Reconstruction of Time-Series Data from ECG PDFs: A Lightweight and Accurate Approach for Clinical Applications
ABSTRACT
Background:
Electrocardiogram (ECG) data are critical for clinical decision-making and cardiology research. However, many ECG systems only provide graphical outputs in PDF format, limiting access to raw time-series data. Existing solutions, such as image processing and deep learning-based digitization, are computationally intensive, error-prone, and often fail to handle diverse ECG formats or noisy data.
Objective:
This study aims to develop and validate a lightweight, accurate method for reconstructing near-original ECG time-series data directly from PDF files, without relying on image processing, deep learning, or intermediate SVG conversion.
Methods:
We utilized ECG PDFs generated by the MUSE system, focusing on the 12-lead ECG format. Using PdfFileAnalyzer, we directly parsed path objects embedded in the PDF files to extract waveform coordinate data. Each lead's signal was reconstructed by mapping the x- and y-coordinates into time and voltage values, calibrated against the reference 1 mV marker within each PDF. The method was validated using 5,000 randomly selected ECG PDFs with available original time-series data as ground truth. We evaluated reconstruction accuracy using mean absolute error (MAE), root mean square error (RMSE), and Pearson correlation coefficients.
Results:
The reconstructed ECG waveforms demonstrated high fidelity to the original signals, with an average MAE of ~0.008, RMSE of ~0.01, and Pearson correlation coefficients consistently approaching 1.0 across all 12 leads. Each reconstructed signal contained exactly 5,000 data points, matching the resolution of the ground truth. The complete dataset was processed in 36 minutes on a standard desktop computer with an Intel i7 CPU and 4 GB RAM, highlighting the method's computational efficiency. Attempts to extract signals via PDF-to-SVG conversion were unsuccessful in multiple cases, further supporting the robustness of the proposed direct parsing method.
Conclusions:
This study introduces a novel, resource-efficient approach for reconstructing ECG time-series data directly from PDF files by parsing vector path objects. The method provides near-perfect reconstruction accuracy without the need for complex computational infrastructure or large training datasets. It enables scalable ECG data extraction for clinical diagnostics and research, facilitating broader ECG data accessibility and promoting innovations in cardiovascular care.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.