Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Currently submitted to: Journal of Medical Internet Research

Date Submitted: Jul 20, 2026
Open Peer Review Period: Jul 20, 2026 - Sep 14, 2026
(currently open for review)

Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.

Test-Retest Reliability of Smartwatch-Derived Features for Longitudinal Monitoring: An Observational Cohort Study

  • Hans Lennard Schneidewind; 
  • Tim Fellerhoff; 
  • Jan Franco; 
  • Mitra Tewes; 
  • Caroline Reßing; 
  • Felix Wichum; 
  • Claudia Eickhoff; 
  • Simon B. Eickhoff; 
  • Juergen Dukart

ABSTRACT

Background:

Smartwatches are increasingly used for decentralized data collection in clinical research, but the everyday-life settings that make these data attractive also introduce variability. Before a smartwatch-derived feature can support clinical monitoring, its reproducibility must be established: features with low test-retest reliability weaken associations with clinical outcomes and potentially generate non-actionable signals. Reliability is expected to vary by feature type, aggregation window, and data availability, but has not been systematically screened in a clinical cohort.

Objective:

This study aimed to evaluate the test-retest reliability of candidate smartwatch-derived features for longitudinal monitoring in adults with advanced cancer and in healthy controls, and to determine how reliability depends on the temporal aggregation window.

Methods:

In a prospective single-centre observational cohort study, we analysed 8 weeks of Garmin Vivosmart 5 sensor data from 60 adults with advanced cancer, and 20 healthy controls. Out of 80 participants, 77 contributed usable smartwatch data. We examined 35 daily features across seven domains: heart rate variability, heart rate, respiration, oxygen saturation, sleep, activity, and smartwatch-derived stress. Test-retest reliability was quantified as the intraclass correlation coefficient (ICC(2,1)) across adjacent non-overlapping 1-, 3-, and 7-day windows, with 95% CIs from subject-level bootstrap resampling (10,000 resamples). Between-group and therapy-centred contrasts used permutation testing with Benjamini-Hochberg false discovery rate correction.

Results:

Reliability improved with longer aggregation windows in both cohorts. Between 1-day and 7-day windows, median ICC(2,1) increased from 0.53 to 0.73 in controls and from 0.66 to 0.79 in patients. The number of features reaching good-to-excellent reliability, defined as ICC(2,1)≥0.75, increased from 5 of 35 to 15 of 34 in controls and from 12 of 35 to 26 of 35 in patients. Heart rate and heart rate variability features were the most reliable, with 4 of 11 reaching weekly ICC(2,1)≥0.90 in both cohorts. Activity, respiration, sleep and oxygen-saturation features were more sensitive to aggregation window and data availability, showing larger gains from daily to weekly aggregation (e.g., step count ICC increased from 0.30 to 0.73 in controls). Weekly reliability did not differ significantly between patients and controls (median ΔICC=0.036, no feature survived FDR correction, all q>0.05). No feature showed a significant change in reliability around therapy (median ΔICC=0.016, all q>0.05).

Conclusions:

Weekly aggregation improved the reliability of many smartwatch-derived features, but reliability remained feature specific. Heart rate and heart rate-variability features were consistently reliable, whereas selected sleep and oxygen-saturation features displayed only moderate reliability across all aggregation windows. Reliability was comparable across cohorts and stable around therapy, indicating that feature-wise estimates are transferable across these clinical contexts. Feature-level reliability screening is a prerequisite before smartwatch-derived measures are used in clinical monitoring.


 Citation

Please cite as:

Schneidewind HL, Fellerhoff T, Franco J, Tewes M, Reßing C, Wichum F, Eickhoff C, Eickhoff SB, Dukart J

Test-Retest Reliability of Smartwatch-Derived Features for Longitudinal Monitoring: An Observational Cohort Study

JMIR Preprints. 20/07/2026:107218

DOI: 10.2196/preprints.107218

URL: https://preprints.jmir.org/preprint/107218

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.