Accepted for/Published in: Journal of Medical Internet Research
Date Submitted: Jan 20, 2026
Date Accepted: Jun 3, 2026
Digital Phenotyping in Health Research: A scoping review of methods, gaps and opportunities
ABSTRACT
Background:
Wearable and connected digital devices generate unprecedented volumes of real-world data, offering opportunities for monitoring, behavior tracking, and personalized care. Digital phenotyping has emerged as a promising paradigm, yet there is little consensus on how such data should be processed or analyzed. Inconsistent or non-explicit methods risk producing biased or misleading conclusions.
Objective:
This scoping review aims to explore the current analytical practices in digital phenotyping health research, highlighting methodological gaps and the urgent need for shared standards.
Methods:
The PubMed database and Google Scholar were used from September 1, 2024 to January 31, 2025 to conduct a review of studies published up to December 31, 2024. Following the Population-Concept-Context approach, eligible studies involved all human subjects in clinical or community contexts, used wearable digital devices in longitudinal health research, and addressed at least one of six methodological domains: sample size planning, variable selection, data cleaning and preprocessing, digital phenotyping techniques, predictive modeling, or statistical handling of big data.
Results:
A total of 128 studies were included, with the majority published after 2018 (82.8%) and conducted in North America (60.9%). The most frequent devices were activity trackers (40.0%), smartphones (22.5%), accelerometers (13.8%) and smartwatches (11.9%). Marked heterogeneity and under-reporting were observed across all methodological domains. Only 27% of studies reported sample size calculations, 16% described variable selection, and preprocessing strategies such as missing-data handling were rarely documented (22%). Digital phenotyping was explicitly applied in just 8.6% of studies, primarily using clustering methods. Predictive models appeared in 22.7% of studies, but validation, calibration, and uncertainty were seldom addressed. Big data specific statistical challenges were rarely discussed (3.1%).
Conclusions:
Digital phenotyping and wearable-based health research are expanding faster than the methodological standards needed to support them. The review highlights a critical methodological vacuum, with analytical choices that are often implicit, inconsistent, or insufficiently reported. By mapping current practices, this work offers researchers a practical foundation to align models with the type of data they collect, to anticipate the risks linked to specific analytical decisions, and to contribute collectively to the development of transparent and reliable guidelines that will strengthen the scientific and clinical value of digital phenotyping.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.