Currently accepted at: Journal of Medical Internet Research
Date Submitted: Jan 19, 2026
Date Accepted: Jul 7, 2026
This paper has been accepted and is currently in production.
It will appear shortly on 10.2196/91756
The final accepted version (not copyedited yet) is in this tab.
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Neurosift: Quality Assurance for Multimedia Data in Automated Movement Disorder Assessment
ABSTRACT
Background:
Movement disorders affect millions globally, yet access to specialists remains limited. Automated multimedia analysis offers a scalable path for screening and remote monitoring, but unsupervised recordings often suffer from quality issues that compromise model reliability.
Objective:
This study presents NeuroSift, a machine learning framework designed to automatically assess recording quality and task compliance in home-based multimedia data collected for movement disorder assessment.
Methods:
Three experts annotated 2,516 recordings corresponding to three common movement disorder tasks: finger tapping, facial expression, and speech. Based on these annotations, we developed task-specific quality guidelines, designed interpretable features, and trained quality classification models. Inter-rater agreement was evaluated before and after guideline implementation. Interpretable, task-specific features were extracted from each modality, and ordinal classification models were trained to categorize recordings as poor, borderline, or good quality. Model performance was evaluated against expert consensus, and feature importance analyses were conducted to identify common sources of quality degradation.
Results:
The quality guidelines significantly improved inter-rater agreement among experts (???? < 0.001). Across tasks, the quality classification models achieved 77–90% ordinal accuracy against expert consensus when classifying recordings as poor, borderline, or good quality. Feature importance analysis further identifies task-specific failures such as hand visibility, camera angle, or background noise.
Conclusions:
NeuroSift is intended to strengthen trust in predictive models and support reliable deployment of multimedia-based neurological assessment tools at scale. By identifying low-quality or noncompliant recordings prior to downstream analysis, the framework would enhance the robustness, trustworthiness, and scalability of machine learning-based neurological assessment tools. Clinical Trial: None.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.