Previously submitted to: Journal of Medical Internet Research (no longer under consideration since May 14, 2025)
Date Submitted: Apr 12, 2025
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Breaking the Data Bottleneck in Exercise Health AI
ABSTRACT
Artificial intelligence (AI) holds significant potential to transform exercise health through applications like personalized training, injury prevention, and targeted rehabilitation. However, the scarcity of high-quality, real-world data—due to issues like data fragmentation, inconsistent formats, privacy constraints, and confidentiality concerns—severely restricts AI development and deployment in this domain. Learning from advancements in healthcare AI, strategies such as federated learning for privacy-preserving collaboration, synthetic data generation to augment limited datasets, transfer learning to leverage existing models, and structured data-sharing protocols offer viable solutions. Furthermore, an innovative closed-loop data generation process is proposed: leveraging large language model APIs to create extensive synthetic datasets from initial high-quality inputs, followed by rigorous human expert validation and refinement through real-world feedback. This paper argues that overcoming the data bottleneck necessitates an integrated strategy combining these methods. Implementing such comprehensive measures will enhance AI model robustness and generalizability, ultimately accelerating the delivery of tangible benefits to athletes and health practitioners.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.