Accepted for/Published in: Journal of Medical Internet Research
Date Submitted: May 26, 2026
Date Accepted: Aug 18, 2026
Date Submitted to PubMed: Aug 18, 2026
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Representational Veracity in Data Science Health Research: A Companion Viewpoint on Targets, Proxies, Labels, and Descriptors
ABSTRACT
Prevailing ethical oversight of data science health research concentrates on privacy, consent, bias, and fairness. These concerns are necessary but insufficient, because they presuppose that the targets, proxies, labels, classifications, ontologies, and population descriptors already truthfully represent the populations and phenomena they claim to describe. That prior question is often left unexamined. In this Viewpoint, we introduce representational veracity (RV) as a companion concept to the Continuity Trap (CT) framework, grounded in a persistence-based account of representational adequacy developed in the parent framework manuscript. The Continuity Trap names the governance error in which a salient signal of continuity in one domain - technical, administrative, or infrastructural - is treated as sufficient evidence that ethical continuity has been preserved across all domains. In contrast, RV asks whether stored data continue to accurately and responsibly represent the persons, communities, and phenomena from which they were derived — at the point of use, not merely the point of collection. We draw on scholarship in quantification, classification, measurement, critical data studies, algorithmic fairness, and health artificial-intelligence (AI) governance to show that a model may be accurate, reproducible, and formally fair while resting on a representation that is too thin, unstable, or normatively misdirected for the proposed use. Our analysis proceeds through the four-domain representational-veracity architecture (material provenance, informational descriptors, normative authorization, relational community), examines four recurrent failure modes (proxy substitution, category misassignment, label-generation error, descriptor sedimentation), and applies these to a polygenic risk score (PRS) use case. Our argument is not that existing governance mechanisms are wrong but that they need an additional upstream layer. Investigators should be required to justify not only whether their models perform, but also whether their representations are truthful enough for the clinical, public health, or policy claims at hand.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.