Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: Journal of Medical Internet Research

Date Submitted: May 26, 2026
Date Accepted: Aug 18, 2026
Date Submitted to PubMed: Aug 18, 2026

The final, peer-reviewed published version of this preprint can be found here:

Representational Veracity in Data Science Health Research: Targets, Proxies, Labels, and Descriptors

Adebamowo C, Adebamowo SN, Akintola A, Ikhane P, Akintola S, Ogundiran T, Jegede A, Adeyemo O, Callier S, Imam-Tamim MK, Uthman I, BridgELSI Project as part of DSI Africa Consortium BPapoDAC

Representational Veracity in Data Science Health Research: Targets, Proxies, Labels, and Descriptors

J Med Internet Res 2026;28:e102537

DOI: 10.2196/102537

PMID: 42610557

Representational Veracity in Data Science Health Research: Targets, Proxies, Labels, and Descriptors

  • Clement Adebamowo; 
  • Sally N Adebamowo; 
  • Adeola Akintola; 
  • Peter Ikhane; 
  • Simisola Akintola; 
  • Temidayo Ogundiran; 
  • Ayodele Jegede; 
  • Olusegun Adeyemo; 
  • Shawneequa Callier; 
  • Muhammad K Imam-Tamim; 
  • Ibrahim Uthman; 
  • BridgELSI Project as part of DSI Africa Consortium BridgELSI Project as part of DSI Africa Consortium

ABSTRACT

Prevailing ethical oversight of data science health research concentrates on privacy, consent, bias, and fairness. These concerns are necessary but insufficient, because each presupposes an answer to a prior question that is seldom asked directly. That is “do the targets, proxies, labels, classifications, ontologies, and population descriptors on which a current study rests still truthfully represent the persons, populations, and phenomena they are taken to describe, at the point of use rather than the point of collection?” In this viewpoint we name that question representational veracity (RV) and develop it as a construct for upstream ethical review. Our aims are to define RV and derive the domains along which it can be assessed; to demonstrate that it asks something that measurement validity, critical data studies, and algorithmic fairness do not; and to translate it into instruments that review bodies can use. We derive four assessment domains of RV analytically, asking for each transition in the data journey what must remain stable for a stored artifact still to stand for what it originally stood for. The resulting domains are material provenance, informational descriptors, normative authorization, and relational community. These domains interact but do not substitute for one another. Intact provenance cannot repair a poorly chosen target, and a transparent labeling process cannot confer authorization it never had. Drawing on scholarship in quantification, classification, measurement, critical data studies, algorithmic fairness, and health artificial intelligence governance, we show that a model may be accurate, reproducible, and formally fair while resting on a representation that is too thin, too unstable, or too normatively misdirected for the proposed use. We examine four recurrent failure modes, proxy substitution, category misassignment, label generation error, and descriptor sedimentation, anchoring each in a published case, and we present a counterpoint in which better representation reveals rather than conceals inequity. A polygenic risk score (PRS) case study illustrates all four domains and shows how a score can misclassify risk in the populations least represented in its derivation while its code, pipeline, and internal validation statistics remain intact. We then translate the framework into practice using ten reviewer prompts that an editor can paste into a review form, a justification template and scoring rubric provided as appendices, a tiered model that triggers full review only for subgroup, equity, transportability, public health, or clinical implementation claims, and a graded account of what should follow an adverse finding. Our argument is that existing governance mechanisms require an upstream layer. Investigators should be asked to justify not only whether their models perform, but whether their representations are truthful enough for the claims at hand. The intended audience is investigators, informaticians, research ethics committees, institutional review boards, data access committees, funders, regulators, and journal editors.


 Citation

Please cite as:

Adebamowo C, Adebamowo SN, Akintola A, Ikhane P, Akintola S, Ogundiran T, Jegede A, Adeyemo O, Callier S, Imam-Tamim MK, Uthman I, BridgELSI Project as part of DSI Africa Consortium BPapoDAC

Representational Veracity in Data Science Health Research: Targets, Proxies, Labels, and Descriptors

J Med Internet Res 2026;28:e102537

DOI: 10.2196/102537

PMID: 42610557

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.