Currently accepted at: Journal of Medical Internet Research
Date Submitted: Apr 17, 2026
Date Accepted: Jul 23, 2026
This paper has been accepted and is currently in production.
It will appear shortly on 10.2196/98699
The final accepted version (not copyedited yet) is in this tab.
The Continuity Trap in Data Science Health Research
ABSTRACT
Secondary use is now ordinary in data science health research. Electronic health records collected for care become prediction tools and generative-AI inputs; imaging archives become foundation-model corpora; genomic datasets become resources for polygenic risk scores; and legacy biospecimens become renewable cell lines. Governance has responded by emphasizing verifiable instruments: provenance logs, repository approvals, broad-consent forms, data-use agreements, model cards, records of processing, and locality-preserving architectures. These instruments are necessary, but they are not sufficient. We define ethical continuity as the persistence of normatively relevant relationships between the original conditions of data generation or material collection and subsequent downstream uses, such that current uses remain justifiable in light of the expectations, permissions, meanings, and relational obligations present at entrustment. We define the Continuity Trap as a review-stage governance error in which a salient signal of continuity in one domain is treated as sufficient evidence of ethical continuity overall, causing inquiry into other continuity domains to close prematurely. The trap is not ordinary noncompliance, ethics creep, or a demand for universal re-review. It is a cross-domain inference error. We operationalize ethical continuity across provenance, semantics, authorization, and relational standing. We then apply the framework to consent and non-consent settings, including public health surveillance, immunization registries, syndromic surveillance, wastewater surveillance, polygenic risk scores, induced pluripotent stem cells, federated learning, and health-related large language models. The policy implication is trigger-based continuity review: investigators and reviewers should identify the weakest continuity domain at the present data-stage and impose a domain-matched safeguard.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.