Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: JMIR Medical Informatics

Date Submitted: Aug 13, 2025
Date Accepted: Mar 6, 2026

The final, peer-reviewed published version of this preprint can be found here:

From Data Poverty to Data Sovereignty: Operationalizing Gold-Standard Biomedical Datasets in Low- and Middle-Income Countries

Rana S, Bhagat P, Pal D, Pati S, Singh H

From Data Poverty to Data Sovereignty: Operationalizing Gold-Standard Biomedical Datasets in Low- and Middle-Income Countries

JMIR Med Inform 2026;14:e82360

DOI: 10.2196/82360

PMID: 42574669

PMCID: 13456305

From Data Poverty to Data Sovereignty: Operationalising Gold-Standard Biomedical Datasets in LMICs

  • Shweta Rana; 
  • Pranshu Bhagat; 
  • Debnath Pal; 
  • Sanghamitra Pati; 
  • Harpreet Singh

ABSTRACT

The rapid proliferation of artificial intelligence (AI) in clinical medicine contrasts sharply with the persistent underrepresentation of low- and middle-income countries (LMICs) in AI datasets. Current FDA-approved AI tools predominantly utilize data from high-income nations, with fewer than 4% reporting geographic or racial diversity. This geographic skew leads to significant performance degradation when used in LMIC populations, perpetuating an equity crisis where health burdens are highest. Addressing this disparity, we highlight Medical Imaging Datasets for India (MIDAS) initiative as a viable model to transition LMICs from "data poverty" to "data sovereignty." MIDAS employs a rigorous, four-domain Dataset Quality Matrix to ensure representativeness, documentation, technical fidelity, and governance, creating openly benchmarked, gold-standard datasets tailored to local contexts. Initial releases, including datasets for oral and dural lesions, demonstrate feasibility and practical value for developing robust, generalizable AI models. Further, we propose a multilateral South-South Data Commons structured around three foundational pillars: a harmonized dataset-grading rubric, distributed custodial governance, and outcome-linked incentives. This infrastructure supports local stewardship, encourages global collaboration, and ensures quality benchmarks drive financial incentives for dataset expansion and diversity. This proposed framework not only positions LMICs as autonomous data stewards but also enhances global AI equity. By institutionalizing quality control, interoperability, and outcome accountability, LMICs can transform from passive data consumers into active contributors of essential, trustworthy, and globally relevant biomedical datasets.


 Citation

Please cite as:

Rana S, Bhagat P, Pal D, Pati S, Singh H

From Data Poverty to Data Sovereignty: Operationalizing Gold-Standard Biomedical Datasets in Low- and Middle-Income Countries

JMIR Med Inform 2026;14:e82360

DOI: 10.2196/82360

PMID: 42574669

PMCID: 13456305

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.