Accepted for/Published in: JMIR Formative Research
Date Submitted: Feb 3, 2025
Date Accepted: May 14, 2026
Lessons from Building a Large, Public, HIV-Related Database in Support of Ending the HIV Epidemic Initiative
ABSTRACT
The HIV epidemic remains a national priority within the United States (U.S.), and the Ending the HIV Epidemic (EHE) Initiative has renewed the call for expanded prevention and treatment strategies capable of reducing new HIV infections 90% by 2030. Data are crucial for understanding HIV-related needs, barriers to care, and the effectiveness of interventions. However, while several publicly available datasets exist, few integrate multiple domains such as HIV outcomes, social determinants of health (SDOH), and community-level factors. The lack of unified data and difficulty linking datasets across these domains hampers efforts to tailor HIV management and treatment strategies. As a result, existing datasets often remain siloed and difficult to analyze in combination, limiting potential for comprehensive research. Integrating data across geographic and thematic strata offers significant potential to strengthen our understanding of community factors driving HIV prevention and treatment. In this viewpoint, we describe our experience building a unified compilation of publicly available HIV and community data to identify factors influencing HIV outcomes and interventions. The completed database comprises 242 variables drawn from eight public sources mapped across clinic, ZIP code, county, and state geographies. The build required approximately 350 total project hours (~250 core build hours for data sourcing, quality control, and cleaning, plus ~100 hours for meetings and project management) and revealed an initial spot-check error rate of approximately 33%--a rate we attribute primarily to reliance on manual data entry. Our central argument is that teams undertaking similar builds should: (1) invest heavily in source identification before construction begins; (2) adopt automated data engineering tools from the outset rather than defaulting to manual entry; (3) establish a shared data dictionary before the first variable is entered; and (4) build quality control throughout the workflow. By sharing the approach used to develop this database and making the final resource publicly accessible through the Yale Center for Methods in Implementation and Prevention Science (CMIPS), we aim to reduce barriers to data access and encourage similar data integration efforts. Consolidating HIV, SDOH, and contextual variables into a unified data source is a critical step towards enabling deeper, more comprehensive analysis and supporting ongoing efforts to end the HIV epidemic in the EHE era.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.