Accepted for/Published in: Journal of Medical Internet Research
Date Submitted: May 6, 2023
Date Accepted: May 13, 2024
Subtyping Social Determinants of Health in All of Us: Network Analysis and Visualization Approach
ABSTRACT
Background:
Social determinants of health (SDoH), such as financial resources and housing stability, account for between 30-55% of people’s health outcomes. While many studies have identified strong associations among specific SDoH and health outcomes, most people experience multiple SDoH that impact their daily lives. Analysis of this complexity requires the integration of personal, clinical, social, and environmental information from a large cohort of individuals that have been traditionally underrepresented in research, which is only recently being made available through the All of Us research program. However, little is known about the range and response of SDoH in All of Us, and how they co-occur to form subtypes, which are critical for designing precision interventions.
Objective:
To address two research questions: (1) What is the range and response to survey questions related to SDoH in the All of Us dataset? (2) How do SDoH co-occur to form subtypes, and what are their risk for adverse health outcomes?
Methods:
For Question-1, an expert panel analyzed the range of SDoH questions across the surveys with respect to the 5 domains in Healthy People 2030 (HP-30), and analyzed their responses across the full All of Us data (n=372,397, V6). For Question-2, we used the following steps: (1) due to the missingness across the surveys, selected all participants with valid and complete SDoH data, and used inverse probability weighting to adjust their imbalance in demographics compared to the full data; (2) an expert panel grouped the SDoH questions into SDoH subdomains for enabling a more consistent granularity; (3) used bipartite modularity maximization to identify SDoH biclusters, their significance, and their replicability; (4) measured the association of each bicluster to three outcomes (depression, delayed medical care, emergency room visits in the last year) using multiple data types (surveys, electronic health records, and zip codes mapped to Medicaid expansion states); and (5) the expert panel inferred the subtype labels, potential mechanisms that precipitate adverse health outcomes, and interventions to prevent them.
Results:
For Question-1, we identified 110 SDoH questions across 4 surveys, which covered all 5 domains in HP-30. However, the results also revealed a large degree of missingness in survey responses (1.76%-84.56%), with later surveys having significantly fewer responses compared to earlier ones, and significant differences in race, ethnicity, and age of participants of those that completed the surveys with SDoH questions, compared to those in the full All of Us dataset. Furthermore, as the SDoH questions varied in granularity, they were categorized by an expert panel into 18 SDoH subdomains. For Question-2, the subtype analysis (n=12,913, d=18) identified 4 biclusters with significant biclusteredness (Q=0.13, random-Q=0.11, z=7.5, P<0.001), and significant replication (Real-RI=0.88, Random-RI=0.62, P<.001). Furthermore, there were statistically significant associations between specific subtypes and the outcomes, and with Medicaid expansion, each with meaningful interpretations and potential precision interventions. For example, the subtype Socioeconomic Barriers included the SDoH subdomains not employed, food insecurity, housing insecurity, low income, low literacy, and low educational attainment, and had a significantly higher odds ratio (OR=4.2, CI=3.5-5.1, P-corr<.001) for depression, when compared to the subtype Sociocultural Barriers. Individuals that match this subtype profile could be screened early for depression and referred to social services for addressing combinations of SDoH such as housing insecurity and low income. Finally, the identified subtypes spanned one or more HP-30 domains revealing the difference between the current knowledge-based SDoH domains, and the data-driven subtypes.
Conclusions:
The results revealed that the SDoH subtypes not only had statistically significant clustering and replicability, but also had significant associations with critical adverse health outcomes, which had translational implications for designing targeted SDoH interventions, and for designing decision-support systems. Furthermore, these SDoH subtypes spanned multiple SDoH domains defined by HP-30 revealing the complexity of SDoH in the real-world, and aligning with influential SDoH conceptual models such as by Dahlgren-Whitehead. Finally, the code for our subtyping pipeline is currently available on the All of Us workbench consisting of generalizable and scalable machine learning methods, which can be used to periodically rerun the analysis as the All of Us dataset grows for analyzing subtypes related to SDoH and beyond.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.