NIH’s All of Us Program: Broadening Research Diversity and Reducing Health Disparities

TL/DR –

The All of Us Research Program, funded by the National Institutes of Health, aims to increase diversity in biomedical research participants and reduce health disparities by collecting data from individuals historically underrepresented in such research. As part of its data collection process, the program administers surveys to participants to gather demographic and socioeconomic information, such as race, ethnicity, income, and education. The program also collected data from electronic health records and genetic information, which was used in combination with the survey data to analyze the influence of social determinants on health outcomes, focusing on nine chronic conditions with high prevalence, cost, and medical actionability.


All of Us Research Program

The All of Us (AoU) Research Program, funded by the National Institutes of Health (NIH), was initiated in May 201818. This program focuses on enhancing biomedical research by increasing the diversity of study participants and reducing health disparities. Information collected includes demographic characteristics, socioeconomic determinants of health, health outcomes, behaviors, as well as electronic health records (EHRs), and genetic information. The latest release, Version 8 (V8), has covered 633,547 individuals.

SDoH Instruments

As part of AoU’s data collection to promote precision health, surveys are administered to participants19. The Basics survey, developed at the start, gathers core demographic and socioeconomic data, including self-reported race, ethnicity, income, and education. A SDoH Task Force developed a scientifically valid survey to collect data on SDoH17. In V8, 259,189 participants completed some questions on the SDoH survey. AoU also uses a Health Care Access & Utilization (HCAU) survey, completed by 305,857 participants.

In addition, AoU provides region-specific SDoH measures through the American Community Survey (ACS). Area-level measures include economic stability, education, neighborhood environment, and health care access. AoU also uses the Nationwide Community Deprivation Index (NCDI), which helps explain over 60% of total variance in census-tract-level measurements from the ACS20.

Cohorts

Participants who completed the Basics survey and had sufficient EHR completeness are included in the study. Among these, 125,295 had linked area-level SDoH data along with individual-level education, income, and household size data (“SES Cohort”). A subset (54,313) completed the individual-level SDoH and HCAU surveys with at least a 60% response rate across all five Healthy People 2030 domains (“Individual SDoH Cohort”)

.

Race and ethnicity (SIRE) were categorized into three groups: non-Hispanic Black (NHB), non-Hispanic White (NHW), and Hispanic (HS). Using this categorization, two additional sub-cohorts were created: the “SES-SIRE Cohort” (117,535) and the “Individual SDoH-SIRE Cohort” (51,265).

SDoH Domain Development

Using the Basics, SDoH, and HCAU surveys, SDoH survey scores were calculated21. Cronbach’s alpha was calculated to ensure internal consistency within related survey items. A total of 24 unique individual-level SDoH constructs were obtained from these surveys, with a varying degree of missingness22.

Disease Definition Algorithms

Nine chronic conditions were analyzed due to their high prevalence, cost, and actionability. Disease definitions were adapted from EHR-based algorithms from the Electronic Medical Records and Genomics (eMERGE) network, covering asthma, atrial fibrillation (Afib), breast cancer, chronic kidney disease (CKD), coronary heart disease (CHD), hypercholesterolemia (HCL), prostate cancer, and type 1 and type 2 diabetes (T1D, T2D).

Covariates

Age at last event in the EHR record, Sex/Gender, record depth, and visit frequency were recorded. Participants identifying as cisgender males were excluded from breast cancer models, and cisgender females from prostate cancer models.

Imputation

Missing data was imputed using ten imputations (ridge = 0.001) using the mice package in R (v.3.17.0). All 24 SDoH variables, covariates (age, age2, Sex/Gender, visit frequency, record depth), and SIRE were used as predictive variables.

Evaluation of Performance of SIRE, SES, and SDoH Measures in Disease Prediction Models

To evaluate predictive value of SIRE, SES, and SDoH measures, standard logistic regression models were fitted for prevalence of each chronic condition. Models were evaluated using area under the curve (AUC) and a Bonferroni threshold of P < 3.70 ×10-4 and P < 9.26 ×10-4 was used to adjust for the number of comparisons.

PheWAS

PheWAS can be used to scan for associations with groups of International Classification of Diseases (ICD) codes (phecodes)28. A targeted PheWAS was conducted in each cohort. In the Individual SDoH Cohort, associations were tested across the five domains and the area-level metrics for a total of ten predictors. 513 unique diseases were tested, resulting in a Bonferroni-corrected threshold of 9.750 ×10-6 (0.05/5130). In the SES Cohort, associations were tested across the percent of poverty threshold, education, and the area-level metrics for a total of seven predictors. 621 unique diseases were tested, resulting in a Bonferroni-corrected threshold of 1.15 ×10-5 (0.05/4347).

Ethical Approval and Data Access

This study used data from the AoU Research Program Researcher Workbench. All participants provided informed consent at program enrollment. Access to Registered and Controlled Tier data was limited to authorized researchers who completed required ethics and data use training and agreed to AoU Data User Code of Conduct and related data access policies. All analyzes were conducted within the secure AoU Researcher Workbench environment in accordance with program policies designed to protect participant privacy and confidentiality.


Read More Health & Wellness News ; US News