Accepted for/Published in: JMIR Formative Research
Date Submitted: Nov 24, 2025
Date Accepted: Jun 11, 2026
A Systematic Study of Cohort Selection Criteria and Their Impact on Machine Learning Model Performance and Demographic Disparities in COVID-19 Outcomes
ABSTRACT
Background:
Cohort selection criteria play a critical role in shaping machine learning (ML) model performance and the equity of clinical outcome predictions across demographic groups. In practice, cohort definitions are often influenced by arbitrary or inconsistent data processing decisions, which can introduce bias and limit the generalizability of ML models. During the COVID-19 pandemic, rapid cohort construction further increased concerns about transparency and fairness in ML-based analyses.
Objective:
This study aimed to systematically examine how cohort selection and data processing decisions influence ML performance and demographic equity in predicting COVID-19–related in-hospital mortality.
Methods:
Using data from the National COVID Cohort Collaborative (N3C), we evaluated two sets of cohorts. Set 1 consisted of 16 cohorts derived from four primary data processing decisions: COVID-19 case identification, inpatient inclusion, diagnosis date selection, and admission timestamp availability. Set 2 expanded this design to 64 cohorts by additionally applying provider ID and location ID filtering. Model performance was assessed using area under the receiver operating characteristic curve (AUC) across multiple training–testing cohort combinations. Three ML models—logistic regression, random forest, and gradient boosting—were evaluated using three analytical approaches: maximum AUC classification, direct AUC regression, and AUC gap analysis. Performance was further examined across demographic subgroups defined by gender, race, and ethnicity.
Results:
Cohort selection decisions had a substantial impact on ML model performance. Admission time inclusion or exclusion emerged as the most influential factor in Set 1 and consistently affected model accuracy across analytical approaches. In Set 2, this decision remained important, while additional criteria, particularly provider ID filtering, also significantly influenced results. The importance of specific decisions varied across models and evaluation strategies. Analyses across demographic subgroups showed that data processing decisions affected predictive performance differently by gender, race, and ethnicity.
Conclusions:
Seemingly minor cohort selection and data processing decisions can meaningfully affect both predictive accuracy and demographic equity in ML-based COVID-19 outcome prediction. These findings highlight the risk of bias introduced by arbitrary cohort definitions and underscore the need for transparent, standardized, and equity-aware cohort selection practices to support fair and reproducible ML research in healthcare. Clinical Trial: Not applicable. This study is a retrospective observational analysis of de-identified electronic health record data and does not constitute a clinical trial.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.