A simplified metric to streamline between-group fairness assessment for predictive models: Algorithm Evaluation
ABSTRACT
Background:
Accurate and equitable risk prediction is essential for clinical decision support. Commonly used discrimination metrics, such as the concordance index (CI) and area under the receiver operating characteristic curve (AUC), are often evaluated within subgroups to assess fairness. However, these within-group metrics fail to capture whether models rank individuals equitably across groups, potentially obscuring systematic disparities.
Objective:
To develop and evaluate novel fairness-oriented discrimination metrics for clinical risk prediction that address limitations of within-group and pairwise across-group approaches.
Methods:
We examined theoretical properties of existing U-statistic–based metrics, including concordance index (CI) and area under the receiver operating characteristic curve (AUC), when applied to subgroups. We highlighted the distinction between within-group discrimination (ranking within a subgroup) and group-level discrimination (ranking relative to the broader population). Building on this framework, we proposed group-level extensions of CI and AUC that summarize subgroup-specific performance in a single interpretable measure. We then applied these metrics to the PREVENT equation, a recently developed model for atherosclerotic cardiovascular disease (ASCVD).
Results:
Traditional subgroup-specific CI and AUC captured within-group but not group-level discrimination, obscuring inequities in clinical decision-making. Existing cross-group approaches (e.g., xCI, xAUC) addressed this limitation but became computationally and interpretively burdensome with multiple subgroups due to pairwise comparisons. Our proposed metrics provided a streamlined alternative, yielding one summary statistic per subgroup while retaining sensitivity to cross-group ranking disparities. Applied to PREVENT, these metrics revealed differences in subgroup performance not apparent from within-group evaluations.
Conclusions:
By distinguishing between within-group and group-level discrimination, our framework clarifies a common source of misinterpretation in fairness evaluation. The proposed group-level extensions of CI and AUC provide practical, interpretable tools for evaluating fairness in clinical prediction models, enabling more transparent and equitable risk assessment.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.