Previously submitted to: JMIRx Med (no longer under consideration since Apr 22, 2024)
Date Submitted: Jun 13, 2023
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Random effects adjustment in machine learning models for cardiac surgery risk prediction: a benchmarking study
Background:
With the growing interest in big data and its leverage through the use of ML approaches that are not limited by linear statistical assumptions, the number of clinical variables can theoretically increase exponentially in large national datasets such as the National Adult Cardiac Surgery Audit (NACSA). One issue of using such a large dataset is that there may exist systematic differences in relationships of variables across different hospital sites. In addition, “hidden” confounders not included in the list of variables limit the interpretability of any ML scores that would be built, making such scores based mainly on associations rather than enabling causal interpretations in relation to the outcome.
Objective:
The objectives were to i) develop and compare methods for incorporating hospital sites as random effects in ML to improve the performance of cardiac surgery risk prediction; ii) to separately build Bayesian network models to capturing aetiological insights from the risk factors identified as improving prediction in step i), by considering these with and without the context of other potential confounders.
Methods:
A retrospective analysis of prospective routinely gathered data, on 227,087 adult patients undergoing cardiac surgery with mortality rate of 2.76%, in the UK between 2012-2019 was conducted. We temporally split the dataset into two cohorts: Training/Validation (n = 157196; 2012-2016) and Holdout (n = 69891; 2017-2019). ML mortality prediction models were built using Tree ensemble algorithms (Xgboost, RF HE, GPBoost) through different encodings of the random effects variable, with and without importance based variable selection. These were assessed using Clinical Effectiveness Metric (CEM) for overall performance and for performance within each of CEM’s constituent metrics. Confounding and potential causal relationships between covariates and outcomes were evaluated using Bayesian Network analysis.
Results:
For non-variable selected (NVS) risk scores with 102 variables, Xgboost with adjustment for hospital variation was superior to Xgboost without adjustment (p < 2e-16). Both NVS and the 18 variables selected (VS) Xgboost with adjustment for hospital variation risk scores were superior to the Xgboost (ES II 18 variables) model (p < 6.3e-15), with NVS Xgboost with adjustment for hospital variation having the best performance, followed by the VS Xgboost with adjustment for hospital variation (CEM Difference: 0.0150 and 0.0023, respectively).
Conclusions:
We have identified an ML-adjusted risk score comprising 102 variables that increases risk stratification performance on the hold-out dataset, removing the need to perform variable selection and reduction. This paves the way for further research that utilises this new set of variables with hospital-based adjustments for the safer selection of patients undergoing cardiac surgery.
Clinicaltrial:
International Registered Report:
RR2-https://doi.org/10.1101/2023.06.08.23291129
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.