Accepted for/Published in: JMIR Medical Informatics
Date Submitted: Nov 25, 2025
Date Accepted: Jul 15, 2026
Predicting Call Abandonment in a Healthcare Call Center Using Non-Personal Operational Data: Machine Learning Study
ABSTRACT
Background:
Call abandonment is a critical barrier to patient access in healthcare call centers, yet predictive modeling efforts are limited by strict privacy regulations that restrict use of personal or behavioral data. Whether abandonment can be accurately predicted using only anonymized operational metrics remains unclear.
Objective:
This study evaluated the feasibility, performance, and operational utility of machine learning models trained exclusively on non-personal, routinely collected call center metrics to predict call abandonment across distinct organizational phases.
Methods:
We analyzed 1,037,363 call records from a large academic healthcare system spanning four operational periods marked by workflow changes and skill consolidation. Features included temporal variables, skill identifiers, and rolling operational metrics (in-queue time, occupancy, handle time, after-call work time, and active agents). Random Forest and CatBoost models were trained on three phases defined as: (T1 (Jan–Apr 2023; original workflows), T2 (May–Aug 2023; post skill consolidation, cross-training, and new workflows), and T3a (Sep–Dec 2023; optimized processes) using five-fold cross-validation with three imbalance-handling strategies (none, class weighting, SMOTE). Temporal generalizability was assessed by evaluating all models on all four phases, with T3b serving as an unseen holdout set. Performance was evaluated using AUC, PR-AUC, Brier score, and calibration error. SHAP values quantified feature contributions.
Results:
Across all training phases and algorithms, adding operational metrics improved ROC-AUC by 0.10–0.14 versus models using only temporal and skill features. The best configuration was a CatBoost model trained on T2 with operational metrics and no imbalance correction (T3b AUC 0.80, PR-AUC 0.093, ECE 0.005, Brier 0.026). Models trained solely on pre-intervention data (T1) generalized poorly to post-intervention periods when restricted to temporal and skill features (AUC 0.32–0.38) but achieved AUC ≈0.79 on T3b when operational metrics were included. SHAP analysis consistently identified in-queue time as the dominant predictor, with after-call work time, handle time, occupancy, and number of logged-in agents comprising the remaining top features. Abandonment declined from 8.7% in T1 to 2.8% in T3b; model-based analyses of temporal features (day of week and hour of day) showed highest risk on Mondays and between 11:00 and 16:00. Skill-level analyses showed marked improvement in high-volume imaging teams with high abandonment.
Conclusions:
Call abandonment in healthcare call centers can be accurately predicted using non-personal operational data alone, enabling privacy-compliant deployment of machine learning. Queue and staffing metrics provide the strongest predictive signal, and models must be retrained following major workflow changes to preserve generalizability. These findings support the use of interpretable, operationally grounded models to enhance resource allocation and improve patient access to care.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.