Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: JMIR Aging

Date Submitted: Mar 3, 2026
Date Accepted: Aug 13, 2026

The final, peer-reviewed published version of this preprint can be found here:

Development and Temporal External Validation of a Parsimonious, Interpretable Machine Learning Model for Predicting 6-Month Mortality in Long-Term Care Facilities: A Retrospective Cohort Study

Tsai YC, Wu YL, Yu YC, Chen YW, Chen YW, Yang YC, Lo YT

Development and Temporal External Validation of a Parsimonious, Interpretable Machine Learning Model for Predicting 6-Month Mortality in Long-Term Care Facilities: A Retrospective Cohort Study

JMIR Aging 2026;9:e94567

DOI: 10.2196/94567

PMID: 42748398

Development and Temporal External Validation of a Parsimonious, Interpretable Machine Learning Model for Predicting 6-Month Mortality in Long-Term Care Facilities: A Retrospective Cohort Study

  • Yun-Cheng Tsai; 
  • Yi-Lin Wu; 
  • Yung-Chen Yu; 
  • Yi-Wen Chen; 
  • Yi-Wen Chen; 
  • Yi-Ching Yang; 
  • Yu-Tai Lo

ABSTRACT

Background:

Early mortality after long-term care facility (LTCF) admission is common, yet prognostic tools are often derived from Western minimum data-set-based cohorts or require hospital electronic health record linkages that are unavailable at intake in many LTCFs. There is also limited evidence on explainable, admission-feasible machine learning -based prognostication in Asian LTCF settings.

Objective:

To develop and temporally externally validate an interpretable machine learning model for predicting 6-month all-cause mortality among older adults newly admitted to LTCFs in Taiwan using routinely collected admission data.

Methods:

We conducted a retrospective cohort study using the JUBO Long-Term Care Database, a nationwide private administrative registry covering 636 LTCFs (37.43% of the long-term care facilities in Taiwan). We included residents aged ≥65 years with first-time LTCF admission and prespecified nonoverlapping cohorts for temporal validation: development (January 1, 2020–December 31, 2023; n = 23,901) and external validation (January 1–December 31, 2024; n = 6216). The outcome measure was death within 180 days of admission. We compared a nonlinear ensemble model (HybridXGBRF) with seven other algorithms, including tree-based and linear benchmarks. Discrimination (AUROC), classification metrics (accuracy, precision, recall, and F1), and calibration (Brier score and calibration plots) were assessed. Model interpretability was examined using Shapley Additive Explanations (SHAP).

Results:

In the development cohort, 5272/23,901 (22.1%) residents died within six months; in the 2024 validation cohort, 1781/6216 (28.7%) died. In internal cross-validation, the HybridXGBRF model achieved the highest AUROC (0.875, 95% CI 0.862–0.889) and the lowest Brier score (0.109). In the temporal external validation, the HybridXGBRF model maintained strong discrimination (AUROC 0.878, 95% CI 0.866–0.889), with an accuracy of 0.851 and an F1 score of 0.572. Calibration plots indicated a close agreement between the predicted and observed risks across most probability ranges in both cohorts, with a mild divergence at higher predicted risks. The SHAP analysis identified frequent hospitalizations within six months, activities of daily living impairment, and weight loss as the most influential predictors. The model showed stable AUROC across sex and age strata (0.88–0.89), and maintained high discrimination in the subgroup of residents with improving activities of daily living scores (AUROC: 0.90), a population where mortality risk is often underestimated by conventional clinical assessments.

Conclusions:

An interpretable machine learning model using routinely collected admission data provide accurate mortality predictions without requiring hospital-based electronic health record data linkages, and generalized well under temporal validation in Taiwan. This practical tool can serve as a decision-support system to prompt early palliative care discussions and enhance person-centered care planning in LTCF settings. Clinical Trial: Not applicable


 Citation

Please cite as:

Tsai YC, Wu YL, Yu YC, Chen YW, Chen YW, Yang YC, Lo YT

Development and Temporal External Validation of a Parsimonious, Interpretable Machine Learning Model for Predicting 6-Month Mortality in Long-Term Care Facilities: A Retrospective Cohort Study

JMIR Aging 2026;9:e94567

DOI: 10.2196/94567

PMID: 42748398

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.