Previously submitted to: Journal of Medical Internet Research (no longer under consideration since Sep 12, 2022)
Date Submitted: Sep 5, 2022
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Personalized Progressive Federated Learning with Leveraging Client-Specific Vertical Features: Model Development and Validation
ABSTRACT
Background:
Federated learning (FL) is used to build models across distributed clients. Personalized federated learning (PFL) focuses on the training of personalized models for adaptation to diverse data distributions among clients. Horizontal federated learning (HFL) enables clients to train a global model based on distributed samples of the same feature space, whereas vertical federated learning (VFL) enables clients to train a global model based on the distributed features of the same sample. However, conventional HFL cannot leverage vertically partitioned features to increase the model complexity, and VFL requires all clients to share many overlapping sample-ids.
Objective:
In this study, we propose a personalized progressive federated learning (PPFL) model, a multi-model PFL approach that allows the leveraging of vertically partitioned client-specific features that can vary from client to client.
Methods:
PPFL personalizes the FL model by leveraging client-specific vertical features. PPFL’s performance was evaluated using two datasets: the Physionet Challenges 2012 dataset and a real-world dataset composed of eICU data and Highly Intensive Care Unit (HICU) data from Severance Hospital, Seoul, South Korea. Using these two datasets, we compared the performance and explainability of in-hospital mortality and length of stay task prediction between our model and the FedAvg and local models based on the model's accuracy and Area Under ROC curve (AUROC). We also compared the loss of the prediction task on in-hospital mortality during PPFL training with transfer learning to evaluate the effect of personalization mechanisms on PPFL.
Results:
The PPFL showed 0.867 accuracy and 0.807 AUROC, which are the highest scores compared to client-specific local models and FedAvg algorithms in the in-hospital mortality prediction task. In the length of stay prediction task, PPFL also showed an AUROC of 0.873. This is also the highest score among the compared models. Among common features, age and usage of mechanical ventilation features showed high SHAP values of 0.5 or more and 0.3 or more, respectively, and among vertical features, features related to vital sign (Glasgow Coma Scale, heart rate and oxygen) showed high SHapley Additive exPlanations values. Compared to transfer learning, PPFL showed higher learning performance by manifesting a stable decrease in the loss rate during training. Finally, PPFL achieved an AUROC of 0.955, the highest score compared to FedAvg and other local models in real-world data settings.
Conclusions:
We proposed the PPFL algorithm to personalize federated algorithms for heterogeneously distributed clients and expand the feature space for client-specific vertical feature information. PPFL allows the integration of common globally learned common information and client-specific vertical information to achieve personalized inference while building a robust model.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.