Previously submitted to: JMIR Medical Informatics (no longer under consideration since Mar 11, 2022)
Date Submitted: Oct 4, 2021
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
A comparison of distributed machine learning methods for the support of "Many Labs" collaborations in computational modelling of decision making
Background:
Deep learning models are potentially useful for modeling the complex learning processes and decision-making strategies used by humans. Such neural network models make fewer assumptions about the underlying mechanisms compared to cognitive models, the latter being used widely in the field of computational psychiatry. These less-restrictive models provide greater experimental flexibility in terms of applicability and are increasingly investigated in this field. However, this comes at the cost of involving a larger set of parameters requiring significantly more data for effective learning. This presents practical challenges given that most cognitive experiments involve relatively small numbers of subjects. Laboratory collaborations are a natural way to increase overall dataset size. However, data sharing barriers between laboratories as necessitated by data protection regulations encourage the search for alternative methods to enable collaborative data science.
Objective:
Distributed learning, especially federated learning (FL), which supports the preservation of data privacy, is a promising method for addressing this issue. To verify the reliability and feasibility of applying FL to train neural networks models used in the characterization of decision making, we conducted experiments on a real-world, “many-labs” data pool including experiment datasets from ten independent studies.
Methods:
The performance of single models trained on single laboratory datasets was poor. This unsurprising finding supports the need for laboratory collaboration to train more reliable models. To that end we evaluated four collaborative approaches. The first approach represents conventional centralized learning (CL-based) and is the optimal approach but requires complete sharing of data which we wish to avoid. The results however establish a benchmark for the other three approaches, federated learning (FL-based), incremental learning (IL-based), and cyclic incremental learning (CIL-based).
Results:
We evaluate these approaches in terms of prediction accuracy and capacity to characterise human decision-making strategies. The FL-based model achieves performance most comparable to that of the CL-based model.
Conclusions:
This indicates that FL has value in scaling data science methods to data collected in computational modelling contexts when data sharing is not convenient, practical or permissible.
Clinicaltrial:
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.