Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Temperature-based calibration of pretrained language models for automated systematic review screening of biomedical literature
ABSTRACT
Background:
Systematic reviews are essential to evidence-based health research, but manually screening titles and abstracts is time-consuming and labour-intensive. Pretrained language models (PLMs) offer automation potential, yet these models may be overconfident in their decision due to suboptimal calibration, especially when facing distribution shifts. This limits their trustworthiness and practical application.
Objective:
To evaluate and compare three temperature scaling methods for improving the calibration of PLMs used in systematic review screening, including a novel review-specific approach based on the Mahalanobis distance.
Methods:
Using a dataset of 8,608 systematic reviews comprising over 540,000 abstracts, BERT and BioBERT models were fine-tuned for screening tasks. Calibration performance was assessed using expected calibration error (ECE) and negative log-likelihood (NLL) across three methods: 1. Standard temperature scaling, 2. Parameterized temperature scaling (PTS), and 3. A novel review-specific method adapting temperature based on Mahalanobis distance between a review’s embedding and the training distribution centroid.
Results:
All three calibration methods improved performance compared to uncalibrated models. Standard temperature scaling achieved the best calibration for in-distribution topics, while PTS tended to overfit validation data. The proposed Mahalanobis distance–based method outperformed others under distribution shift, showing lower calibration errors for out-of-distribution reviews.
Conclusions:
Review-specific calibration enhances model reliability for automated systematic review screening. By aligning model confidence with empirical probabilities, calibrated PLMs support better uncertainty handling and more principled stopping rules, reducing dependence on arbitrary heuristics in technology-assisted reviews.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.