Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Currently submitted to: JMIR AI

Date Submitted: Nov 10, 2025

Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.

Temperature-based calibration of pretrained language models for automated systematic review screening of biomedical literature

  • Gary C.K. Chan; 
  • Estrid He; 
  • Karin Verspoor

ABSTRACT

Background:

Systematic reviews are essential to evidence-based health research, but manually screening titles and abstracts is time-consuming and labour-intensive. Pretrained language models (PLMs) offer automation potential, yet these models may be overconfident in their decision due to suboptimal calibration, especially when facing distribution shifts. This limits their trustworthiness and practical application.

Objective:

To evaluate and compare three temperature scaling methods for improving the calibration of PLMs used in systematic review screening, including a novel review-specific approach based on the Mahalanobis distance.

Methods:

Using a dataset of 8,608 systematic reviews comprising over 540,000 abstracts, BERT and BioBERT models were fine-tuned for screening tasks. Calibration performance was assessed using expected calibration error (ECE) and negative log-likelihood (NLL) across three methods: 1. Standard temperature scaling, 2. Parameterized temperature scaling (PTS), and 3. A novel review-specific method adapting temperature based on Mahalanobis distance between a review’s embedding and the training distribution centroid.

Results:

All three calibration methods improved performance compared to uncalibrated models. Standard temperature scaling achieved the best calibration for in-distribution topics, while PTS tended to overfit validation data. The proposed Mahalanobis distance–based method outperformed others under distribution shift, showing lower calibration errors for out-of-distribution reviews.

Conclusions:

Review-specific calibration enhances model reliability for automated systematic review screening. By aligning model confidence with empirical probabilities, calibrated PLMs support better uncertainty handling and more principled stopping rules, reducing dependence on arbitrary heuristics in technology-assisted reviews.


 Citation

Please cite as:

Chan GC, He E, Verspoor K

Temperature-based calibration of pretrained language models for automated systematic review screening of biomedical literature

JMIR Preprints. 10/11/2025:87477

DOI: 10.2196/preprints.87477

URL: https://preprints.jmir.org/preprint/87477

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.