Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Previously submitted to: JMIR Medical Informatics (no longer under consideration since Sep 13, 2024)

Date Submitted: Feb 18, 2024

Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.

Considerations for health care institutions training large language models on electronic health records

  • Weipeng Zhou; 
  • Danielle Bitterman; 
  • Majid Afshar; 
  • Timothy A. Miller

ABSTRACT

Background:

Large language models (LLMs) are usually built (pretrained) on general domain data. To adapt LLMs better to health care needs, it might be beneficial for health care institutions to create LLMs pretrained on their own EHR data. Nevertheless, there are concerns related to the training budget, the type of pretraining (pretraining from scratch vs. pretraining on top of pretrained LLMs), and the EHR database size.

Objective:

The objective of this manuscript is to address these concerns and to help decision-makers at health care institutions better understand opportunities and challenges for LLM development.

Methods:

We use published work on empirical experience in pretraining LLMs, and establish a relationship between the training budget, LLM size, pretraining database size and training time.

Results:

We found that pretraining a modest-sized LLM (13 billion parameters) from scratch requires around 800 GB of EHR data and costs 127,000 USD. For a 65 billion parameter LLM, the EHR requirement is 4000 GB, costing 3 million USD. In comparison, continued pretraining using 500 GB EHR on a 65 billion parameter pretrained LLM would cost 400,000 USD.

Conclusions:

Building private LLMs on EHR data can be challenging due to budget and database size limitations. Most health care institutions will not have the required resources to train their own LLMs from scratch, and even continued pretraining will be challenging for many. This study provides a framework for health care institutions when deciding about the deployment of LLMs in their systems.


 Citation

Please cite as:

Zhou W, Bitterman D, Afshar M, Miller TA

Considerations for health care institutions training large language models on electronic health records

JMIR Preprints. 18/02/2024:57484

DOI: 10.2196/preprints.57484

URL: https://preprints.jmir.org/preprint/57484

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.