JMIR Preprints #70481: How to design, create and evaluate an instruction-tuning dataset for large language model training in healthcare: a tutorial from a clinical perspective.

Current Preprint Settings

(as selected by the authors)

1. When the manuscript is submitted, allow peer review from:

(a) Anybody (open community peer review)
(b) Editor-selected reviewers (closed peer review)

2. When the manuscript is submitted, display the preprint PDF to:

(a) Anybody, anytime
(b) Logged-in users only
(c) Anybody, anytime (title and abstract only)
(d) No one

3. When the manuscript is accepted, display the accepted manuscript PDF to:

(a) Anybody, anytime
(b) Logged-in users only
(c) Anybody, anytime (title and abstract only)
(d) No one

How to design, create and evaluate an instruction-tuning dataset for large language model training in healthcare: a tutorial from a clinical perspective.

Wojciech Nazar;
Grzegorz Nazar;
Aleksandra Kamińska;
Ludmila Danilowicz-Szymanowicz

ABSTRACT

High-quality data is critical in healthcare, forming the cornerstone for accurate diagnoses, effective treatment plans, and reliable conclusions. Similarly, high-quality datasets underpin the development and performance of large language models (LLMs). Among these, instruction-tuning datasets (ITDs) used for instruction fine-tuning have been pivotal in enhancing LLM performance and generalization capabilities across diverse tasks. This tutorial provides a comprehensive guide to designing, creating, and evaluating ITDs for healthcare applications. Written from a clinical perspective, it aims to make the concepts accessible to a broad audience, especially medical practitioners. Key topics include identifying useful data sources, defining the characteristics of well-designed datasets, and crafting high-quality instruction-input-output examples. We explore practical approaches to dataset creation, examining the advantages and limitations of three primary methods: fully manual construction by expert annotators, fully synthetic generation using AI, and an innovative hybrid approach in which experts create the initial dataset and AI generates additional data. Moreover, we discuss strategies for metadata selection and human evaluation to ensure the quality and effectiveness of instruction-tuning datasets. By integrating these elements, this tutorial provides a structured framework for developing ITDs. It bridges technical and clinical domains, supporting the continued interdisciplinary advancement of AI in medicine. Additionally, we address the limitations of current practices and propose future directions, emphasizing the need for a global, unified framework for ITDs. We also argue that artificial general intelligence (AGI), if realized, will not replace empirical research in medicine. AGI will depend on human-curated datasets to process and apply medical knowledge. At the same time, ITDs will likely remain the most effective method of supplying this knowledge to AGI, positioning them as a critical tool in AI-driven healthcare.

Citation

Please cite as:

Nazar W, Nazar G, Kamińska A, Danilowicz-Szymanowicz L

How to Design, Create, and Evaluate an Instruction-Tuning Dataset for Large Language Model Training in Health Care: Tutorial From a Clinical Perspective

J Med Internet Res 2025;27:e70481

DOI: 10.2196/70481

PMID: 40100270

PMCID: 11962319

Download PDF

Request queued. Please wait while the file is being generated. It may take some time.

JMIR Publications

JMIR Preprints

Accepted for/Published in: Journal of Medical Internet Research

Date Submitted: Dec 23, 2024

Date Accepted: Feb 7, 2025

(closed for review but you can still tweet)

How to design, create and evaluate an instruction-tuning dataset for large language model training in healthcare: a tutorial from a clinical perspective.

ABSTRACT

Citation