Accepted for/Published in: Journal of Medical Internet Research
Date Submitted: Jan 5, 2026
Date Accepted: Aug 25, 2026
Automated Brief Hospital Course Summarization in Cardiac Surgery Using a Lightweight Large Language Model Based Framework: Development and Evaluation Study on MIMIC-IV
ABSTRACT
Background:
Physician documentation requirements are a known contributor to clinician burnout, with the manual creation of brief hospital course (BHC) summaries being particularly time-consuming. Automating BHC summarization may mitigate this workload and reduce documentation errors. However, current natural language processing (NLP) methods are often limited to single-document inputs, and large language models (LLMs) face privacy and deployment challenges. Furthermore, existing methods often require manual information extraction and struggle to maintain temporal accuracy.
Objective:
We developed and evaluated LiteMedDoc, a lightweight, locally deployable LLM-based framework to automatically generate BHC summaries without model fine-tuning. Our objective goal was to determine whether clinically useful clinical summaries could be generated under strict privacy and resource constraints, making automated summarization feasible in real-world hospital environments.
Methods:
LiteMedDoc is a modular pipeline built upon an 8-billion-parameter open-source LLM (Llama 3.1). It features three specialized modules: a Static Dynamic Information Hierarchy (SDIH) module to condense multi-source inputs and structure clinical events chronologically; a Similar Document Retrieval-Augmented Generation (SD-RAG) module that retrieves contextually relevant prior case summaries; and a Self-adaptive Feedback Optimization (SFO) module used offline for prompt optimization. All processing was performed locally without any model fine-tuning. We evaluated the framework on a retrospective cohort of 4,538 coronary artery bypass grafting (CABG) surgery cases from the MIMIC-IV database. A held-out test set of 403 cases was used to generate BHC summaries. The model-generated summaries were compared to reference BHCs using eight standard NLP metrics covering lexical overlap (e.g., BLEU-4, ROUGE), semantic similarity (BERTScore, METEOR), and clinical relevance (AlignScore, MEDCON). Additionally, 15 cardiac surgeons conducted a clinical evaluation of a sample of model-generated summaries, rating them on completeness, correctness, readability, conciseness, and global quality using a 5-point Likert scale.
Results:
Without any model training, LiteMedDoc achieved strong performance across individual automated metrics, nearly matching a fine-tuned model and exceeding a 70-billion-parameter model on all metrics. Additionally, in a within-database cross-domain evaluation on lobectomy cases, the framework maintained encouraging performance after prompt adaptation and outperformed both the base model and the CABG-fine-tuned model. Surgeons rated the AI-generated summaries above the prespecified acceptability threshold across all domains (mean scores ≥3.0), specifically praising their structure and conciseness.
Conclusions:
By integrating three specialized modules, the proposed framework offers a practical, locally deployable solution for clinician-in-the-loop BHC draft generation under privacy and resource constraints, and holds promise for improving documentation efficiency and enhancing information continuity during care transitions.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.