Previously submitted to: JMIR Medical Informatics (no longer under consideration since Jan 07, 2026)
Date Submitted: Dec 15, 2025
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
FUSE-Sepsis: A Multi-Modal Fusion Framework for Early Sepsis Prediction Using LLM Fine-Tuning and Tabular Learning
ABSTRACT
Background:
Early prediction of sepsis in the ICU remains challenging because multivariate clinical time-series are noisy, highly imbalanced, and often incompletely documented. Traditional machine-learning models such as gradient boosting perform well on structured data but struggle with complex temporal and contextual reasoning, while large language models (LLMs) excel at reasoning but are not naturally designed for numerical time-series and are expensive to fully fine-tune in vertical clinical domains.
Objective:
This study aimed to develop and evaluate FUSE-Sepsis, a multimodal fusion framework that leverages the complementary strengths of tabular machine-learning models and parameter-efficiently fine-tuned LLMs for early sepsis prediction. We sought to determine whether LoRA-tuned, knowledge-enhanced LLMs, combined with traditional models, can improve discriminative performance and clinical utility compared with state-of-the-art biomedical language models and conventional baselines.
Methods:
We used the PhysioNet 2019 “Early Prediction of Sepsis” ICU dataset, containing over 1.55 million hourly records from 40,336 patients. Engineered temporal features were derived from vital signs, laboratory measurements, and demographics. CatBoost and logistic regression served as structured-data backbones, while LLaMA3-8B, Qwen-14B, and DeepSeek-6B were fine-tuned with Low-Rank Adaptation and enhanced via retrieval-augmented prompts that serialize clinical features into text. Model outputs were integrated through an optimized weighting scheme with post-processing modules for dynamic thresholding, hysteresis filtering, false-alarm control, and trend boosting. Performance was assessed using AUROC, AUPRC, standard classification metrics, utility scores, and average early-warning time.
Results:
LoRA-tuned LLMs substantially outperformed biomedical baselines. The best balanced configuration, LLaMA3-8B with LoRA, achieved sensitivity 0.926, specificity 0.997, AUROC 0.966, AUPRC 0.954, and a utility score of 0.519, providing clinically actionable early warnings on average 4.6 hours before sepsis onset. Qwen-14B attained AUROC 0.995 and AUPRC 0.997 but with lower sensitivity (0.826). All LLM-based models showed markedly higher utility than PubMedBERT and BioClinicalBERT, exceeding their practicality scores by more than 50%.
Conclusions:
FUSE-Sepsis bridges traditional tabular learning and modern LLM fine-tuning to deliver robust, accurate, and computationally efficient early sepsis prediction. Our findings show that general-purpose LLMs, when adapted with LoRA and informed by engineered clinical features, can surpass both conventional machine-learning approaches and domain-specific biomedical language models on structured ICU time-series. This framework offers a promising paradigm for deploying foundation models in critical-care decision support and other multivariate clinical time-series applications.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.