Accepted for/Published in: Journal of Medical Internet Research
Date Submitted: Feb 6, 2026
Date Accepted: Jul 15, 2026
Large Language Models for Heterogeneous Data Mining in Liver Disease: Framework Development and Retrospective Validation Study
ABSTRACT
Background:
Background:
Differentiating liver disease subtypes is challenging due to overlapping clinical features. While traditional mining methods rely on manual feature engineering, Large Language Models (LLMs) offer a new paradigm for integrating complex, multi-source clinical data.
Objective:
Objective:
To investigate the potential of LLMs in clinical data mining and evaluate their ability to support diagnostic reasoning across different granularities of liver disease classification.
Methods:
Methods:
This retrospective study analyzed data from 7,543 patients with liver disease (2010-2025). We utilized three LLMs (Qwen3, Huatuo-o1, II-Medical) to extract semantic embeddings from unstructured clinical narratives. To assess the model's depth of representation, we constructed a hierarchical classification framework: a three-class task (Autoimmune Liver Disease [AILD] vs. Drug-induced Liver Injury [DILI] vs. Chronic Hepatitis B [CHB]) followed by a more granular four-class task, where AILD includes the constituent subtypes autoimmune hepatitis (AIH) and primary biliary cholangitis (PBC). These LLM-derived features were fused with raw laboratory data to develop and compare machine learning classifiers.
Results:
Results:
The LLM-encoded framework demonstrated robust potential in both tasks, simplifying the modeling workflow by reducing manual variable selection. In the three-class task, the fusion models effectively captured the broad features of AILD. In the refined four-class task, the models maintained high discriminative power, proving that LLM embeddings can capture the subtle semantic nuances required to differentiate closely related subtypes like AIH and PBC. While direct zero-shot reasoning by LLMs showed exploratory value for diagnosis and medication, the embedding-based hybrid models provided more stable and accurate outcomes for complex multiclass identification.
Conclusions:
Conclusions:
LLMs possess significant potential for clinical data mining, offering a scalable strategy to integrate unstructured text with structured data. By successfully addressing both broad and fine-grained liver disease classification tasks, this framework provides a feasible path for developing intelligent decision support tools in real-world clinical settings.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.