Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: Journal of Medical Internet Research

Date Submitted: Feb 6, 2026
Date Accepted: Jul 15, 2026

The final, peer-reviewed published version of this preprint can be found here:

Large Language Models for Heterogeneous Data Mining in Liver Disease: Framework Development and Retrospective Validation Study

Zhang H, Li X, Fang K, Ma Y, Li L, Yan H, Liu Y, Yu Y, Wang J

Large Language Models for Heterogeneous Data Mining in Liver Disease: Framework Development and Retrospective Validation Study

J Med Internet Res 2026;28:e92921

DOI: 10.2196/92921

PMID: 42696525

Large Language Models for Heterogeneous Data Mining in Liver Disease: Framework Development and Retrospective Validation Study

  • Haiping Zhang; 
  • Xinming Li; 
  • Kechi Fang; 
  • Yinxue Ma; 
  • Lijuan Li; 
  • Huiping Yan; 
  • Yanmin Liu; 
  • Yanhua Yu; 
  • Jing Wang

ABSTRACT

Background:

Background:

Differentiating liver disease subtypes is challenging due to overlapping clinical features. While traditional mining methods rely on manual feature engineering, Large Language Models (LLMs) offer a new paradigm for integrating complex, multi-source clinical data.

Objective:

Objective:

To investigate the potential of LLMs in clinical data mining and evaluate their ability to support diagnostic reasoning across different granularities of liver disease classification.

Methods:

Methods:

This retrospective study analyzed data from 7,543 patients with liver disease (2010-2025). We utilized three LLMs (Qwen3, Huatuo-o1, II-Medical) to extract semantic embeddings from unstructured clinical narratives. To assess the model's depth of representation, we constructed a hierarchical classification framework: a three-class task (Autoimmune Liver Disease [AILD] vs. Drug-induced Liver Injury [DILI] vs. Chronic Hepatitis B [CHB]) followed by a more granular four-class task, where AILD includes the constituent subtypes autoimmune hepatitis (AIH) and primary biliary cholangitis (PBC). These LLM-derived features were fused with raw laboratory data to develop and compare machine learning classifiers.

Results:

Results:

The LLM-encoded framework demonstrated robust potential in both tasks, simplifying the modeling workflow by reducing manual variable selection. In the three-class task, the fusion models effectively captured the broad features of AILD. In the refined four-class task, the models maintained high discriminative power, proving that LLM embeddings can capture the subtle semantic nuances required to differentiate closely related subtypes like AIH and PBC. While direct zero-shot reasoning by LLMs showed exploratory value for diagnosis and medication, the embedding-based hybrid models provided more stable and accurate outcomes for complex multiclass identification.

Conclusions:

Conclusions:

LLMs possess significant potential for clinical data mining, offering a scalable strategy to integrate unstructured text with structured data. By successfully addressing both broad and fine-grained liver disease classification tasks, this framework provides a feasible path for developing intelligent decision support tools in real-world clinical settings.


 Citation

Please cite as:

Zhang H, Li X, Fang K, Ma Y, Li L, Yan H, Liu Y, Yu Y, Wang J

Large Language Models for Heterogeneous Data Mining in Liver Disease: Framework Development and Retrospective Validation Study

J Med Internet Res 2026;28:e92921

DOI: 10.2196/92921

PMID: 42696525

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.