Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Currently submitted to: JMIR Bioinformatics and Biotechnology

Date Submitted: Aug 11, 2026
Open Peer Review Period: Sep 1, 2026 - Oct 27, 2026
(currently open for review)

Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.

Privacy-Preserving Automated Systematic Reviews via Edge-Hosted Large Language Models: Proof-of-Concept and Cross-Validation Study

  • Yosuke Sato

ABSTRACT

Background:

The exponential growth of biomedical literature has created a critical bottleneck in translational research. Conventional systematic reviews are constrained by manual curation, human screening bias, and time-intensive data extraction. While cloud-based artificial intelligence (AI) offers high-throughput processing, its application in medical informatics introduces severe data privacy risks and potential stochastic hallucinations

Objective:

To engineer, deploy, and validate a secure, fully automated natural language processing (NLP) pipeline utilizing an edge-hosted large language model (LLM) to accelerate systematic reviews, using preventive regenerative medicine interventions for the "Mibyo" (systemic resilience) state as a proof-of-concept.

Methods:

A 120-billion parameter open-weights LLM was deployed on a high-performance, strictly offline edge computing infrastructure (NVIDIA DGX SPARK). To ensure deterministic and reproducible data extraction, the generation temperature was set to 0.0. A zero-shot prompting framework, combined with a custom Python-based regular expression parser, was engineered to autonomously extract predefined therapeutic metrics (e.g., cell source, Senescence-Associated Secretory Phenotype suppression) from 3,228 PubMed abstracts into a structured JSON format. To validate the architectural necessity of massive parameter scale, an offline inter-model cross-validation was conducted against a lightweight local secondary model (0.5-billion parameters).

Results:

The AI-driven pipeline achieved a 100% parsing success rate with zero data loss. Inter-model cross-validation revealed that while the 0.5-billion parameter model achieved an apparent accuracy of 92%, it yielded a Cohen κ of 0.0 due to majority class collapse on highly imbalanced biomedical datasets. This mathematically proved the necessity of the 120-billion parameter model for nuanced text extraction. As a proof-of-concept, the 120-billion parameter pipeline autonomously identified 878 Mibyo-relevant studies and stratified an adipose-derived stem cell cohort (n=182). Downstream statistical synthesis of the AI-structured data successfully identified critical translational trends, mathematically proving via logistic meta-regression that cell-free exosomes yield a significantly higher probability of systemic resilience compared with whole-cell therapies (odds ratio 4.27; P < .001), while automatically mapping unexploited delivery routes (e.g., intranasal).

Conclusions:

The proposed edge-hosted LLM infrastructure provides a secure, highly accurate, and scalable paradigm for biomedical text mining. By automating data structuration and overcoming the cognitive limitations of lightweight local models, this computational approach drastically accelerates evidence synthesis and identifies hidden translational gaps, demonstrating significant utility for metascience and next-generation medical research.


 Citation

Please cite as:

Sato Y

Privacy-Preserving Automated Systematic Reviews via Edge-Hosted Large Language Models: Proof-of-Concept and Cross-Validation Study

JMIR Preprints. 11/08/2026:109300

DOI: 10.2196/preprints.109300

URL: https://preprints.jmir.org/preprint/109300

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.