Accepted for/Published in: JMIR Medical Informatics
Date Submitted: Apr 7, 2026
Open Peer Review Period: Apr 24, 2026 - Jun 19, 2026
Date Accepted: Aug 25, 2026
(closed for review but you can still tweet)
Locally Deployed Large Language Models for AI-Assisted Prescription Review: A Crossover Study of Human-AI Collaboration in Hospital Pharmacy
ABSTRACT
Background:
Pharmacists’ prescription review is key to medication safety, but rising outpatient prescriptions and expanding formularies leave less time per case, increasing error risk. Large language models (LLMs) show promise, yet two barriers hinder routine use: most systems rely on cloud-based commercial models, risking data breaches by transmitting protected patient information externally; and retrieval-augmented generation (RAG) pipelines used to reduce hallucinations depend on complex text vectorization and vector databases, which hospitals with limited IT resources struggle to build and maintain.
Objective:
To evaluate the feasibility and usefulness of a locally deployed, knowledge-augmented LLM as a decision-support tool within pharmacist-led outpatient prescription review.
Methods:
The open-source Qwen3-14B model was deployed on a hospital intranet server using the Ollama framework. A structured knowledge base derived from drug package inserts served as the principal reference and supported lightweight knowledge augmentation through exact-match injection. A two-period crossover design was used: two pharmacists independently reviewed 213 outpatient prescriptions under both unaided and AI-assisted conditions, yielding paired unaided and collaborative results for each prescription. Plain artificial intelligence (AI) review and knowledge-augmented AI review were also run on the same 213-prescription test set to quantify hallucination rates, and standalone AI performance was assessed against the reference standard. The reference standard was established by independent consensus of two supervising pharmacists, with disagreements adjudicated by a deputy chief pharmacist. Accuracy, sensitivity and specificity were compared between conditions using paired McNemar tests, and review time using the Wilcoxon signed-rank test.
Results:
Overall review accuracy was 97.2% (95% CI 94.0%-98.7%; 207/213) in the human-AI collaborative condition versus 82.6% (95% CI 77.0%-87.1%; 176/213) in the pharmacist-alone condition (P<.001). Sensitivity was 98% (61/62) versus 55% (34/62), and the false-negative rate fell from 45% to 2% (P<.001). Specificity did not differ significantly (96.7% vs 94.0%; P=.29). Knowledge augmentation reduced the model hallucination rate from 19.7% (42/213 prescriptions) to 4.7% (10/213), an absolute reduction of 15.0 percentage points (relative reduction 76.2%; P<.001). Collaborative review reduced the mean per-prescription review time from 2.33±0.97 to 1.12±0.49 minutes (approximately 51.9% reduction; Wilcoxon signed-rank Z = -12.65, P < .001, r = 0.87).
Conclusions:
A locally deployed, knowledge-augmented LLM used as a pharmacist-supervised pre-screening tool was associated with substantially higher accuracy and sensitivity in outpatient retrospective prescription review, while keeping all prescription data within the hospital network. Locally deployed open-source models may offer hospitals a privacy-preserving and practical route to AI-assisted pharmacy decision support.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.