Currently submitted to: Journal of Medical Internet Research
Date Submitted: Sep 2, 2026
Open Peer Review Period: Sep 3, 2026 - Oct 29, 2026
(currently open for review)
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Auditable Longitudinal Diagnostic Alignment for Psychiatric Follow-Up Using a Large Language Model and Rule-Based Strategy Engine: Model Development and Evaluation Study
ABSTRACT
Background:
Psychiatric follow-up requires clinicians to reconcile established diagnoses with evolving symptoms, standardized scale scores, treatment response, mental status findings, and longitudinal records. Large language models (LLMs) can interpret unstructured clinical narratives but may inconsistently extract exact evidence or apply temporal and clinical constraints. Combining semantic interpretation with deterministic evidence processing may provide more auditable support for this setting.
Objective:
This study aimed to develop and evaluate an auditable hybrid framework for aligning a 5-category psychiatric diagnostic profile with documented current-visit labels during follow-up and to quantify the contribution of a rule-based strategy engine.
Methods:
This retrospective single-center study used 40,724 deidentified records from 658 psychiatric outpatients (November 2017-January 2026): 7,974 clinical notes, 15,500 scale records, 14,746 diagnosis records, and 2,504 voice transcripts, examination reports, and patient-information records. Among 656 patients with clinical notes, 634 had confirmed current-visit labels; 508 were assigned to development and 126 to holdout testing. The target was a multi-label profile covering sleep, anxiety, depression, bipolar disorder, and other psychiatric conditions. The framework combined semantic phenotyping with DeepSeek V4 Flash, deterministic extraction of scale scores, medications, and prior diagnoses, and a strategy engine with 19 soft strategies and 1 hard rule. We compared the hybrid framework with LLM-only output and a parser-feature XGBoost baseline. Outcomes were label-level agreement, category F1 scores, and per-patient exact-match agreement. We estimated 95% CIs using 10,000 bootstrap iterations and used the McNemar test for paired hybrid and LLM-only predictions.
Results:
On the holdout set, label-level agreement was 94.3% (95% CI 92.4%-96.0%) for the hybrid framework and 85.9% (95% CI 83.2%-88.6%) for LLM-only output. Per-patient exact-match agreement was 82.5% (95% CI 76.2%-88.9%) and 54.8% (95% CI 46.0%-63.5%), respectively (χ²=36.05; P<.001). The parser-feature XGBoost baseline achieved 88.9% label-level agreement. Hybrid category F1 scores ranged from 0.53 for other conditions to 0.97 for sleep. Removing the prior-diagnosis continuity strategy reduced agreement from 94.3% to 86.2%; removing any other soft strategy changed agreement by 0.0-0.3 percentage points. Prior and current category flags were identical in 84.9% of comparable cases, and performance was not estimated separately for the 15.1% with label additions, removals, or substitutions.
Conclusions:
The hybrid framework improved agreement with documented follow-up labels while preserving an auditable evidence-reconciliation process. However, most of the incremental gain reflected prior-diagnosis continuity. The findings support interpreting the framework as a follow-up alignment and auditing tool rather than as evidence of de novo diagnostic accuracy or diagnostic-change detection. External and prospective evaluation, particularly in first-visit patients and patients with changing diagnoses, is required.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.