Currently submitted to: JMIR Infodemiology
Date Submitted: Aug 27, 2026
Open Peer Review Period: Sep 13, 2026 - Nov 8, 2026
(currently open for review)
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Auditing Evolving Public Health Communication with Human-in-the-Loop AI: Development and Validation of a Longitudinal Natural Language Processing Framework
ABSTRACT
Background:
Public health agencies often have to issue guidance before scientific uncertainty has resolved. As evidence, risk, and policy change, earlier and later messages can remain visible together without the explanation that connects them, creating what we describe as an expectation gap between the stability people need in order to acts and the revision that responsible institutions sometimes need to make. This creates a longitudinal infoveillance challenge: identifying potentially difficult transitions within large institutional communication records while preserving the context needed for interpretation.
Objective:
We introduce the Public Health Message Consistency System (PHMCS), a human-in-the-loop framework for examining explanatory continuity: whether potentially difficult transitions in an institution’s longitudinal communication record can be located efficiently enough for contextual review.
Methods:
PHMCS combines corpus-quality filtering, supervised COVID-19 relevance classification, ordered within-agency temporal pairing, Sentence-Transformer retrieval, neural and symbolic contradiction signals, and rationale-bearing large-language-model screening. Development used a 1,000-pair collection for error analysis, protocol refinement, and adjudication; after protocol freeze, an independent held-out test set of 500 previously unseen real-message pairs was used for evaluation.
Results:
Preprocessing retained 103,801 messages, including 43,549 classified as COVID-19 relevant. Overlapping 3-month temporal windows produced 7,318,027 unique within-agency dyads after duplicate-pair removal. Masked and Permuted Pretraining (MPNet) retrieval retained 357,957 semantically comparable dyads comprising 24,343 unique messages, and the finalized protocol prioritized 6,212 dyads (1.74%) for review. On the independent held-out test set of 500 pairs, weighted F1 was 0.975; only seven reference cases were contradictions, limiting inference about rare-class performance.
Conclusions:
PHMCS does not treat changing guidance as institutional failure. It makes a longitudinal record reviewable by locating transitions where the path from an earlier message to a later one may warrant explanation, while leaving final interpretation to human reviewers. As an infoveillance application, PHMCS provides a scalable way to surface such transitions for subsequent contextual review rather than treating automated screening as final adjudication. Clinical Trial: null
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.