Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Safety Risk Prioritization for Patient-Facing Agentic Voice AI in Clinical Care: A Two-Round Modified Delphi Study
ABSTRACT
Background:
Patient-facing agentic voice artificial intelligence (AI) is moving from pilot to production across US health systems, with autonomous agents managing chronic disease check-ins, post-acute follow-up, and quality outreach. Unlike diagnostic models or ambient documentation tools, these agents initiate patient contact, interpret verbal responses in real time, and act with limited synchronous clinician oversight. Existing AI governance frameworks were not designed for autonomous systems of this kind. We previously developed a qualitative risk framework as a multi-stakeholder voice AI safety taskforce.
Objective:
This study aimed to establish multidisciplinary expert consensus on the relative priority of safety risks for patient-facing agentic voice AI, stratified by clinical use case, to support risk-proportionate oversight by health systems, payers, and vendors.
Methods:
A 14-member expert panel spanning 9 constituencies, including clinical practice, health system leadership, clinical informatics, industry, payer operations, patient safety research, and patient advocacy, rated that framework through a two-round modified Delphi conducted across 2025 and 2026. The instrument comprised 21 risk categories across 4 themes (6 agent-level, 5 data-level, 5 patient-level, 5 clinician-level), each rated against 5 clinical use cases, yielding 105 risk-by-use-case cells. Panelists scored likelihood and impact on 1-to-5 scales anchored in ISO 31000 descriptors and case-specific vignettes. Dual consensus was defined a priori as at least 70% of raters within 1 point of the median on both dimensions, per the RAND/UCLA Appropriateness Method. Round 2 non-responder ratings were carried forward, with prespecified sensitivity analyses.
Results:
All 14 panelists completed Round 1 (2940 ratings); 11 of 14 completed Round 2. Dual consensus was reached on 95 of 105 cells (90.5%) after Round 1 and all 105 (100%) after Round 2. Of the 10 Round 1 non-consensus cells, 8 reflected impact disagreement alone, 1 likelihood alone, and 1 both. By theme, clinician-level risks ranked highest (median likelihood by impact 10.5, IQR 9.0-12.0), followed by agent-level (9.5, IQR 9.0–12.0), patient-level (9.0, IQR 8.0–12.0), and data-level (9.0, IQR 6.0–10.5). Hypertension monitoring and post-stroke follow-up ranked highest among use cases; HEDIS gap closure ranked lowest. Identity mis-verification accounted for 3 of the 10 Round 1 non-consensus cells. Theme rank order was preserved in sensitivity analyses excluding framework developers and excluding non-responders.
Conclusions:
A multidisciplinary panel reached complete consensus on a use-case-stratified prioritization of safety risks for patient-facing agentic voice AI. Clinician-level risks ranked highest, directing attention toward the clinician interface rather than agent-internal failure modes, and risk profiles varied substantially across use cases, indicating uniform governance is unlikely to be efficient. These priorities may inform risk-proportionate oversight, pre-deployment testing, and vendor contracting, although prospective validation against observed incidents is required before they can be used as predictive estimates of harm.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.