Theoretical Exploration of Error Thresholds for Clinical AI Decision Support in Nursing: An Exploratory Simulation Study Grounded in Human–AI Reliance Data
ABSTRACT
Background:
Clinical artificial intelligence (AI) decision support is being introduced into nursing practice; however, existing large language models demonstrate only moderate accuracy on complex clinical tasks, raising questions about the level of accuracy required for safe clinical use across varying levels of clinician experience and task complexity.
Objective:
To develop an empirically calibrated simulation model of human–AI reliance and error in nursing decision-making and estimate the AI accuracy required to achieve clinically acceptable error rates.
Methods:
A linear reliance model with coefficients for AI accuracy (A), clinician experience (E), and task complexity (C) was calibrated using weighted least squares against nine empirical data points from three independent randomized experiments on AI-assisted decision-making (N = 3,502; Lu and Yin, 2021; Yin et al., 2019). Predicted error was computed as reliance × (1 − A) across a 27-cell factorial design. Study-level bootstrap (2,000 iterations) quantified calibration uncertainty. AI accuracy was also evaluated on 100 clinical case-based questions.
Results:
Empirical LLM accuracy was 72.0% (95% CI 62.1%–80.5%), with significant degradation in high-complexity cases (χ²(1) = 4.44, p = 0.035). Calibration placed β_A at 0.201 (bootstrap 95% CI 0.023–0.234; P(β_A > 0) = 1.000). For the novice × high-complexity combination, the minimum AI accuracy values required to achieve error rates <10% and <20% were 0.89 and 0.78, respectively (bootstrap 95% CIs 0.88–0.90 and 0.75–0.79). The observed 72.0% LLM accuracy predicts an error rate of approximately 25% in this high-risk condition.
Conclusions:
Current LLM accuracy is insufficient for high-complexity nursing decision support by novice clinicians, requiring accuracy approaching or exceeding 90% for clinically acceptable error rates. The A_threshold framework provides a decision-theoretic tool for evaluating the minimum AI accuracy requirement by user-and-task profile. Behavioral validation in nursing contexts remains an essential next step.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.