Currently submitted to: JMIR Human Factors
Date Submitted: Jul 20, 2026
Open Peer Review Period: Jul 21, 2026 - Sep 15, 2026
(currently open for review)
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Evaluating Human factors Design in Clinician-Facing AI Applications: A Scoping Review
ABSTRACT
Background:
Artificial intelligence (AI)-enabled tools have become increasingly designed for and implemented in clinical settings. However, empirical research on the human factors components of the human-AI dyad using these tools remains inconsistently evaluated and reported.
Objective:
To characterize the state of the literature evaluating human factors design elements in relation to clinician perceptions and use of AI tools.
Methods:
A scoping review was conducted using Arksey and O’Malley’s framework and reported following PRISMA-ScR guidelines. PubMed, Scopus, PsycINFO, and IEEE Xplore were searched for English-language studies (2020-2025). Eligible studies evaluated clinician-facing AI tools with outcomes evaluating either human factors design elements and/or clinician perception. Screening and data extraction were completed by a single reviewer. Extracted variables included study characteristics, AI type, tool purpose, participant type, sample size, explanation or transparency methods, and primary outcomes.
Results:
Of the 1,234 studies identified, 1,011 unique publications were screened and eighty-five studies met inclusion criteria. Most were published in 2023 or later and conducted in North America or Europe. Machine-learning (ML) dominated (80%, n=68), with fewer evaluating large language models (LLMs; 9.4%, n=8) or natural language processors (NLP; 3.5%, n=3). Tools supported diagnosis (n=33), prognosis/risk prediction (n=24), or treatment decision support (n=16). Sample sizes were typically small (median = 17). Usability and user experience were the most common primary outcomes (n=33), while trust, adoption, workflow fit, and interpretability were less common. Feature importance (n=26) and saliency maps (n=16) were applied inconsistently and primarily to ML models. No study systematically evaluated explainability for LLM tools.
Conclusions:
Evaluations of clinician-facing AI are heterogeneous, with inconsistent terminology and measurement strategies. Standardized terminology, consistent measurement strategies, and rigorous human-centered evaluations, particularly for LLM-based applications, remain needed to support safe clinical deployment. Clinical Trial: Not registered
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.