Previously submitted to: JMIR Human Factors (no longer under consideration since Dec 09, 2025)
Date Submitted: Apr 27, 2025
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
SafetyScan: Validating a Software Tool for Human-on-the-Loop Oversight of GenAI Behavioral Counseling Conversations
ABSTRACT
Background:
As generative AI (GenAI) is increasingly used in behavior interventions, ensuring the safety and appropriateness of AI-generated statements and detecting high-risk patient disclosures are critical.
Objective:
We developed and tested SafetyScan, a large language model (LLM)-assisted transcript scanner designed to support human-on-the-loop (HOTL) oversight by identifying and categorizing potentially concerning content in counseling encounters.
Methods:
To validate SafetyScan, we analyzed transcripts from a completed study where a GenAI conducted motivational interviewing with young adults reporting hazardous alcohol consumption. We analyzed 37 transcripts previously confirmed by human raters to contain no concerning content (serving as negative controls). We then created 37 positive test cases by inserting one randomly generated concerning statement into each of these original transcripts, drawn from five predefined categories: client safety (e.g., suicidal ideation), unsupported treatment recommendation, inappropriate counselor comment, discriminatory language, and invented client history. SafetyScan analyzed all 74 transcripts.
Results:
SafetyScan demonstrated strong performance with a sensitivity of 0.92, specificity of 1.00, and an F1 score of 0.96. It correctly identified 34 of the 37 inserted concerning statements (true positives) and correctly cleared all 37 benign transcripts (true negatives), resulting in 3 false negatives and 0 false positives. The three missed statements were two inappropriate counselor responses and one unsupported treatment recommendation, none judged as critical safety failures.
Conclusions:
These findings suggest SafetyScan offers a feasible and accurate solution for assisting human oversight of GenAI-driven counseling, supporting safer deployment of conversational AI systems in clinical and research contexts. Clinical Trial: NA
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.