Previously submitted to: Journal of Medical Internet Research (no longer under consideration since Jun 21, 2024)
Date Submitted: Mar 13, 2024
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Artificial Intelligence Based Chatbots as Professional Medical Consultants in Oral Surgery: Are they reliable? Comparative Study
ABSTRACT
Background:
With advances in artificial intelligence (AI) technologies, AI-based chatbots have become promising tools for generating medical information. Given their widespread use and easy accessibility, it is important to comprehensively evaluate the quality, accuracy, and safety of the information they produce to ensure their effectiveness as reliable sources in healthcare.
Objective:
The purpose of the study is to examine the reliability of chatbots and their role as professional consultants.
Methods:
64 questions were generated, including systemic diseases and medications for which professional consultation is often requested and common conditions that may raise concerns during performing oral surgery. The questions were posed to ChatGPT 3,5 and Claude-instant at 2 sessions with 1 week interval. The answers were recorded and rated by 2 experienced oral surgeons by using 2 evaluation metrics: A modified DISCERN tool, a subset of the original, was used to evaluate the answers in terms of quality, and the Likert scale(LS) was used to evaluate the answers in terms of accuracy (6 point LS) and completeness (3 point LS). Statistical analyses, including intraclass correlation, Mann-Whitney U test, skewness and kurtosis coefficients calculations were conducted to assess and compare the performance of the chatbots.
Results:
In terms of intrarater agreement, ChatGPT demonstrated a high level of quality and accuracy in both sessions. Additionally, quality scores of ChatGPT was found statistically significantly higher than Claude instant.
Conclusions:
The evaluated chatbots exhibited a remarkable capacity to generate valuable medical content. Our particular view is that, although they are currently insufficient to serve as a single source of information, AI-based chatbots will gain greater acceptance among healthcare professionals in the near future and can provide a quick solution to the high demand for medical care.
Citation
The author of this paper has made a PDF available, but requires the user to login, or create an account.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.