Accepted for/Published in: JMIR Medical Education
Date Submitted: Jan 2, 2026
Date Accepted: Jul 12, 2026
Date Submitted to PubMed: Jul 12, 2026
Prompt Framing and Evidence Requirements for AI-Generated Educational Responses in Dental Education: Factorial Experimental Study
ABSTRACT
Background:
Large language models (LLMs) are increasingly used in health professions education; however, the role of prompt design in shaping their outputs remains poorly understood in clinical training contexts. In dentistry, where accuracy, patient safety, and procedural reasoning are essential, the impact of instructional framing and evidence requirements on Artificial Intelligence (AI) -generated guidance is particularly relevant.
Objective:
This study examined how these prompt characteristics affect the factuality, tone, stance, and structural features of ChatGPT-5 responses related to cavity preparation.
Methods:
A 2 × 2 factorial experiment manipulated instructional framing (patient-centered vs. skill-centered) and evidence requirement (evidence-required vs. no evidence) across 40 prompts. Five trained raters independently coded each response on multiple dimensions, including factuality, confidence tone, stance orientation, hedging, citations, safety notices, and response length. Inter-rater reliability was assessed using intra-class correlations and Fleiss’ κ. Consensus scores were analyzed using ANOVAs and chi-square tests.
Results:
Evidence requirement significantly increased factuality (p = .015), citation presence (95% vs. 25%; p < .001), and response length (p < .001). It showed marginal effects on hedging and confidence tone. Instructional framing exerted a strong influence on stance orientation (p = .006) but did not significantly affect factuality or length. No refusals occurred, and safety notices were infrequent across conditions. Inter-rater reliability was high for citation presence but low/variable for several subjective measures, especially factuality and safety notices.
Conclusions:
Prompt design meaningfully shapes both the substance and presentation of LLM-generated educational content in dentistry. Evidence requirements enhance reference-inclusion and perceived rigor but introduce verbosity, while framing influences the model’s stance. Deliberate, pedagogically aligned prompt engineering is crucial to support safe and effective AI use in dental education.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.