Currently submitted to: JMIR Formative Research
Date Submitted: Sep 21, 2026
Open Peer Review Period: Sep 21, 2026 - Nov 16, 2026
(currently open for review)
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
From Prompt Quality to Output Quality: A Blinded Human Evaluation of AI-Generated Responses to Healthcare Research Tasks
ABSTRACT
Background:
Background:
Generative artificial intelligence is increasingly used in higher education and health research, but fluent outputs can conceal evidential and methodological weaknesses. In Phase 1 of this project, automated prompt enhancement substantially improved audit-rated prompt quality across 10 authentic healthcare education research tasks, especially failure handling. Whether downstream AI-generated research responses were themselves of sufficient quality remained unresolved.
Objective:
Objective:
This formative Phase 2 study aimed to characterize the quality of AI-generated responses to common healthcare research tasks using blinded independent human scoring, and to identify measurement and governance issues that should be addressed before condition-level claims are made about automated prompt enhancement.
Methods:
Methods:
Thirty coded responses represented 10 matched healthcare research tasks, with 3 responses per task corresponding to the wider study's raw, researcher-authored, and automatically enhanced prompt conditions. An independent research assistant, blinded to prompt condition, scored every response from 1 to 5 on accuracy, completeness, methodological rigor, academic quality, and practical utility (maximum total 25). Descriptive statistics, task-level profiles, and structured analysis of the assessor's free-text comments were undertaken. Because the response-code-to-condition key remained sealed for this analysis, no condition-level efficacy comparison was inferred.
Results:
Results:
The mean total score was 19.93/25 (SD 1.87; median 20; range 14-22). Completeness had the highest mean (4.17/5), followed by accuracy (4.07), methodological rigor (4.00), academic quality (3.97), and practical utility (3.73). Academic quality showed a pronounced ceiling effect, with 29 of 30 responses scoring 4/5. Mixed-methods research design had the highest task mean (21.67/25), whereas Discussion writing had the lowest (17.33/25). Eight responses that withheld substantive completion when essential evidence was absent had mean accuracy of 5.00/5 and methodological rigor of 4.50/5, but completeness of 2.88/5. Four responses that constructed hypothetical or illustrative content had higher completeness (4.00/5) but lower accuracy (3.50), rigor (2.75), and practical utility (2.75).
Conclusions:
Conclusions:
Blinded human assessment showed that polished academic presentation was not a reliable proxy for evidential integrity or practical usefulness. The results reveal a completeness paradox: when necessary evidence is absent, a response may appear more complete only by becoming less defensible. Output assurance should therefore reward appropriate noncompletion, uncertainty disclosure, evidence fidelity, and human accountability. This stage does not establish that automated prompt enhancement improves downstream outputs; condition unblinding, multiple independent raters, repeated generations, and verified generation metadata are required for definitive validation.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.