Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Currently submitted to: JMIR Formative Research

Date Submitted: Sep 14, 2026
Open Peer Review Period: Sep 15, 2026 - Nov 10, 2026
(currently open for review)

Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.

Automated prompt enhancement for AI-assisted research in higher education: A controlled pilot study in healthcare education

  • Nicole Pollock

ABSTRACT

Background:

1. Introduction 1.1 Generative artificial intelligence in higher education research Generative artificial intelligence (GenAI) has moved rapidly from experimental use into the everyday practices of higher education. Its applications now extend beyond teaching and assessment to academic writing, literature searching, research design, data interpretation, policy analysis and administrative work. Reviews of artificial intelligence in higher education have documented expanding use across disciplines and institutional functions, while also identifying persistent gaps in educator-focused and context-sensitive evidence (Crompton & Burke, 2023; Zawacki-Richter et al., 2019). The speed of adoption has created a practical tension: GenAI is sufficiently accessible to shape academic work, but its outputs remain probabilistic, fallible and sensitive to the way tasks are framed. This tension is particularly important in research. A fluent response may contain inaccurate claims, fabricated references, unsupported interpretations or inappropriate disclosure of sensitive material (Alkaissi & McFarlane, 2023; Farquhar et al., 2024; Walters & Wilder, 2023). Dai and Chan (2026) argue that higher education guidance has concentrated more heavily on teaching, learning and assessment than on the complex ethical and methodological decisions involved in research. Their work with postgraduate researchers shows that actionable, task-sensitive support is needed across literature review, data analysis, interpretation and academic writing. Prompt quality is therefore not merely a technical matter. It is an upstream condition of how researchers communicate purpose, evidence expectations, methodological boundaries, data-governance requirements and uncertainty to an AI system. Healthcare education provides a useful case domain because its academic tasks combine general higher education research practices with high expectations for evidence, professional accountability and data protection. The present study does not evaluate clinical decision-making or patient-facing use. Instead, it examines authentic academic knowledge-work tasks undertaken by a healthcare education researcher within a university context. The disciplinary setting provides ecological specificity, while the task categories; literature review, study design, qualitative analysis, critical appraisal, academic writing and policy analysis—are recognisable across many higher education fields. 1.2 Prompt engineering as artificial intelligence literacy Prompt engineering refers to the deliberate construction of instructions and contextual information intended to elicit a useful response from a generative model. A prompt may define the task, audience, role, scope, evidence base, output format, constraints, evaluative criteria and actions to take when required information is absent. Lee and Palmer’s (2025) systematic review concluded that prompt engineering is increasingly relevant to higher education curricula and that frameworks can help align GenAI interactions with educational goals. Walter (2024) similarly places prompt engineering alongside artificial intelligence (AI) literacy and critical thinking as a capability needed for meaningful use of AI in higher education. AI literacy provides a stronger educational frame than prompt-writing technique alone. Ng et al. (2021) conceptualise AI literacy through four connected dimensions: knowing and understanding AI; using and applying AI; evaluating and creating with AI; and engaging with ethical issues. From this perspective, an effective research prompt is not simply longer or more detailed. It demonstrates an understanding of the system’s limitations, applies the tool to a bounded task, creates conditions for critical evaluation and addresses ethical responsibilities. Prompt enhancement may therefore function as a scaffold for AI literacy when it makes these dimensions visible and actionable. Conversely, it may undermine literacy if users treat a polished prompt as a substitute for judgement or assume that automated improvement guarantees a reliable answer. This distinction matters for institutional practice. Chan’s (2023) ecological framework locates effective AI integration across pedagogical, governance and operational dimensions. Research-specific prompt guidance similarly sits at the intersection of professional development, data governance, scholarly integrity and infrastructure. Automated prompt-auditing systems may contribute to this landscape by identifying omissions and proposing structured revisions, but their educational and governance value requires empirical evaluation rather than vendor claims. 1.3 Theoretical framework: Prompts as sociotechnical and cognitive artefacts The study is interpreted through an integrated AI-literacy, sociotechnical and distributed-cognition framework. Sociotechnical systems theory holds that outcomes arise from the interaction of people, practices, organisational arrangements and technologies rather than from a technical artefact in isolation (Trist & Bamforth, 1951). Applied to GenAI in higher education, prompt quality is one element within a wider system that includes the researcher’s disciplinary expertise, institutional rules, the model and interface, available evidence, data environment, human review and the consequences of using the output. This perspective guards against technological determinism: changing the prompt may improve task specification, but cannot independently assure the quality of the research process. Distributed cognition conceptualises representations and artefacts as carrying part of the cognitive work of a system (Hutchins, 1995). A structured prompt externalises aspects of the researcher’s reasoning—such as methodological choices, inclusion criteria, evidence boundaries and required checks—into the interaction. Domain experts may encode substantive and disciplinary knowledge, while an automated enhancer may add procedural safeguards that users overlook. Cognitive load theory offers a complementary explanation: externalising multiple task constraints may reduce avoidable working-memory demands during complex academic work (Sweller, 1988). Human–automation research introduces a necessary caution. Perceived fluency and consistency can encourage automation misuse and over-reliance (Parasuraman & Riley, 1997). Reason’s (2000) systems approach suggests that safe practice depends on layered defences that make foreseeable failure more visible. In research prompting, such defences include explicit uncertainty statements, source-verification requirements, data minimisation, stop conditions and human decision authority. The theoretical proposition examined in this pilot is therefore not that automation replaces expertise, but that domain expertise and automated scaffolding may contribute different, complementary forms of prompt quality

Objective:

1.4 Research gap, aim and research questions Existing higher education literature establishes the importance of prompt engineering and AI literacy, but comparatively little empirical work evaluates tools that audit or enhance prompts before they are submitted to a GenAI system. Studies in academic and professional contexts show that prompting strategies can materially affect response structure and task performance, while novice users often struggle to translate intent into effective instructions (Giray, 2023; Meskó, 2023; Zamfirescu-Pereira et al., 2023). Evidence is particularly limited for research workflows, where the consequences of fabricated sources, unsupported analysis and disclosure of research data extend beyond immediate task performance. In addition, large language model (LLM)-based evaluators can exhibit self-preference, creating a risk of circularity when related systems both improve and score an artefact (Panickssery et al., 2024). The wider project was designed to evaluate both prompt quality and the quality and reproducibility of downstream AI-generated research outputs. This paper reports the first, bounded phase: a controlled pilot evaluation of prompt quality. The study asked: (1) Does automated enhancement improve the audit-rated quality of researcher-authored prompts for higher education research tasks? (2) Which dimensions of prompt quality and risk are most affected? (3) What limitations or new risks emerge from enhancement? and (4) How can the findings inform a provisional prompt-quality framework for higher education research?

Methods:

2. Methods 2.1 Design and setting A controlled, within-task methodological evaluation compared three prompt conditions across ten authentic academic research tasks. The primary comparison was between an original researcher-authored prompt (Condition A) and the same prompt after automated enhancement (Condition B). A third set of short one-line prompts was included as an exploratory raw baseline. The unit of analysis for the primary comparison was the matched prompt pair (ten pairs; twenty prompts). Across conditions, the task topic was held constant and only the formulation of the instruction varied. The study was conducted on 9–10 July 2026. It represents Phase 1 of a broader programme that originally proposed a larger prompt sample, generation and expert assessment of downstream outputs, repeated runs and consistency testing. The pilot was undertaken to establish feasibility, characterise prompt-level change and refine a quality-assurance framework before the resource-intensive output phase (Fig1). Accordingly, the analysis is descriptive and the framework is provisional. Raw baseline (n = 10) Exploratory comparison Original prompts Condition A (n = 10) Automated enhancement → Enhanced prompts Condition B (n = 10) All prompts audited using the standardised multi-agent rubric Primary matched comparison: Condition A versus Condition B Fig. 1 Phase 1 study design and prompt conditions 2.2 Task sample and prompt conditions The purposive task sample was drawn from the routine work of a university academic in nursing and healthcare education. It included two research-design tasks, two critical-appraisal tasks, two academic-writing tasks, two policy-analysis tasks, one literature-review task and one qualitative-data-analysis task. The sample was designed to span generative, analytic and evaluative academic work rather than to represent all disciplines or research methods. Original prompts were written by an experienced researcher and included domain-specific roles, named methodological approaches, evaluative criteria and intended outputs. Enhanced prompts were produced by passing each original prompt through the evaluated system and extracting the rewritten version verbatim; no manual editing was applied after enhancement. Raw prompts were single-sentence requests resembling novice or time-pressured use. Table 1 summarises the conditions. Table 1 Prompt conditions used in the controlled pilot Condition Purpose Construction n Raw baseline Exploratory context Single-sentence request with little role, scope, structure or failure guidance. 10 Original (A) Primary comparator Researcher-authored prompt containing disciplinary knowledge, methods and evaluative criteria. 10 Enhanced (B) Primary intervention Condition A prompt automatically audited and rewritten; output used verbatim. 10 Note. Condition A versus Condition B was the prespecified primary comparison; the raw baseline was exploratory. 2.3 Automated audit and enhancement system The evaluated system was a commercial prompt-auditing and enhancement platform [name withheld for double-blind peer review]. Audits used Claude Sonnet 4.6 at temperature 0. The pipeline made six model calls per prompt. A primary critic scored each prompt from 0 to 24 across eight dimensions, including clarity, output specification, contextual sufficiency, role, constraints, reasoning, examples and failure handling. A failure-mode agent then attempted to identify ways in which the instruction could yield misleading, unsafe or unusable outputs. Three specialist agents assessed quality, security and business or operational impact, each on a 0–12 scale. Their assessments were synthesised into a combined 0–36 score. A critical-risk override could force a Poor rating when a zero-scored safety-critical dimension coincided with maximum potential impact, irrespective of the arithmetic total. For original prompts, the system also generated an enhanced rewrite. The evaluation suite returned eight of eight checks as passing immediately before data collection. The primary ratings were defined by the system as Poor (<50%), Needs Improvement (50–74%) and Good (≥75%). The study treated these thresholds as properties of the evaluated rubric rather than independently validated educational standards. The complete rubric, per-prompt reports and software configuration should accompany any replication. 2.4 Outcomes and analysis The primary outcome was the matched difference in critic score (0–24) between each original prompt and its enhanced version, reported as raw points, percentage of the maximum and rating band. Secondary outcomes were the combined multi-agent score (0–36), task-level score changes, failure-handling scores, critical-risk overrides and qualitative risk findings. Raw-prompt scores were used only to contextualise the contribution of researcher expertise before automated enhancement. Means, absolute differences and percentage-point changes were calculated. Because the pilot involved only ten matched tasks, inferential testing and effect-size estimation were not undertaken. No AI-generated research outputs were produced; therefore accuracy, depth, criticality, usefulness, consistency and reproducibility were not evaluated. The theoretical framework was applied after scoring as an abductive lens for interpreting patterns rather than as part of the scoring algorithm. 2.5 Ethics, data governance and reflexivity The study used researcher-authored prompts, synthetic task descriptions and automated audit outputs. It involved no human participants, patient records, student data, identifiable interview transcripts or clinical decision-making. Ethics approval and consent were therefore not applicable to this methodological evaluation. Privacy and data governance were nevertheless treated as outcomes because research prompts may invite users to enter confidential or identifiable data into externally hosted systems. The author team combined higher education domain expertise and technical expertise in the evaluated platform. One author created the commercial platform evaluated in this study. Because the platform is commercially available and may be purchased by readers, this constitutes a competing interest that could influence design, interpretation or reporting. To reduce promotional framing, the enhanced prompts were used without manual polishing; adverse findings were retained; claims were limited to audit-rated prompt quality; and the manuscript explicitly examines circularity and self-preference. Independent human and cross-model validation remain necessary.

Results:

3. Results 3.1 Aggregate scores In the primary original-to-enhanced comparison, the mean critic score increased from 12.7/24 (52.9%) to 19.8/24 (82.5%), an absolute gain of 7.1 points and 29.6 percentage points. Five original prompts were rated Poor and five Needs Improvement; none reached Good. All ten enhanced prompts reached Good. The mean combined multi-agent score increased from 24.2/36 (67.2%) to 28.0/36 (77.8%). The exploratory raw baseline averaged 5.9/24 (24.6%) on the critic score and 20.1/36 (55.8%) on the combined multi-agent score. The gradient from raw to original to enhanced suggests that disciplinary expertise and automated enhancement contributed successive gains, although the study was not designed to identify causal mechanisms. Table 2. Aggregate audit scores by prompt condition. Condition Critic mean /24 Critic % Rating distribution Multi-agent mean /36 Multi-agent % Raw baseline 5.9 24.6 Poor 10; Needs Improvement 0; Good 0 20.1 55.8 Original (A) 12.7 52.9 Poor 5; Needs Improvement 5; Good 0 24.2 67.2 Enhanced (B) 19.8 82.5 Poor 0; Needs Improvement 0; Good 10 28.0 77.8 Percentages are expressed relative to the maximum score for each audit scale. Figure 2. Mean audit scores by prompt condition. The raw baseline was exploratory. The original-to-enhanced comparison was primary. 3.2 Matched task-level changes Every task improved on the combined multi-agent score after enhancement. Improvements ranged from one point for the second critical-appraisal and academic-writing tasks to nine points for the second policy-analysis task. The literature review, first critical appraisal and first academic-writing task each improved by five points. Table 3 presents the matched comparisons. One result demonstrates the importance of risk-sensitive interpretation. The original policy-analysis prompt scored 28/36 (77.8%), which would ordinarily appear favourable, but was force-rated Poor because a zero-scored failure-handling dimension coincided with maximum potential impact. The enhanced version scored 30/36 and did not trigger the same override. Aggregate scores alone would have concealed the original prompt’s lack of instructions for outdated, absent or conflicting evidence. Table 3. Matched multi-agent scores for original and enhanced prompts. Academic research task Original /36 Enhanced /36 Change Literature review 25 30 +5 Research design 1 25 29 +4 Critical appraisal 1 20 25 +5 Academic writing 1 22 27 +5 Policy analysis 1 28* 30 +2 Qualitative data analysis 22 25 +3 Research design 2 26 29 +3 Critical appraisal 2 25 26 +1 Academic writing 2 26 27 +1 Policy analysis 2 23 32 +9 *The original policy-analysis prompt was force-rated Poor by the critical-risk override despite its arithmetic score. 3.3 Failure handling and residual weaknesses Failure handling produced the most consistent dimension-level result. All ten original prompts scored 0/3, as did all ten raw prompts, while all ten enhanced prompts scored 3/3. Enhanced instructions more often required the model to identify absent inputs, disclose uncertainty, avoid fabricated citations, request clarification or state when a task could not be completed safely. The finding indicates that methodological expertise did not automatically translate into explicit planning for AI failure. Enhancement did not resolve every weakness. Worked examples remained absent from several prompts, limiting the model’s access to the expected depth and style. Some tasks still lacked source material that only a researcher could provide. The weakest enhanced academic-writing prompt reached the Good threshold but did not move comfortably beyond it, illustrating that structure and constraints cannot substitute for evidence or exemplars. 3.4 Adverse and security-relevant findings The most important adverse finding concerned the qualitative-data-analysis prompt. The original prompt contained no anonymisation or personally identifiable information instructions and received a moderate data-exposure concern. The enhanced prompt added a placeholder inviting the user to insert full transcripts but did not add anonymisation, data-minimisation or approved-platform requirements. Its security assessment therefore deteriorated. This result shows that aggregate prompt improvement can coexist with a more serious weakness in a specific domain. Prompts that accepted pasted studies or external documents were also vulnerable to embedded or malformed instructions, a form of prompt-injection exposure. The audit identified this risk but the enhanced versions did not consistently introduce robust input-separation or instruction-hierarchy safeguards. Automated enhancement was therefore better at improving explicit task specification than at guaranteeing data protection or security.

Conclusions:

4. Discussion 4.1 Principal findings and contribution This controlled pilot found a substantial increase in audit-rated prompt quality after automated enhancement of authentic higher education research prompts. The primary critic score rose by nearly thirty percentage points and every enhanced prompt reached the system’s Good band. The strongest result was not the aggregate mean but the uniform addition of failure handling. At the same time, a privacy deterioration in one enhanced prompt demonstrates why prompt quality must be treated as multidimensional and why automated improvement cannot be equated with responsible practice. The study contributes to higher education educational-technology research in three ways. First, it provides empirical evidence in a field where prompt engineering has often been discussed conceptually or through frameworks (Lee & Palmer, 2025; Walter, 2024). Second, it extends attention from classroom prompting to academic research work, complementing Dai and Chan’s (2026) call for task-sensitive guidance for researchers. Third, it proposes a provisional framework that connects prompt design to AI literacy, scholarly integrity and institutional governance. 4.2 Automated enhancement as an AI-literacy scaffold Within Ng et al.’s (2021) framework, the original prompts displayed considerable “Use and Apply” capability: they specified disciplinary roles, methods and analytic expectations. Their shared failure-handling deficit, however, suggests weaker operationalisation of “Evaluate and Create” and “AI Ethics”. The enhanced prompts more often required uncertainty disclosure, source verification and clarification, making evaluative and ethical requirements explicit. Automated enhancement can therefore be interpreted as a scaffold that exposes dimensions of AI literacy that experienced academics may understand in principle but omit when translating a task into an instruction. The term scaffold is important. In education, scaffolding should support capability development and be gradually internalised, rather than producing dependency. If researchers simply accept enhanced prompts without understanding the rationale for each addition, the tool may improve immediate specification while doing little to develop durable AI literacy. Institutions should therefore use prompt audits as learning artefacts: researchers should compare versions, explain why changes matter, identify inappropriate additions and revise the final prompt themselves. Such reflective activity would connect technical prompting with critical thinking and ethical judgement, consistent with higher education accounts of AI literacy (Dai & Chan, 2026; Walter, 2024). 4.3 Complementary expertise and distributed cognition The raw–original–enhanced gradient supports an interpretation of complementary expertise. Researcher authorship roughly doubled the critic score compared with the raw baseline by adding disciplinary scope, named methods, roles and evaluative criteria. Automated enhancement added a similar further gain, particularly through procedural structure and failure guidance. Domain expertise and automation were therefore additive rather than redundant in this sample. Distributed cognition helps explain this pattern. The prompt operates as an external representation through which disciplinary and procedural reasoning are coordinated across a human–AI system. The researcher contributes situated knowledge of what the task means, which evidence is relevant and what counts as a defensible conclusion. The automated system contributes a repeatable checklist of possible omissions. Neither contribution is sufficient alone. A generic prompt lacks disciplinary content; an automated rewrite cannot invent missing evidence, legitimate methodological decisions or context-specific ethical approval. The educational implication is that prompt competence should be taught as the orchestration of human expertise, external representations and verification practices, not as a search for a universal formula. 4.4 Failure handling as metacognitive and epistemic practice The zero-to-three shift in failure handling is theoretically and practically significant. Researchers commonly write instructions for successful task completion, not for conditions in which the AI lacks evidence, receives incomplete data or cannot verify a claim. Explicit failure handling functions as a metacognitive prompt: it asks the system and user to recognise the boundary between what is requested and what can be warranted. In academic work, this boundary is epistemic rather than merely operational. Examples include requiring a model to flag uncertain references, distinguish supplied evidence from background knowledge, request missing source material, stop rather than fabricate an analysis and identify when a conclusion exceeds the data. These practices do not remove hallucination or error, but they create visible points for human review. They also align with systems-based error management, in which safeguards are designed around predictable vulnerabilities rather than idealised performance (Reason, 2000). Failure handling should therefore be included in researcher development and institutional GenAI guidance as a core component of critical AI literacy. 4.5 Institutional governance and professional learning The findings have implications for universities seeking to move from high-level principles to actionable support. Chan’s (2023) ecological approach suggests that GenAI integration requires alignment across pedagogy, governance and operations. Prompt-quality assurance can contribute at each level. Pedagogically, it can support staff and postgraduate researcher development. From a governance perspective, prompts can document intended purpose, human authority, evidence expectations and data restrictions. Operationally, standardised review may help institutions embed approved tools, records and escalation pathways into research workflows. However, the privacy finding cautions against relying on an aggregate score or vendor-generated “Good” label. A university framework should require domain-specific safety gates. Prompts involving student records, interview transcripts, unpublished manuscripts, assessment data or commercially sensitive information need explicit data-minimisation and platform-approval checks. The same applies to documents that may contain embedded instructions. Institutions should also avoid making automated enhancement compulsory without evaluating accessibility, disciplinary fit, equity and the risk of deskilling. A practical professional-development activity would ask participants to write a prompt for an authentic task, compare it with an enhanced version, audit both against a shared framework and decide which changes to retain. The emphasis should be on explanation and judgement, not merely score improvement. Such an approach could support doctoral training, academic writing development, research methods education, continuing professional development and institutional research-integrity programmes. 4.6 Implications for curriculum, assessment and equitable implementation For educational technology practice, prompt quality should be developed through curriculum rather than treated as an informal skill acquired by trial and error. Research methods modules, doctoral training and staff development can use matched prompt comparison as a form of deliberate practice. Learners can identify what changed between an initial and enhanced instruction, connect each change to a methodological or ethical rationale, and then justify a final version. This shifts attention from collecting prompt templates to developing transferable judgement about purpose, evidence, uncertainty and accountability. It also provides a concrete way to integrate the “know and understand”, “use and apply”, “evaluate and create” and ethical dimensions of AI literacy described by Ng et al. (2021). Assessment of this capability should focus on process evidence rather than the apparent sophistication of a final prompt. A highly elaborate instruction can still be unsafe, epistemically weak or misaligned with the research question. Suitable assessment artefacts could include a prompt rationale, a record of revisions, identification of failure states, verification of cited evidence and a short reflection on what remained under human control. Such artefacts would allow educators to evaluate critical use without rewarding verbosity or dependence on a particular commercial platform. They could also support transparent disclosure of GenAI use in dissertations and research outputs, consistent with UNESCO’s (2023) emphasis on human agency, inclusion and accountability. Implementation also has equity implications. Automated enhancement may reduce barriers for researchers who are new to GenAI, working in an additional language or managing complex tasks under time pressure. Yet access to premium tools, confidence in challenging automated recommendations and familiarity with disciplinary conventions are unevenly distributed. Universities should therefore provide platform-neutral frameworks, examples from multiple disciplines, accessible training and alternatives for users who cannot or do not wish to use a commercial system. Support should be designed with students, researchers, librarians, learning developers, information-governance specialists and disability services rather than imposed as a single technical solution. Finally, prompt enhancement should sit within existing institutional risk management. The NIST AI Risk Management Framework emphasises governance, mapping context, measuring risk and managing identified harms (National Institute of Standards and Technology, 2023). In a UK setting, prompts involving personal data also require alignment with data-protection principles and approved processing environments (Information Commissioner’s Office, 2023). A high aggregate quality score should never override a failed privacy, security, intellectual-property or research-integrity gate. Institutions could operationalise this principle through tiered review: low-risk exploratory tasks may use a brief checklist, whereas work involving participant data, assessment decisions, unpublished findings or external documents should require explicit approval, data-minimisation and human verification. The educational value of a tool is therefore greatest when it makes governance requirements discussable and reviewable, not when it obscures them behind a single score. 4.7 Provisional Higher Education Research Prompt Quality Assurance Framework The original project proposed a healthcare research framework. The findings support reframing this as a provisional Higher Education Research Prompt Quality Assurance Framework (HE-RPQAF), with healthcare education as its first disciplinary application. The eight domains in Table 4 combine the empirical audit dimensions with AI-literacy and governance considerations. The framework is a human-readable developmental and review aid, not a validated compliance instrument. Table 4. Provisional Higher Education Research Prompt Quality Assurance Framework. Domain Review question AI-literacy and governance function 1. Purpose and alignment Is the academic purpose, intended user, audience and required outcome unambiguous? Connects tool use to legitimate educational or research goals. 2. Context and scope Are the setting, population, timeframe, definitions, inclusions, exclusions and source materials sufficient? Reduces task drift and unsupported assumptions. 3. Human role and accountability Is the AI role bounded and are human review, authorship, decision authority and responsibility explicit? Protects agency and prevents inappropriate delegation. 4. Evidence and source integrity Does the prompt require traceable evidence, prohibit invented sources and distinguish supplied material from model knowledge? Supports scholarly integrity and verification. 5. Critical and evaluative reasoning Does the task require synthesis, counter-evidence, limitations, bias and justified conclusions? Develops evaluate-and-create capability rather than passive acceptance. 6. Output specification Are format, length, academic level, headings, citation expectations and reporting requirements clear? Improves consistency, usability and transparency. 7. Failure handling and uncertainty Does the prompt state what to do when information is missing, conflicting, unverifiable or unsafe? Creates stop, clarify and limitation pathways. 8. Ethics, privacy and security Are data minimisation, anonymisation, approved-platform use, bias, intellectual property and injection risks addressed? Operationalises ethical AI literacy and institutional governance. The framework requires prospective testing for content validity, usability, inter-rater reliability and association with downstream output quality. 4.8 Strengths and limitations A strength of the study was the matched-task design: the academic task remained constant while prompt formulation changed. All conditions used the same audit settings, enhanced prompts were analysed without post-processing, and adverse findings were retained. The inclusion of a risk override also demonstrated why high average scores should not automatically determine acceptability. The limitations are substantial. The sample comprised ten tasks from one UK researcher in one disciplinary context. Prompts were purposively selected and conditions were processed in separate, non-randomised batches. The study did not include students, postgraduate researchers or staff as learners and therefore did not measure AI-literacy development. The scorer and enhancer shared a model lineage and rubric, creating possible circularity and self-preference; although LLM-based evaluators can align with human judgements in some natural-language-generation tasks, they remain sensitive to model and rubric design (Liu et al., 2023; Panickssery et al., 2024). No independent human raters assessed prompt quality, no inter-rater reliability was available and the proprietary rubric limits reproducibility unless disclosed. Most importantly, the study assessed prompts rather than the outputs they produce. A higher prompt score cannot be assumed to generate a more accurate literature review, a more valid thematic analysis or a better educational outcome. The findings demonstrate improved explicitness according to the evaluated rubric, not improved truth, learning, research quality or impact. Generalisation beyond English-language healthcare education research requires independent, international and multi-disciplinary testing. 4.9 Future research The next phase should recruit multiple researchers and institutions and include 50–100 prompts across disciplines, career stages and languages. Original and enhanced prompts should be submitted to the same GenAI models under controlled settings and repeated to assess variability. Blinded domain experts should rate downstream outputs for accuracy, completeness, relevance, criticality, evidence fidelity, usefulness, privacy and safety. Human ratings should be compared with automated scores and alternative tools. The HE-RPQAF should be evaluated as both a measurement framework and an educational intervention. Studies could assess content validity, inter-rater reliability and predictive validity, then test whether structured prompt-review workshops improve researchers’ AI literacy, calibration of trust and ability to detect unsafe enhancement. Longitudinal research should examine whether users internalise the framework or become dependent on automated rewriting. Cross-cultural work is needed because institutional rules, research ethics, language practices and conceptions of authorship vary internationally. 5. Conclusion This pilot provides promising evidence that automated prompt enhancement can strengthen the audit-rated quality of AI-assisted research instructions in higher education. In authentic healthcare education tasks, enhancement increased mean critic scores from 52.9% to 82.5%, improved every matched multi-agent score and consistently added failure handling that was absent from researcher-authored prompts. The findings suggest that disciplinary expertise and automated scaffolding can contribute complementary forms of value. The positive case is nevertheless bounded. One enhanced prompt increased privacy risk, and the study did not evaluate downstream answers or learning. Prompt-enhancement systems should therefore be used as critical AI-literacy and governance aids, not as substitutes for disciplinary expertise, research ethics or source verification. A structured system such as PromptMatrix can make missing context, uncertainty and failure states more visible, particularly when its recommendations are reviewed against the HE-RPQAF and revised by an accountable researcher. With independent human validation and multi-institutional testing, prompt-quality assurance could become a useful component of responsible GenAI professional development and research governance across higher education.


 Citation

Please cite as:

Pollock N

Automated prompt enhancement for AI-assisted research in higher education: A controlled pilot study in healthcare education

JMIR Preprints. 14/09/2026:111925

DOI: 10.2196/preprints.111925

URL: https://preprints.jmir.org/preprint/111925

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.