Accepted for/Published in: Journal of Medical Internet Research
Date Submitted: Apr 22, 2026
Open Peer Review Period: Apr 22, 2026 - Jun 17, 2026
Date Accepted: Jun 23, 2026
(closed for review but you can still tweet)
LLM-generated Lay-language Protocols for Molecular Tumor Board Patients: Evaluation of Quality and Clinical Usability
ABSTRACT
Background:
Molecular tumor boards (MTBs) generate highly technical recommendations. The language used in their protocols is rarely accessible to patients. Lay-language patient protocols could support patient-clinician communication, yet manual production is difficult to sustain in high-volume oncology settings. Large language models (LLMs) may offer scalable drafting assistance, yet clinical usability remains largely uninvestigated under real-world deployment constraints. Existing evaluations rely predominantly on synthetic data or closed-source models that are incompatible with strict data protection requirements.
Objective:
This study evaluated whether open-weight LLMs can provide clinically usable drafting support for German MTB patient protocols under real-world deployment constraints and developed a transferable evaluation framework for patient-facing text generation.
Methods:
Eight open-weight LLMs were evaluated under zero-shot (A1) and one-shot (A2) prompting with constrained decoding, which ensures section-schema compliance. Automatic evaluation used ROUGE-1, BERTScore-F1, WSTF4, and DistilBERT-based complexity using a corpus of 316 MTB protocols and 47 expert-written patient protocols. For expert evaluation, seven medical oncologists evaluated 50 protocols from the best-performing model across three ISO 9241-11 usability dimensions using fine-grained error annotation, perceived post-editing effort (PPEE), and net promoter score (NPS). Critical errors were defined as bearing the risk of patient harm.
Results:
Llama-3.3-70B-Instruct achieved the strongest automatic performance. Across models, A2 significantly improved most automatic metrics compared to A1. However, expert usability evaluation showed the opposite picture: the proportion of protocols containing at least one critical error doubled under A2 (40% vs. 20%) compared with A1, and the dominant error type shifted from language (37%) errors to factual errors (48%). Overall, 6.1% of the annotated paragraphs contained errors. Median PPEE was 2 (low) and median NPS was 7. Detractors (46%) outweighed promoters (29%), which signals clinical hesitation toward routine adoption.
Conclusions:
Prompting strategies that improve automatic metrics can simultaneously increase the number of critical errors. Surface-level metric gains were, therefore, insufficient proxies for clinical safety. Nonetheless, the low paragraph-level error rate and favorable PPEE suggest that structured open-weight LLM generation may be a useful drafting aid in a clinician-supervised setting. The proposed evaluation framework establishes a text-quality-focused basis for future assessment of patient-facing LLM applications in real-world clinical settings.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.