Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: Journal of Medical Internet Research

Date Submitted: Apr 22, 2026
Open Peer Review Period: Apr 22, 2026 - Jun 17, 2026
Date Accepted: Jun 23, 2026
(closed for review but you can still tweet)

The final, peer-reviewed published version of this preprint can be found here:

LLM-Generated Lay-Language Protocols for Molecular Tumor Board Patients: Evaluation of Quality and Clinical Usability

Pakull TMG, Bender N, Benson S, Fleischhauer A, Alsara M, Gromke T, Hilser T, Kaminski K, Pogorzelski M, Prasuhn N, Rosery V, Schadendorf D, Schuler M, Wiesweg M, Zaun G, Horn PA, Friedrich CM, Pretzell I

LLM-Generated Lay-Language Protocols for Molecular Tumor Board Patients: Evaluation of Quality and Clinical Usability

J Med Internet Res 2026;28:e99136

DOI: 10.2196/99136

LLM-generated Lay-language Protocols for Molecular Tumor Board Patients: Evaluation of Quality and Clinical Usability

  • Tabea Margareta Grace Pakull; 
  • NoĆ«lle Bender; 
  • Sven Benson; 
  • Anke Fleischhauer; 
  • Mohammad Alsara; 
  • Tanja Gromke; 
  • Thomas Hilser; 
  • Katharina Kaminski; 
  • Michael Pogorzelski; 
  • Nicola Prasuhn; 
  • Vivian Rosery; 
  • Dirk Schadendorf; 
  • Martin Schuler; 
  • Marcel Wiesweg; 
  • Gregor Zaun; 
  • Peter Alexander Horn; 
  • Christoph Matthias Friedrich; 
  • Ina Pretzell

ABSTRACT

Background:

Molecular tumor boards (MTBs) generate highly technical recommendations. The language used in their protocols is rarely accessible to patients. Lay-language patient protocols could support patient-clinician communication, yet manual production is difficult to sustain in high-volume oncology settings. Large language models (LLMs) may offer scalable drafting assistance, yet clinical usability remains largely uninvestigated under real-world deployment constraints. Existing evaluations rely predominantly on synthetic data or closed-source models that are incompatible with strict data protection requirements.

Objective:

This study evaluated whether open-weight LLMs can provide clinically usable drafting support for German MTB patient protocols under real-world deployment constraints and developed a transferable evaluation framework for patient-facing text generation.

Methods:

Eight open-weight LLMs were evaluated under zero-shot (A1) and one-shot (A2) prompting with constrained decoding, which ensures section-schema compliance. Automatic evaluation used ROUGE-1, BERTScore-F1, WSTF4, and DistilBERT-based complexity using a corpus of 316 MTB protocols and 47 expert-written patient protocols. For expert evaluation, seven medical oncologists evaluated 50 protocols from the best-performing model across three ISO 9241-11 usability dimensions using fine-grained error annotation, perceived post-editing effort (PPEE), and net promoter score (NPS). Critical errors were defined as bearing the risk of patient harm.

Results:

Llama-3.3-70B-Instruct achieved the strongest automatic performance. Across models, A2 significantly improved most automatic metrics compared to A1. However, expert usability evaluation showed the opposite picture: the proportion of protocols containing at least one critical error doubled under A2 (40% vs. 20%) compared with A1, and the dominant error type shifted from language (37%) errors to factual errors (48%). Overall, 6.1% of the annotated paragraphs contained errors. Median PPEE was 2 (low) and median NPS was 7. Detractors (46%) outweighed promoters (29%), which signals clinical hesitation toward routine adoption.

Conclusions:

Prompting strategies that improve automatic metrics can simultaneously increase the number of critical errors. Surface-level metric gains were, therefore, insufficient proxies for clinical safety. Nonetheless, the low paragraph-level error rate and favorable PPEE suggest that structured open-weight LLM generation may be a useful drafting aid in a clinician-supervised setting. The proposed evaluation framework establishes a text-quality-focused basis for future assessment of patient-facing LLM applications in real-world clinical settings.


 Citation

Please cite as:

Pakull TMG, Bender N, Benson S, Fleischhauer A, Alsara M, Gromke T, Hilser T, Kaminski K, Pogorzelski M, Prasuhn N, Rosery V, Schadendorf D, Schuler M, Wiesweg M, Zaun G, Horn PA, Friedrich CM, Pretzell I

LLM-Generated Lay-Language Protocols for Molecular Tumor Board Patients: Evaluation of Quality and Clinical Usability

J Med Internet Res 2026;28:e99136

DOI: 10.2196/99136

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.