Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: JMIR Human Factors

Date Submitted: Jan 29, 2026
Date Accepted: Aug 17, 2026

The final, peer-reviewed published version of this preprint can be found here:

What Users Say About Reimbursable Digital Therapeutics in Germany: Large-Scale App Store Review Analysis Using a Large Language Model

Kinast B, Rohde H, Schreiweis B, Ulrich H

What Users Say About Reimbursable Digital Therapeutics in Germany: Large-Scale App Store Review Analysis Using a Large Language Model

JMIR Hum Factors 2026;13:e92415

DOI: 10.2196/92415

PMID: 42849059

What Users Say About Reimbursable Digital Therapeutics in Germany: A Large-Scale App Store Review Analysis Using a Large Language Model

  • Benjamin Kinast; 
  • Henrik Rohde; 
  • Björn Schreiweis; 
  • Hannes Ulrich

ABSTRACT

Background:

In 2019, Germany has introduced a unique regulatory framework for Digital Therapeutics (DTx) known as DiGAs, with the goal of integrating evidence-based DTx into statutory healthcare. DTx are approved for statutory health insurance reimbursement if the manufacturers can show evidence of health improvement, better coordination of care process, easier access to healthcare services, or the promotion of health literacy in a controlled study setting. Systematic investigations have revealed shortcomings in the evidence provided by manufacturers.

Objective:

The study examines patients experience and evaluate DTx in public App Stores by analyzing sentiments and thematic categories. In addition, it explores if the use of large language model (LLM) can provide effectively support the categorization and sentiment analysis of the patients’ reviews.

Methods:

First, a list of approved DTx is collected and limited to those with mobile apps. Patients’ reviews were extracted from the public app stores using a tailored Python script. In the second step, the sentiments and topics of the patients’ reviews were categorized into ten predefined categories using ChatGPT-4o. To ensure the quality, the LLM-based analysis was verified through manual validation.

Results:

In total, 44 mobile DiGAs were included and were analyzed. After data extraction and cleansing, the final dataset comprises 4,328 patients' reviews containing at least one interpretable statement, resulting in 9,439 valid and interpretable statements. A systematic validation of the automated classification demonstrated exceptionally high model performance with 99% accuracy for sentiment classification (F1-scores of 1.00 for positive and 0.99 for negative categories) and 95% accuracy for category classification with an average F1-score of 0.95. While the categories Overall Impression and Effectiveness scored particularly well, patients were most negative for topics related to the login and registration process as well as technical malfunctions.

Conclusions:

The findings highlight both the potential and current limitations of reimbursable DTx from the patients’ perspective. While patients value the therapeutic benefits and content quality of many DTx, technical functionality and usability, particularly during login and registration, are frequently criticized. The study also demonstrates that LLM-assisted analysis combined with human-in-the-loop validation offers an efficient approach for structuring patient feedback at scale. This combined approach could serve as a valuable complement to traditional clinical evaluation in future assessments of digital health applications.


 Citation

Please cite as:

Kinast B, Rohde H, Schreiweis B, Ulrich H

What Users Say About Reimbursable Digital Therapeutics in Germany: Large-Scale App Store Review Analysis Using a Large Language Model

JMIR Hum Factors 2026;13:e92415

DOI: 10.2196/92415

PMID: 42849059

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.