Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: JMIR AI

Date Submitted: Jun 3, 2025
Date Accepted: Jun 29, 2026

The final, peer-reviewed published version of this preprint can be found here:

Benchmarking AI-Powered Translation of the EQ-5D-5L Patient-Reported Outcome Measure Using Automated Metrics: Comparative Evaluation Study

Vashisht H, Ward T, Muehlhausen W

Benchmarking AI-Powered Translation of the EQ-5D-5L Patient-Reported Outcome Measure Using Automated Metrics: Comparative Evaluation Study

JMIR AI 2026;5:e78485

DOI: 10.2196/78485

PMID: 42550950

Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.

Benchmarking AI-Powered Translation for eCOA: A Comparative Evaluation of EQ-5D-5L Translations Using Machine Translation Metrics

  • Himanshu Vashisht; 
  • Tomás Ward; 
  • Willie Muehlhausen

ABSTRACT

Background:

Patient-reported outcome (PRO) measures are pivotal in clinical research for collecting patient insights directly. With the increasing global reach of clinical trials and the rise of AI, ensuring accurate and culturally sensitive translation of PROs, specifically electronic clinical outcome assessments (eCOAs), is paramount.

Objective:

This study aims to evaluate the quality and comparability of AI-powered translation services for the EQ-5D-5L, a widely used generic PRO instrument, across multiple languages using various machine translation metrics.

Methods:

This study used a comparative evaluation design. The EQ-5D-5L English version was translated into five target languages (Danish, Dutch, French, German, Spanish) by four AI services (Google Translate, GPT-4o, Amazon Translate, DeepL). Translation quality was assessed using four automated machine translation metrics: BLEU, METEOR, COMET, and BLEURT. Statistical analysis, including Kruskal-Wallis H-tests, was employed to compare the performance of the AI services. Spearman's rank correlation was used to assess the consistency between services and metrics.

Results:

The Kruskal-Wallis H-tests revealed statistically significant differences (P<.05) in translation quality across AI services for only one metric—METEOR—specifically for English to French translations. For all other languages and metrics, no statistically significant differences were found, indicating comparable performance among the AI services. Spearman's rank correlation coefficients showed moderate to strong positive correlations (ranging from 0.5 to 0.9) between the scores generated by different AI services across most metrics and languages, suggesting a good level of consistency. However, some variability was observed, particularly for lower-resource languages. The sample size was 43 sentences translated for each language pair.

Conclusions:

The findings suggest that current AI translation services provide largely comparable and consistent quality for translating the EQ-5D-5L instrument across common European languages based on automated metrics. While some variability exists, particularly for specific language pairs and metrics, the overall high level of agreement supports the potential utility of these tools in facilitating multilingual eCOA deployment. Further research incorporating human evaluation is recommended to validate these findings and explore nuanced aspects of translation quality.


 Citation

Please cite as:

Vashisht H, Ward T, Muehlhausen W

Benchmarking AI-Powered Translation of the EQ-5D-5L Patient-Reported Outcome Measure Using Automated Metrics: Comparative Evaluation Study

JMIR AI 2026;5:e78485

DOI: 10.2196/78485

PMID: 42550950

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.