Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Previously submitted to: Journal of Medical Internet Research (no longer under consideration since May 17, 2025)

Date Submitted: Jan 21, 2025

Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.

Self-Logical Consistency Assessment of Large Language Models for Patient Feedback Classification : Algorithm Development and Validation Study

  • Zeno Loi; 
  • David Morquin; 
  • François-Xavier Derzko; 
  • Xavier Corbier; 
  • Sylvie Gauthier; 
  • Patrice Taourel; 
  • Emilie Prin-Lombardo; 
  • Grégoire Mercier; 
  • Kévin Yauy

ABSTRACT

Background:

Patient satisfaction feedback is crucial for hospital service quality, but manual reviews are time-consuming, and traditional natural language processing methods remain inadequate. Large Language Models (LLMs) show promise but are prone to extrinsic faithfulness hallucinations—fabricated or illogical outputs that limit their reliability in healthcare.

Objective:

This study aimed to evaluate the Self-Logical Consistency Assessment (SLCA), an original method designed to enhance LLM feedback classification reliability by enforcing a logically-structured chain of thought.

Methods:

SLCA uses two validation steps: self-consistency (identifying the most coherent response) and logical consistency (ensuring alignment with the original statement and expert classifications). We evaluated SLCA using GPT-4 and Llama-3.1 405B on 12,600 classifications from 100 patient feedback samples to assess hallucinations, and tested its performance on a 49,140-classification benchmark derived from 1,170 feedbacks.

Results:

SLCA reduced hallucinations among detected categories from 15.80% (168/1063) to 0.51% (4/786) with GPT-4 and from 7.17% (51/711) to 1.67% (10/599) with Llama-3.1, with residual errors confined to the emergency feedback category. On the benchmark, SLCA achieved precision-recall scores of 0.86-0.78 for GPT-4 and 0.84-0.58 for Llama-3.1. These results demonstrate SLCA’s ability to achieve human-level performance across LLMs.

Conclusions:

SLCA offers a scalable, explainable solution for improving LLM classification reliability in healthcare. Its capacity to enhance performance without fine-tuning or additional training data positions it as a valuable tool for analyzing patient feedback and supporting hospital service quality improvement.


 Citation

Please cite as:

Loi Z, Morquin D, Derzko FX, Corbier X, Gauthier S, Taourel P, Prin-Lombardo E, Mercier G, Yauy K

Self-Logical Consistency Assessment of Large Language Models for Patient Feedback Classification : Algorithm Development and Validation Study

JMIR Preprints. 21/01/2025:71571

DOI: 10.2196/preprints.71571

URL: https://preprints.jmir.org/preprint/71571

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.