Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: Journal of Medical Internet Research

Date Submitted: May 11, 2025
Open Peer Review Period: May 12, 2025 - Jul 7, 2025
Date Accepted: Feb 25, 2026
(closed for review but you can still tweet)

The final, peer-reviewed published version of this preprint can be found here:

Impact of Large Language Model–Based AI Tools on Physician-Patient Communication: Systematic Review and Meta-Analysis

Richter S, Buszello CH, Prem M, Willkommen S, Hasani E, Uckermann O, Sandi-Gahun S, Juratli TA, Eyüpoglu IY, Polanski WH

Impact of Large Language Model–Based AI Tools on Physician-Patient Communication: Systematic Review and Meta-Analysis

J Med Internet Res 2026;28:e77307

DOI: 10.2196/77307

PMID: 42536971

Impact of Large Language Model–Based AI Tools on Physician–Patient Communication: A Systematic Review and Meta-Analysis

  • Sven Richter; 
  • Clara H. Buszello; 
  • Markus Prem; 
  • Sophia Willkommen; 
  • Elida Hasani; 
  • Ortrud Uckermann; 
  • Sahr Sandi-Gahun; 
  • Tareq A. Juratli; 
  • Ilker Y. Eyüpoglu; 
  • Witold H. Polanski

ABSTRACT

Background:

Recent advances in large language models (LLMs) such as GPT-3/4 have spurred development of AI chatbots and advisory tools in medicine. These systems are posited to assist or augment physician–patient communication, potentially improving empathy, clarity, and responsiveness. However, their actual impact on communication outcomes remains uncertain.

Objective:

To systematically review and meta-analyze peer-reviewed studies (2020–2025) evaluating how LLM-based interventions affect physician–patient communication, including empathy, clarity, trust, and patient understanding.

Methods:

Following PRISMA 2020 guidelines, we searched PubMed/MEDLINE for studies published from 2020 to 2025 examining LLM or chatbot applications in clinical communication contexts. Eligible designs included randomized, observational, cross-sectional, and qualitative studies. Two reviewers independently screened titles/abstracts, assessed full texts, and extracted data on study design, population, LLM type, communication measures, and outcomes. We conducted a qualitative synthesis and random-effects meta-analysis, reporting pooled standardized mean differences (SMD) or odds ratios (OR) with 95% confidence intervals (CI).

Results:

From 312 records, 10 studies (N=10) were included, all quantitative and predominantly cross-sectional. Populations ranged from patients with chronic conditions to healthcare professionals and laypersons. Outcomes assessed included empathy (7 studies), clarity/information quality (6), satisfaction or usefulness (4), and trust perceptions (2). In six direct comparisons of AI- versus physician-generated responses, LLMs were rated significantly higher in empathy in five studies. One large study found chatbot replies were judged empathetic in 45.1% of cases versus 4.6% for physician replies (OR ~9.8, P<.001). Similarly, ChatGPT-4 answers scored higher in empathy on a 5-point scale than human-written responses (mean 4.18 vs 2.70, P<.001). One neurology study showed higher empathy scores (CARE scale +1.38, P<.01) for ChatGPT answers. Only one study found no significant empathy difference. LLM content was also longer and more information-rich, improving patient-perceived clarity and understanding. On the other hand, GPT-4 simplified pathology reports, increasing patient comprehension scores (7.98 vs 5.23/10, P<.001) and reducing consultation time by 70%. However, AI replies were sometimes less concise or less readable for low-literacy patients. In pooled analyses (4 studies, n=2,604), LLMs showed a large positive effect on empathy (SMD +1.05, 95% CI 0.45–1.65) and improved understanding (SMD +0.82, 95% CI 0.30–1.34). Patient satisfaction results were mixed. No study directly assessed long-term trust.

Conclusions:

Current evidence suggests LLM-based chatbots can enhance physician–patient communication by producing more empathetic, detailed, and understandable responses. These improvements may positively influence patient experience and engagement. However, LLMs may also generate overly lengthy or occasionally inaccurate advice, emphasizing the need for physician oversight. While meta-analytic findings are promising, robust randomized, controlled trials are needed to confirm benefits, assess trust outcomes, and define optimal clinical integration strategies.


 Citation

Please cite as:

Richter S, Buszello CH, Prem M, Willkommen S, Hasani E, Uckermann O, Sandi-Gahun S, Juratli TA, Eyüpoglu IY, Polanski WH

Impact of Large Language Model–Based AI Tools on Physician-Patient Communication: Systematic Review and Meta-Analysis

J Med Internet Res 2026;28:e77307

DOI: 10.2196/77307

PMID: 42536971

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.