Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: JMIR Infodemiology

Date Submitted: Apr 15, 2026
Open Peer Review Period: Apr 27, 2026 - Jun 22, 2026
Date Accepted: Aug 31, 2026
(closed for review but you can still tweet)

The final, peer-reviewed published version of this preprint can be found here:

Identifying Public Response Topics to the Centers for Disease Control and Prevention’s COVID-19 Communications on Social Media: Infoveillance Study Using Large Language Model–Based Rephrasing

Xin W, Yin S, Park A, Ge Y, Chen S

Identifying Public Response Topics to the Centers for Disease Control and Prevention’s COVID-19 Communications on Social Media: Infoveillance Study Using Large Language Model–Based Rephrasing

JMIR Infodemiology 2026;6:e98319

DOI: 10.2196/98319

Identifying Public Response Topics to CDC COVID-19 Communications on Social Media: Infoveillance Study Using Large Language Model–Based Rephrasing

  • Wangjiaxuan Xin; 
  • Shuhua Yin; 
  • Albert Park; 
  • Yaorong Ge; 
  • Shi Chen

ABSTRACT

Background:

Public health agencies increasingly use social media platforms such as X to monitor public responses to official health communications and support timely decision-making during health crises. Public replies to communications from official agencies provide valuable insights into population-level discourse and engagement with health policies. However, analyzing large-scale short-text responses remains challenging because social media replies are often brief, informal, ambiguous, and linguistically fragmented. Topic modeling offers a scalable approach for identifying thematic patterns in such data, but its performance is often limited when applied to social media short texts. Recent advances in large language models provide new opportunities to improve text normalization before topic modeling, but their value and limitations for public health infoveillance remain underexplored.

Objective:

This study develops and evaluates TM-Rephrase, a model-agnostic LLM-based rephrasing framework designed to improve the performance of topic models for short public health social media texts while ensuring rephrasing preserves the original semantic meaning, stance/intent, and tone.

Methods:

We analyzed 25,027 public replies to official CDC posts on X collected between May 2020 and November 2022. TM-Rephrase transformed informal and noisy short texts into more standardized expressions using two prompt-guided schemes: general rephrasing and colloquial-to-formal rephrasing. Original and rephrased texts were analyzed using multiple topic models, with rephrasing generated by Gemini-2.5-flash, GPT-4o-mini, and Mistral-7B-Instruct. Topic quality was evaluated using coherence, uniqueness, redundancy, and diversity. We also conducted an expert semantic-fidelity validation study in which four expert raters evaluated whether rephrased texts preserved the original meaning, stance/intent, and tone.

Results:

TM-Rephrase generally improved topic quality across models, rephrasing schemes, and LLM backbones. For LDA, C_v coherence increased from 0.3094 without rephrasing to 0.5004 with colloquial-to-formal rephrasing. For BERTopic, C_v coherence improved from 0.4078 to 0.4734. Diversity-related metrics also improved in most rephrasing settings. Additional checks using C_NPMI, C_UCI, and U_MASS generally supported the main improvement pattern. Sensitivity analysis across different topic numbers further suggested that the coherence gains were not limited to the main K = 8 setting. Expert validation showed favorable semantic preservation overall, with a mean rating of 3.90 and median of 4.00 out of 5.00, and acceptable inter-rater reliability, ICC(2,k) = 0.736. General rephrasing received stronger semantic-fidelity ratings than colloquial-to-formal rephrasing, suggesting that stronger formalization may improve topic readability while introducing greater risk of changes to the original meaning, stance/intent, and tone.

Conclusions:

TM-Rephrase can enhance the coherence, distinctiveness, and interpretability of topic modeling outputs for short, noisy public health social media data. However, TM-Rephrase should be viewed as a text-normalization preprocessing strategy for topic enhancement rather than a fully neutral replacement for original public discourse. In public health infoveillance, original and rephrased texts should be interpreted together, especially when sarcasm, distrust, uncertainty, anger, or oppositional framing are analytically important.


 Citation

Please cite as:

Xin W, Yin S, Park A, Ge Y, Chen S

Identifying Public Response Topics to the Centers for Disease Control and Prevention’s COVID-19 Communications on Social Media: Infoveillance Study Using Large Language Model–Based Rephrasing

JMIR Infodemiology 2026;6:e98319

DOI: 10.2196/98319

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.