Previously submitted to: JMIR Public Health and Surveillance (no longer under consideration since Feb 11, 2021)
Date Submitted: Nov 14, 2020
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Mask Usage Classification of Tweets and its Application to the COVID-19 Pandemic
ABSTRACT
Background:
Face masks have been recommended or mandated at different points in time during the ongoing COVID-19 pandemic. The effectiveness of masks has been understood to be linked to their adoption. The scale of adoption of masks in the community remains to be understood.
Objective:
Given the popularity of social media, we aim to use tweets (social media posts on Twitter) to analyse discussions mentioning masks and, specifically, tweets that report mask usage, as an indication of mask adoption.
Methods:
We used a repository of tweets from Australia and New Zealand to create a dataset of mask-related tweets posted from 2017 to 2020, specifically tweets containing five mask-related words: face mask, surgical mask, cloth mask, N95 mask and P2 mask. From this dataset, we created a manually annotated dataset of 3016 tweets labeled with mask usage. We first used topic modeling separately on each mask type to understand the context in which these mask types are mentioned. We then proposed mask usage classification: automatic task of predicting if a tweet reports usage or intention of usage of a mask. We experimented with a set of classification approaches based on BERT and trained on the manually annotated dataset. Finally, we applied mask usage classification on the whole dataset of mask-related tweets to analyse trends in mask usage reports in 2020.
Results:
Concerns around availability of masks during the COVID-19 pandemic or bushfires appear in the topics, demonstrating the contexts in which mask types are mentioned. Our classifier for mask usage with oversampled instances results in the best performance with an F-score of 85.02%. When the classifier is applied to our dataset, we observe peaks in mask usage reporting in tweets shortly after key news articles related to masks were published in news websites.
Conclusions:
This work focuses on mask usage-related discussions on social media. We introduce a novel dataset of tweets annotated for mask usage, and a set of classification approaches based on the dataset. When applied to the datasets of tweets mentioning masks, we observed that events such as masks being recommended or being made mandatory in Victoria resulted in an increase in mask usage reports on social media within a few days. The utility of mask usage classification lies in its ability to predict mask usage reporting on Twitter so as to understand the effectiveness of mask advisory during the ongoing COVID-19 pandemic and beyond. Clinical Trial: N/A
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.