Previously submitted to: Journal of Medical Internet Research (no longer under consideration since Jan 09, 2025)
Date Submitted: Oct 11, 2023
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
A COVID-19 Prediction Model based on Symptomatology Google Trends Big Data and its Optimization
ABSTRACT
Background:
The accurate prediction of the development trend of COVID-19 pandemic can provide important guidance for the prevention and control of the pandemic. The Google Trends data of typical symptoms of COVID-19 infection such as loss of smell and taste have a significant correlation with the occurrence and development trend of COVID-19 epidemic.
Objective:
This study is based on Google Trends data of symptomatology, aiming to build and optimize the prediction models of the COVID-19 epidemic.
Methods:
Correlation analysis is carried out on the data of number of confirmed cases from Johns Hopkins University’s COVID-19 database and the symptomatology Google Trends data from the Google Trends platform. SEIR model, ARIMAX model and LSTNet model are established respectively based on the Google Trends data, and their predictive performance are compared.
Results:
There was a strong correlation between different symptomatology Google Trends data such as smell loss, taste loss, cough, fever, sore throat, and shortness of breath (all correlation coefficients > 0.5, P<0.001). During the Delta variants predominated period, there is 2336 new cases every day for each Google Trends data of taste loss increased (P<0.05), during that of Omicron variants predominated, the number is 3361 (P<0.05). It is suggested that the correlation is always strong in the epidemic periods of different SARS-COV-2 variants. Furthermore, Principal Component Analysis (PCA) is conducted, results showed that the cumulative contribution rate reached 93.04%. The Google Trends data after PCA (PCAGT) is introduced into different prediction models to improve the performance, and the results showed the SEIR model could predict the number of daily newly confirmed cases in the next 7 days with an MAE of 5.30×103 and an RMSE of 6.60×103. Comparing the error of the predicted value and the actual value within 5 weeks, after incorporating PCAGT, the error of the ARIMAX model drops from 16.5% to 0.2%, and the error of the LSTNet model drops from 11.3% to 10.1%, and the improvement effect of ARIMAX is better when PCAGT was included. The ARIMAX model incorporating PCAGT has the best prediction performance with a prediction error of only 0.2% within 5 weeks.
Conclusions:
Incorporating symptomology-related Internet big data can generally improve the performance of the COVID-19 prediction models, and the prediction accuracy of the ARIMAX prediction model was the highest, which can be used to predict the trend of a COVID-19 epidemic.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.