Accepted for/Published in: JMIR Medical Informatics
Date Submitted: Jan 21, 2026
Date Accepted: Jun 17, 2026
Influenza-Like Illness Forecasting Using Multi-Source Data: A Comparative Deep Learning Study
ABSTRACT
Background:
Background:
Accurate forecasting of Influenza-like Illness (ILI) is crucial for public health. Integrating novel digital data streams (e.g., internet searches, human mobility) with traditional surveillance can improve accuracy, but optimal modeling frameworks are underexplored.
Objective:
Objective:
To develop and compare multi-source data-driven models for forecasting ILI incidence trends.
Methods:
Methods:
Weekly ILI incidence data and multi-source variables for Hubei Province from 2020 to 2023 were collected. Multisource predictors included mean temperature, relative humidity, air quality index, a synthesized Baidu Search Index (SBI) for influenza-related queries, the Baidu Migration Scale Index (BMS) for population mobility, and the Oxford Stringency Index (SI) for non-pharmaceutical interventions. Predictive models evaluated were: Seasonal Autoregressive Integrated Moving Average (SARIMA), Long Short-Term Memory (LSTM), a hybrid Convolutional Neural Network-LSTM (CNN-LSTM), Transformer, and Random Forest (RF). Models were constructed using an 85:15 training-test split and optimized via grid search with 5-fold cross-validation. Performance was assessed using Mean Absolute Error (MAE), Root Mean Squared Error (RMSE), Mean Absolute Percentage Error (MAPE), and the Coefficient of Determination (R²).
Results:
Results:
All models incorporating multisource data substantially outperformed univariate time-series benchmarks. Among univariate models, CNN-LSTM achieved superior performance (R² = 0.7623) over SARIMA (R² = −0.2645). In multisource configurations, the LSTM model demonstrated the strongest predictive capability and feature integration, attaining an optimal R² of 0.8350 when combining environmental, mobility (BMS), and internet search (SBI) data. The inclusion of the Oxford Stringency Index (SI) consistently and significantly reduced prediction error across all models, with MAPE decreasing by up to 84.87% in the LSTM model. The Baidu Search Index emerged as the most influential single external predictor, notably enhancing model fit. In contrast, model performance varied with feature composition: Transformer excelled with full feature sets, while CNN-LSTM’s efficacy peaked primarily with SBI integration. Over a 26-week projection, the optimal LSTM model forecasts a “rapid decline–gradual decline–stabilization” trend, indicating a return to baseline ILI activity.
Conclusions:
Conclusion: Deep learning models, particularly LSTM, can effectively leverage heterogeneous digital data to improve ILI forecasting. Internet search behavior and policy stringency are critical external predictors. Integrating real-time digital data with conventional surveillance offers a robust approach for enhancing early warning systems and supporting public health decision-making.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.