Skip to main content
Proceedings of the National Academy of Sciences of the United States of America logoLink to Proceedings of the National Academy of Sciences of the United States of America
. 2025 Aug 13;122(33):e2422335122. doi: 10.1073/pnas.2422335122

Ensemble approaches for short-term dengue fever forecasts: A global evaluation study

Skyler Wu a,b,1, Austin G Meyer b,c,1,2, Leonardo Clemente b, Lucas M Stolerman d, Fred Lu b, Atreyee Majumder e, Rudi Verbeeck e, Serge Masyn e, Mauricio Santillana b,f,2
PMCID: PMC12377650  PMID: 40802689

Significance

Dengue fever affects nearly 400 million people annually, creating an urgent need for improved outbreak forecasting tools. We developed ensemble forecasting methods that combine multiple predictive models to forecast dengue cases 1 to 3 mo in more than 180 locations around the world. While individual forecasting models show variable performance across different regions, our ensemble approaches consistently rank among the top performers, offering improved predictions even when individual models falter. Our methods demonstrate particular strength in capturing local dengue dynamics during epidemic periods and maintain effectiveness even when data reporting is delayed or incomplete. This work provides an advancement over previous forecasting approaches, potentially helping public health officials make more informed decisions.

Keywords: dengue, forecasting, machine learning, virus, outbreak

Abstract

Dengue fever, a tropical vector-borne disease, is a leading cause of hospitalization and death in many parts of the world, especially in Asia and Latin America. Where timely dengue surveillance exists, decision-makers can better implement public health measures and allocate resources. Reliable near-term forecasts may help anticipate healthcare demands and promote preparedness. We propose ensemble modeling approaches combining mechanistic, statistical, and machine learning models to forecast dengue cases 1 to 3 mo ahead at the province level across multiple countries. We assess these models’ predictive ability out-of-sample and retrospectively in over 180 locations worldwide, including provinces in Brazil, Colombia, Malaysia, Mexico, Thailand, plus Iquitos, Peru, and San Juan, Puerto Rico, during at least 2 to 3 y. We also evaluate ensemble approaches in a real-time, prospective dengue forecasting platform during 2022–2023, considering data availability limitations. Our ensemble modeling leads to an improvement to previous efforts that may help decision-making in the context of large uncertainties. This contrasts with the variable performance of individual component models across locations and time. No single model achieves optimal predictions across all scenarios, but while ensemble models may not always perform best in specific locations, they consistently rank among the top 3 performing models both retrospectively and prospectively.


Dengue fever is a tropical vector-borne disease threatening an estimated 3.9 billion people (1) over 141 countries (2), with cases doubling every ten years since 1990 (3). Over 390 million infections arise each year worldwide, with severe dengue causing 25,000 deaths annually, mostly in children (2, 4). For many parts of Asia and Latin America, dengue is a leading cause of hospitalization and death, especially among children (5). It also causes more morbidity and mortality than any other arthropod-borne virus (6). While dengue infections often presents with flu-like or otherwise mild symptoms (7), about 1 in 20 people infected with dengue will develop severe dengue, which can further develop into dengue hemorrhagic fever (4)—a very serious condition marked by capillary leakage leading to potentially significant organ impairment, multiorgan failure, and death (4, 79). The ability to forecast dengue fever cases can provide public health officials with a more accurate picture of future disease dynamics, empowering decision-makers to implement public health measures and better allocate limited resources.

Over the past few decades, multiple approaches have been developed for dengue forecasting. Dynamic, mathematical models that incorporate knowledge of dengue virus transmission biology, historical incidence, and climatological factors have been developed to predict the evolution of dengue epidemics (1013). However, the intricate nature of dengue dynamics poses a significant challenge, as mechanistic assumptions usually remain unclear—or hard to quantify—, and acquiring the data to parameterize models often proves impossible. More recently, data-driven methodologies to predict the severity of an upcoming seasonal outbreak—a classification problem—have experienced a surge in popularity due to the increasing availability of epidemiological and exogenous dengue-related data. These include methods such as k-Nearest-Neighbors, Logistic Regression, and various boosting methods (14); or the identification of weather (temperature, rain frequency) patterns that may help anticipate years with high incidence (1517).

Additionally, classical time series methods like SARIMA (and its variants) have been widely explored to forecast confirmed case counts over time(1821), as well as more complex, nonlinear methods such as generalized additive models, artificial neural networks, and exponential smoothing approaches (18, 20, 22, 23). Other studies have explored the feasibility of leveraging Internet-based data sources such as Dengue-related Google search data (1, 24, 25), social media data (26, 27), and Wikipedia access logs (27) as additional predictors of Dengue activity.

The abundance of dengue forecasting methodologies presents a significant challenge for decision-makers and stakeholders who ultimately need a single set of reliable predictions to make informed decisions to protect their communities. Factors such as data availability, computational resources, and the desired level of accuracy further complicate the decision-making process. Therefore, navigating the array of forecasting methodologies requires careful consideration and expertise to ensure effective dengue prediction and response. Additionally, each model has its own limitations (19, 28), including robustness to variability in data quality and sensitivity to different outbreak phases, which can lead to inconsistent predictions.

Ensemble methods that intelligently and adaptively combine the predictions of multiple component models into a single prediction may be more robust alternatives for disease forecasting. For example, previous studies have used averaged and weighted-averaged ensembles combining models such as SARIMA, vector autoregression, neural networks, and linear regression to nowcast and forecast dengue in Brazil and India (29, 30). Similar ensembling approaches combining models such as the Method of Analogues, Holt-Winter models, and Bayesian generalized linear mixed models, among other historical models, have shown promising results in Iquitos, Peru (31), and Vietnam (32). Chakraborty et al. directly combined ARIMA and neural networks to forecast dengue in San Juan, Iquitos, and the Philippines as a composition of linear and nonlinear signals (33). Mahajan et al. use a gradient-boosting superensemble to combine ARIMA, exponential smoothing, and neural network for forecasting dengue in Hong Kong (34). Ensemble models intentionally involving strong and weak learners to reduce overfitting have also shown good predictive performance in Bangkok and Chiang Mai, Thailand (35).

However, the studies referenced above are 1) individually limited in their geographic scope—to study a few locations within a country at a time, 2) they only focus on a few model choices as potential ensemble components, and 3) they were in all cases implemented retrospectively. The latter implies that issues of data availability, data completeness, and other challenges that emerge in real-time forecasting efforts were not at all considered (36). As such, the approaches described are not guaranteed to be generalized, ready-to-use methodologies globally, in real-time and prospective forecasting efforts.

The challenges of the real-time implementation of disease forecasting systems have been documented extensively, and they are still the backbone that motivate multiple research studies (37, 38, 39). In fact, public health systems typically experience severe lags in reporting (40, 41), and reported case counts for a given month may be updated many times in the following months (this is commonly referred to as “backfill”) (25). While some methods for addressing these data quality and backfill issues have been proposed (42), for example—methods that attempt to learn backfill patterns (37, 43) and methods that introduce auxiliary data sources (44) to improve forecasts; we did not include in our real-time prediction pipeline a comprehensive set of approaches to address these issues. Instead, we evaluated the ability of our machine learning-based ensemble methods to lead to improved or consistent forecasts in the presence of these challenges.

The primary prediction task addressed by this manuscript is the short-term forecast of reported dengue fever case counts one to three months ahead in province-sized localities. The main contributions of this work are twofold. First, we retrospectively formulate a family of ensemble system pipelines—composed of multiple structurally heterogeneous individual component models—that generate more robust and accurate forecasts compared to their individual component models. Specifically, we include the following 11 diverse classes of individual component models as potential inputs to our ensembling pipelines: autoregressive (AR); autoregressive with Google Trends data as exogenous covariates [ARGO (45)]; three variants of vector autoregression [NetModel and two variants of VAR (46)]—VAR (Reg.) and VAR (Clust., Reg.); a combination of ARGO and NetModel [ARGONet (38)]; a mechanistic, repurposed dynamically trained SIR; a classical error-trend-seasonality model (ETS); a miniensemble of machine learning methods (Stacked ML); a baseline naive persistence model; and a baseline seasonal model. We exhaustively demonstrate across 180+ province-sized locations that our ensemble models not only incur lower percent absolute errors on average compared to their individual component models (though such improvements may be modest), but also, and more importantly, they produce top performing forecasts (typically in the top three) more consistently than any individual model.

The second contribution of this study is the evaluation of a real-time dengue activity forecasting platform implemented as a prospective tool to identify when and where an upcoming dengue outbreak would be experienced 1, 2, and 3 mo into the future. These predictive efforts were implemented as a decision-making support tool to guide the allocation of resources for prospective clinical trials in all provinces of Brazil, Colombia, Malaysia, Mexico, and Thailand. The ensemble techniques were assessed using a different set of component models that were chosen based on their suitability to be implemented in the presence of multiple data availability and data incompleteness issues. These models included KNN, VAR (Reg.), Support Vector Machines (SVM), and SARIMA. The scope of our study and the consistency of our analyzes suggests that our ensemble approaches are generalizable across geographically diverse locations and individual component models, and thus may become a reliable first choice for forecasting teams who are interested in communicating concisely with public health officials and other decision makers (39).

Results

We evaluated the performances of eleven optimized component models and two selected ensemble models on forecasting reported dengue cases in over 180 province-sized locations worldwide in Brazil, Colombia, Malaysia, Mexico, and Thailand. Specifically, we retrospectively assessed each model’s ability to forecast reported dengue cases 1-, 2-, and 3-mo ahead into the future, with performance quantified using percent absolute error (SI Appendix).

For each location, we partitioned the available data into distinct training and test periods. During the training phase, we engaged in a hyperparameter optimization for each component model, and two versions of optimal hyperparameters for our ensembles: a “Country” set of hyperparameters (using the same set of hyperparameter with best on-average performance across all locations within a country), and a “Overall” (using the set of hyperparameters that, on average, best performed within all the locations in the study). Following this stage, we proceeded to generate out-of-sample forecasts utilizing the designated test period. It is important to note that due to the varied availability of epidemiological data across different regions, the time periods for analysis varied from country to country. For specific details on the analysis period for each country, please consult SI Appendix, Table S1.

We present our main results in the following way: First, we observe the heterogeneity in the performance of each of our component models in Fig. 1, which presents a summary of the performance of each model across each country, and each horizon of prediction. Then, in Fig. 2, we focus specifically on the capacity of the ensemble models to consistently succeed in generating reliable forecasts, independent of the location where they are trained on. Finally, we present an overview of the error reduction of each component model, emphasizing the consistency of the ensemble techniques to reduce error across locations.

Fig. 1.

Fig. 1.

Country-specific optimized individual and ensemble model performance rankings. (A) Graphical representations of the 1-, 2-, and 3-mo forecast horizons that we explore in this paper. The red X marks our forecasting target n-months ahead, the blue dots represent the historical cases that we are using as our observed training data (in this case, a 5-mo sliding window), the vertical blue dotted lines represent the limits of our training data range, and the gray silhouette represents the ground truth reported case counts. (BF) Within each country and forecast horizon, the heatmaps show the rankings distribution for each individual and ensemble model’s forecasts in terms of percent absolute error. The geographic maps next to each heatmap indicate the best-performing model in each province, color-coded by the legend at the Bottom of the figure.

Fig. 2.

Fig. 2.

A summary of our prediction tasks and optimized model performances across 187 locations. (A) An example of our fine-tuned country-specific ensemble variants’ forecasts compared to their optimized component models in one selected location—Bahia, Brazil. The gold standard ground truth of reported dengue cases is shown as the gray silhouette. Ensemble predictions are shown in thick, bolded lines, while component models are shown in thinner, colored lines. (B) Heatmaps of the number of locations where each model attained a specific ranking in terms of mean absolute error with respect to the ground truth reported dengue case counts across all 187 locations. (C) Geographical maps of Brazil, Colombia, Malaysia, Mexico, and Thailand showing provinces where either the country-specific or overall ensemble performed in the top 3 rankings for each location (in yellow) and where they did not (in gray).

Component model performance was significantly dependent on the location. Our results, summarized in Fig. 1, present a table with the percent absolute error ranking (PAE) distribution for each model within Sections B through F (one for each country). Each row within the table represents a distinct model, while columns denote their respective rankings (first, second, third, etc.). For instance, for Brazil, the first entry of the table indicates that the model “Ensemble (Country, EW)” secured the lowest PAE, achieving first place in 7 out of 27 locations across the country. The tables showcase the top-performing model in terms of this ranking, listed in descending order. We found that our ensembles, including the “Country” and “Optimized” versions of the “Equal Weights” (EW) ensemble, and “Country” and “Overall” “Performance Based Weights” (PBW) tended to be in the top positions within our ranking tables across all horizons for Brazil, Colombia, Mexico, and Thailand (see Methods for details on the ensemble approaches). In the case of Malaysia, the EW ensemble appeared in the sixth position (almost half of the participating models). The Vector Autoregression (VAR) also consistently appeared in top positions for each of the ranking matrices for each country. On the other hand, the positioning of each of our component models varied from country to country. A country-by-country description, along with additional results for San Juan, Puerto Rico, and Iquitos, Peru can be found in SI Appendix. We relegate Puerto Rico and Peru to SI Appendix, since we only have one location in each of these regions, which does not facilitate interprovince analyses.

Ensemble Models Consistently Achieve the Most Top Performance Rankings Compared to Any Individual Model.

Our second main result is that the country-specific and overall ensembles consistently perform among the top 3 relative to the individual component models. Moreover, averaged across all locations, two ensemble variants incurred the lowest prediction error compared to all other component models. Fig. 2 shows the forecasting performance of both individual and ensemble models across all 187 tested locations. In panel (A), we provide one example emphasizing the advantages of using ensembles over component models. For the 1- and 2-mo ahead horizons, the best country-specific ensemble used EW for all component models when producing an ensemble forecast at each timestep, while at the 3-mo ahead horizon, the best country-specific ensemble used PBW, where the weightings of the component models were determined based on which component models performed the best in the recent past. Specifically, at the 1- and 2-mo ahead forecast horizons, the ensemble model predicted the ground truth (gray silhouette) much more closely and with less variance than that of its component models, which tended to display more significant oscillation relative to the ground truth. At the 3-mo ahead horizon, the ensemble and its component models tended to produce less accurate predictions in this specific example. Given the broad geographical scope of our study, these findings suggest that ensemble models are a suitable default choice for generalizable forecasts.

The heatmaps in Fig. 2B show that the country-specific and overall ensembles had the most locations where they performed in the top 3 rankings in terms of percent absolute error, followed by regularized VAR and clustered + regularized VAR, out of the 187 tested locations. At the 2-mo ahead horizon, country-specific ensembles still garnered the most top 3 rankings, but regularized VAR seemed to garner more top 3 rankings than the overall ensemble. At the 3-mo ahead horizon, the country-specific and overall ensembles again attained the greatest number of top 3 rankings, followed by regularized VAR. While ensembles may not always be the absolute winner in every tested location, overall, they consistently placed on the top 3 rankings. From the geographical maps in panel (C), ensemble models performed in the top 3 rankings in a vast majority of all 187 locations tested, as indicated by most of the maps being colored yellow. We found only a few consistent exceptions to this rule primarily in central Brazil and Colombia. These maps again corroborate the finding that ensembles are an effective option for stakeholders.

Ensembling Also Yields Improvement in Error Distributions across Locations.

We analyzed the potential error reduction of our ensemble models with respect to individual models. Our results are displayed in Fig. 3, which shows the percent absolute error distributions for the individual and the ensemble models in the 187 locations tested. SI Appendix, Fig. S1 shows the same percent absolute error distributions of the individual and ensemble models within each country and forecast horizon.

Fig. 3.

Fig. 3.

Overall error distributions for optimized individual and ensemble models across all 187 tested locations. (AC) Ridgeline plots show percent absolute error distributions at the 1-, 2-, and 3-mo horizons, respectively. Side tables record the mean percent absolute error incurred.

Overall performances.

Despite the minor reduction in some cases, the country-specific and overall ensembles had the lowest mean percent absolute errors compared to all other models across all forecast horizons. Fig. 3AC displays the error distributions for each model, across the 3 different horizons of forecast. In all cases, we can observe that the ensemble models were placed as the top performers, reaching values of 38.5%, 54.5%, and 62.7% in terms of percent absolute error (% AE). The next best models were the VAR variants (one incorporating regularization, and other implementing both regularization and clustering), reaching values of 40%, 55%, and 64%, for each respective task.

Performance by country.

SI Appendix, Fig. S1 shows the performance per country. Notably, an ensemble variant achieved the lowest mean percent absolute error in 14 out of 15 tested combinations of forecast horizon and country. The sole exception was Colombia at 2-mo ahead, with clustered + regularized VAR having achieved the lowest mean percent absolute error. However, the mean percent absolute errors and the shapes of the errors’ distributions are very similar for the top-performing models. Nonetheless, our analyses shown in this figure still confirm that the ensembles were greater than the sum of their component models, even when looking only within a particular country.

High-performing locations.

While Fig. 3 present overall forecast error distributions across all provinces, certain locations exhibit consistently strong predictive performance. In Fig. 4, we highlight 20 such locations, each represented by two panels: i) a time-series plot of reported dengue cases (gray lines) overlaid with ensemble forecasts (colored lines) at 1-, 2-, and 3-mo-ahead horizons; and ii) a corresponding scatter plot comparing predicted versus observed values with corresponding R2. For example, locations such as Nonthaburi (Thailand), Goias (Brazil), and Meta (Colombia) show notably high correlation between predictions and ground truth, particularly at the shorter lead times (1 to 2 mo ahead). These examples illustrate that, despite relatively high average percent absolute errors in some provinces, the ensemble can capture local dengue dynamics quite accurately under favorable epidemiological and data conditions.

Fig. 4.

Fig. 4.

Ensemble forecast in the top 20 high-performing locations. Each row presents data for a specific province, including a time-series panel (Top) and a scatter plot panel (Bottom). In the time-series panels, the gray line indicates the reported dengue case counts, while the colored lines represent ensemble forecasts at 1-, 2-, and 3-mo-ahead horizons. In the scatter plots, the x-axis is the observed dengue incidence and the y-axis is the predicted incidence for each horizon, with the solid diagonal line denoting perfect agreement. The corresponding R2 values for each prediction are also provided. Locations such as Nonthaburi (Thailand), Goias (Brazil), and Meta (Colombia) demonstrate notably strong predictive accuracy, particularly at shorter lead times.

Prospective Study: A Real-Life Application and Analysis of Our Methodology.

In this section, we evaluate the performance of our forecasts generated by a real-time dengue activity forecasting platform. This platform was designed as a prospective tool to predict where dengue outbreak activity would occur 1, 2, and 3 mo into the future. There are two primary differences between our retrospective and prospective studies. First, the prospective study was conducted in a real-world scenario where future ground truth data was unknown, whereas in the retrospective study, we had access to the entire time series beforehand. Second, due to reporting delays in dengue case data, our prospective predictions did not utilize the most up-to-date epidemiological data. In contrast, the retrospective predictions were made using the most complete datasets available. This evaluation covered various locations in Brazil, Colombia, Mexico, Thailand, and Panama from May 2022 to July 2023. The forecasting models included SVM, K-Nearest Neighbors (KNN), a regularized version of VAR, and Seasonal Autoregressive Integrated Moving Average (SARIMA). Our ensemble models comprised a weighted ensemble, which assigned weights based on the mean squared error score of each model over the past six months, and a winner-takes-all ensemble, which selected the prediction from the “best” base model with the least mean squared error over the past 3 mo. Additionally, we used Persistence as our baseline model.

Fig. 5 and Table 1 show a summary of the number of times our ensemble’s mean squared error was among the top 3 best over all the locations within a Country (the analysis was repeated over each forecasting horizon). We also present the overall error reduction of each model with respect to persistence (ERRORmodelERRORpersistence) using a set of violin plots. We conducted the analysis for each location and each country (Fig. 6).

Fig. 5.

Fig. 5.

Top performer count. A visualization of the number of times a model scored a value of MSE rated among the top 3 best, ordered from Left (models with the highest count) to Right (models with the lowest count). Our results show that the weighted ensemble (blue), VAR (green), and the winner-takes-all approach are frequently among the top performers (leftmost side).

Table 1.

Summary of the number of times a model scored a Mean Squared Value among the top 3 best, across every location, and every horizon

Top 3 overview
Weighted Ensemble 3,318
VAR (regularized) 2,989
Winner Takes All 2,310
SARIMA 1,638
KNN 1,519
SVM 1,512
Persistence 658

Our results show that our weighted ensemble, VAR and Winner takes all approach are among the top performers.

Fig. 6.

Fig. 6.

Error reduction with respect to the Persistence model. Summary of the error reduction of each model with respect to the Persistence baseline model. Each plot represents a different horizon (columns) and country (rows). Each violin plot visualizes a summary of the error reduction scores (ERRORmodelERRORpersistence) for a single model. Models are ordered from worst (Left) to best (Right). The gray dashed line represents the value of persistence and serves as a reference to validate whether a model improved over the baseline.

Top 3 analysis.

Fig. 5 shows the number of times a model reached within the top 3 mean squared error reductions, for each country and each horizon. Each barplot represents the performance of a model within a Country, for a different prediction horizon. We can see VAR appearing eight times on the leftmost side (40%), and our weighted ensemble appearing seven times (35%). Our Winner-Takes-all approach appeared only two times on the leftmost side (10%) but had a comparable count with the top model whenever it was on second position (see Thailand in horizon 1, Panama in Horizon 1 and 4, Brazil in horizon 1, and Mexico in Horizon 1 and 2). Overall, most of our base models and ensembles had a higher count than persistence, with exception to Colombia in horizon 1, Panama in horizon 2, and Mexico in Horizon 1. Table 1 shows a total count over all locations, and all horizons. Our weighted ensemble had a total of 3,318 counts, followed by VAR with 2,989 and Winner-Takes-All with 2,310 counts.

Overall error reduction.

Fig. 6 exhibits a summary of the error reduction for the analyzed models with respect to the Persistence model (ERRORmodelERRORpersistence). The violin plots are ordered so that the best model is to the rightmost side, and a dashed horizontal line plotted at y=1 (y=1 means the error of our model is equal to the error of Persistence) serves as a reference to know whether a model consistently beats persistence or not.

The weighted ensemble (WE), vector autoregression, and the winner-takes-all (WTA) ensemble were the three models that most frequently scored within top-3 error reduction. We observe that the weighted ensemble had median error reduction within the top 3 at every location and time horizon, except Panama in horizon 1. Regularized VAR scored the biggest error reduction in Colombia and Thailand for horizons 1, 2, and 3, 4. Although less frequently, the Winner-Takes-All ensemble also remained within the top 3 performances with exception to Thailand in horizon 4, Panama in horizon 3 and 4, and Mexico in horizon 3 and 4.

Impact of reporting delays in our model’s performance.

In conducting the prospective analysis, our forecasts were generated in a real-time scenario where the ground truth for each location was not fully reported at the time of prediction. The performance of our models were therefore likely affected by backfill issues in the data. Fig. 7 illustrates our forecasts within the region of Sergipe, Brazil. At the time of prediction, the available information on confirmed cases (depicted in dark gray) differed from the most recently reported data (shown in light gray), which we employed as our ground truth for final metrics and error scores. Such backfill issues significantly impact real world applications as our models are trained solely on the information available at the given point in time.

Fig. 7.

Fig. 7.

Visualization of the bias embedded in our models given reporting delays. (A) Bias in Sergipe, Brazil. (B) Bias in Amazonas, Colombia. (C) Bias in Baja California Sur, Mexico. (D) Amnat Charoen, Thailand. The official reports known at the time of prediction, shown in dark gray, is the only information available to our models at the time of prediction. After several months, the ground truth changes due to backfill efforts based on the most recent reports are shown in light gray for each location.

High-performing locations in our prospective study.

In SI Appendix, Fig. S2, we highlight 20 high-performing locations in our prospective study, each represented by two panels: i) a time-series plot of reported dengue cases (gray lines) overlaid with ensemble forecasts (colored lines) at 1-, 2-, 3-, and 4-mo-ahead horizons; and ii) a corresponding scatter plot comparing predicted versus observed values. Locations such as Bahia (Brazil), Kalasin (Thailand), and Khon Kaen (Thailand) show notably high correlation between predictions and ground truth, particularly at the shorter lead times (1 to 2 mo ahead). These examples show that, even though some provinces exhibit relatively high average percent absolute errors, our ensemble can still accurately capture local dengue dynamics when epidemiological and data conditions are favorable.

These “high-performing” provinces were selected post hoc to illustrate the upper bound of ensemble skill. Preliminary exploratory checks suggest that provinces with uninterrupted multiyear surveillance and minimal postrelease data revision are far more likely to appear in the top tier, whereas geographical area shows no consistent relationship to forecast accuracy. Because those diagnostics can be confirmed only after inspecting the incoming data stream, we cannot yet identify such provinces a priori; developing lightweight, real-time data-quality indicators remains an important direction for future work.

Discussion

Dengue is a leading cause of hospitalization and death for many people around the world (5), and with cases doubling every ten years (3). An essential component of dengue control is disease forecasting. Enhancing the accuracy and robustness of predictive models, particularly across multiple diverse geographical localities, empowers public health institutions to adopt a more proactive approach toward curbing the spread of the disease.

We tackled the task of predicting reported dengue fever case counts one, two, and three months ahead in various province-sized locations around the worldwide. Since individual model performances typically fluctuate across locations, there is a need for more robust and generalizable forecast models. In this context, our first contribution was the development of a family of ensemble models that retrospectively produced more accurate forecasts than their components across a broad range of geographically and socially diverse locations. Specifically, we investigated eleven types of data-driven, statistical, and mechanistic models as potential components of our ensemble. We also explored three ensembling mechanisms—performance-based weights, winner-takes-all, and simple average—and found that our ensemble models achieved lower percent absolute errors across 180+ geographically diverse locations. Our second contribution was a real-time dengue forecasting platform to predict when and where outbreaks will occur 1, 2, and 3 mo in advance. Our forecasts were made in real-time without complete ground truth for each location. In fact, this task is not the same as performing retrospective studies, since dengue case data are typically updated months later (a problem commonly referred to as “backfill”). In this challenging scenario, our ensemble models still emerged as the top performers, producing better forecasts and reducing error compared to individual components.

Our ensemble models were good predictors of dengue across the world. This fact is especially relevant because, as shown in Fig. 1, no individual model consistently achieved the lowest error. By contrast, while country-specific and overall ensembles were not always the best-performing models at a given location and forecast horizon, they almost always incurred the most top 3 performers compared to any of the component models. As shown in Fig. 3, our ensemble variants incurred the lowest error averaged across all 180+ tested locations in terms of percent absolute error compared to the component models. Even looking within a particular country and forecast horizon, as shown in SI Appendix, Fig. S1, ensemble models achieved the lowest error averaged across all locations within that country in 13 out of 15 combinations of country and forecast horizon. The two exceptions were Colombia at 2- and 3-mo ahead. Furthermore, in 9 out of 15 country-horizon combinations, both the country-specific and overall ensembles achieved the top 2 lowest errors averaged across locations. While ensemble models improve performance relative to individual models, research efforts should continue to explore ways to improve the accuracy of individual models’ long-term forecasts (2 to 3 mo). Inconsistent data availability, disease counts’ reporting delays and underreporting, among other factors, complicate this endeavor especially in real-time forecasting efforts. In our prospective study, the WE model was overall the top performer in terms of low mean squared error, followed by VAR and the WTA ensemble (Table 1). In terms of error reduction with respect to the persistence model, WE and WTA also consistently performed within top-3 performers across locations.

It is worth addressing the excellent performance of our VAR model in both retrospective and prospective studies. Examining our results more granularly, from SI Appendix, Fig. S1 we observe that in Colombia, at the 2-mo ahead horizon, VAR (Clust., Reg.) achieved more top 3 rankings than both ensemble variants. At the 3-mo ahead horizon, VAR (Clust., Reg.) not only outperforms the overall ensemble in terms of the number of top 3 rankings but achieves more top 1 rankings than both ensemble variants. Similarly, at least one VAR model also outperforms at least one ensemble variant in terms of the number of top 3 rankings in all three forecast horizons of both Malaysia and Thailand. We hypothesize that VAR’s stellar performance in Colombia, Malaysia, and Thailand can be significantly attributed to the fact that these three countries are “province-dense” in the sense that individual provinces are relatively geographically small and, by extension, extremely close to each other. For example, Thailand has 77 provinces compacted into a relatively small total surface area. In contrast, Brazil has 27 provinces spread out across a much larger area. The consequence of this geographical difference is that population centers between Brazilian provinces are much farther apart, and thus network effects are much weaker than their Thai counterparts. As such, VAR is much more effective in Thailand than Brazil because there are significantly stronger network effects between provinces to capture in our models.

We also investigated component models that were not exhaustively hyperparameter tuned but rather deployed straight out of the box, which we refer to as “standard” models. For details, we refer the reader to our SI Appendix. As shown in SI Appendix, Figs. S7 and S9, while our standard component models are nearly all unable to outperform our naive persistence baseline in terms of percent absolute error as averaged across locations, our ensemble models composed of these standard component models outperformed the naive persistence baseline consistently. As such, our ensembling approach can take relatively weak, unoptimized learners and output a much stronger and more robust prediction. In this sense, the ensemble still performs better than its components.

From SI Appendix, Tables S5–S7, we observe that when forecasting 1- and 2-mo ahead, the ensembling method most commonly employed (albeit plurality, not majority) was the equal weights method, followed by the performance-based weights model. At 3-mo ahead, however, the performance-based weights mechanism was employed in most countries, including the overall ensemble. There does not appear to be a clear trend with respect to the ensemble training window sizes used to fit the performance-based weights and winner-takes-all ensembles.

Since the success of our forecasts is measured by achieving a lower percent absolute error than the naive persistence baseline model, ensembling enables us to include the naive persistence model itself as a component. As shown in our standard model results in SI Appendix, we observe that when working with standard, non-fine-tuned component models, nearly all of the best country-specific and overall ensembles were composed of the naive and seasonal basic models, coupled with one or two other models. From a bias–variance tradeoff perspective, the naive persistence model has very low variance, given its absence of tunable parameters. While other component models may overfit to noise and thus incur large errors, the naive persistence, by being simple, provides a stable component to the ensemble and thus allows the ensemble to outperform the other models, including its components.

In this work, we selected PAE as our primary evaluation metric due to its interpretability for public health officials. However, deriving PAE from mean absolute error (MAE) in other studies presents challenges, given differences in evaluation periods, limited access to raw data, and variability in reported mean case counts at the time of prediction. As such, a comprehensive head-to-head comparison between our models and previous studies is beyond the scope of this work but remains an important future research direction. To provide some context for our ensemble model performance, we conducted a targeted comparison against the best-performing models from Johansson et al. (47) for a subset of Mexican states. The results, shown in SI Appendix, Tables S56–S58, present PAE values from Johansson’s models alongside our ensembles, with best-performing models highlighted. While acknowledging that these comparisons involve different prediction periods, our ensemble models exhibit comparable performance for 1-mo-ahead forecasts. Notably, for 2- and 3-mo-ahead forecasts, our ensembles consistently achieve lower errors, demonstrating substantial improvements in predictive accuracy. It is also important to highlight that Johansson’s models were specifically optimized for Mexico, whereas our ensemble approach was designed for generalization across multiple countries worldwide.

To place our results in the context of broader infectious disease forecasting efforts, we compared the forecast accuracy of our dengue ensembles to that of the US CDC’s FluSight ensemble, which provides weekly influenza forecasts at 1- to 4-wk lead times. Despite the differences in disease dynamics and temporal resolution, we found that the average prediction error of our dengue forecasts was slightly lower than those reported by FluSight. This comparison, presented in SI Appendix, Fig. S12, highlights that a PAE of approximately 40 to 60% for 1- to 3-mo dengue forecasts is consistent with the performance standards observed in operational forecasting systems. These findings emphasize that, even in the face of dengue’s inherent variability and longer forecast horizons, our ensemble approach achieves a level of accuracy comparable to real-world benchmarks.

One limitation of our work is that of the eight nonbasic models that we include as potential components into the ensemble, six of them—AR, ARGO, ARGONet, NetModel, VAR (regularized), VAR (clustered + regularized)—can be interpreted as belonging to a common family tree of linear models involving autoregressive terms. In the future, one could consider including more expressive but also more heavily parameterized models such as Random Forests (48) and neural networks (48, 49) into our ensemble lineup to potentially increase performance. However, as explored in ref. 49, heavily overparameterized neural networks may underperform compared to simpler regression models at short-term disease forecasting tasks. With the exception of ARGO and ARGONet, all of our models were trained exclusively using historical dengue-reported case counts. Future work could include models that leverage climate data and earth observations into our ensemble lineup, as explored in ref. 32.

In terms of further limitations and outlook, from a public health perspective, it would be ideal if forecasters were able to predict the entire expected burden of Dengue for a season in a given location. However, this is currently far from realization at the global scale we address in this work. Thus, our study was designed to benchmark operational 1- to 3-mo dengue forecasts across the widest set of provinces yet assembled. Because prospective surveillance is currently available for only a single transmission season (2022–2023), our evaluation cannot capture lower-frequency drivers such as climate anomalies, serotype turnover, or secular trends in mosquito control. Likewise, the incidence-only models retrained on short sliding windows do not include explicit climate or immunological covariates; abrupt trend reversals or localized climate shocks that fall outside the recent training window may therefore elude prediction. Extending the prospective record over multiple years and incorporating lightweight climate or serotype indicators are critical next steps if longer seasonal horizons are desired. Complementary, region-specific case studies that leverage longer historical records could also probe two questions we could not tackle at global scale: i) what training-window length is optimal under different transmission regimes, and ii) how much additional skill comes from climate or serotype covariates. Such focused analyses-including mechanistic simulation experiments under well-known climate modes such as the El Niño-Southern Oscillation-lie beyond the page scope of the present work, but our open-source forecasting pipeline is expressly designed to support them in future collaborations.

We acknowledge that ensemble skill decays rapidly beyond the one-month mark (Fig. 4B), and by three months the median absolute error roughly doubles relative to 1-mo forecasts. While longer lead times would surely aid seasonal-scale resource planning, most dengue-control interventions (e.g., insecticide procurement, community outreach, larval-source reduction) can be mobilized within 4 to 12 wk. Accordingly, we focused our evaluation on 1- to 3-mo horizons, which align with operational time frames adopted in existing real-time forecasting hubs such as CDC FluSight. Extending useful lead time will likely require hybrid incidence-plus-covariate models or region-specific mechanistic approaches, which are both promising avenues for future work.

Future studies could explore classification tasks of predicting whether a given location will experience an outbreak by thresholding our case count predictions. Methods like DT-SIR, while prone to overpredicting at outbreak peaks, are still excellent at capturing the outbreak progression trend. Another interesting research avenue involves combining regression and classification components together within ensembles. For example, one can consider an ensemble setup containing both regression models (predicting the number of dengue reported case counts) and classification models (predicting whether an outbreak will occur in the next months). Future work could also involve combining ensembles together into superensembles to further reduce variance.

Methods

Data Sources.

We trained our models using two primary data modalities: reported case counts and internet search trends. Weekly and monthly dengue case data were collected from official surveillance systems across Brazil, Thailand, Colombia, Malaysia, Mexico, Peru, and Puerto Rico. Sources included SINAN, InfoDengue, and respective Ministries of Health. Data were reformatted, time-stamped, and aggregated to province and monthly levels for consistency. Detailed procedures, including optical character recognition extraction and verification steps, are provided in SI Appendix. We also accessed the Google Health Trends API to obtain monthly dengue-related search frequencies at the provincial level. When provincial-level data were unavailable, national-level trends served as proxies. In each country, we selected the 10 terms most correlated with local dengue activity based on preevaluation Pearson correlations to prevent signal leakage. These were used as covariates in models incorporating search data. More information is presented in SI Appendix.

Individual Models and Fitting Strategies.

We constructed a diverse set of forecasting models. Autoregression (AR) models use recent case counts in both standard (AR(4)) and optimized (L = 24 with LASSO regularization) variants. ARGO models augment AR with Google Trends features, with configurations mirroring AR. NetModel adds information from neighboring provinces’ case counts to the AR framework, with log transformations and manual lag selections in the optimized version. ARGONet ensembles ARGO and NetModel by averaging their predictions. ETS models were implemented using sktime, selecting preprocessing steps via cross-validation. VAR models, regularized and optionally clustered by geography, were fit per country. We also developed a stacked machine learning model with Elastic Net and SVM base learners. Finally, DT-SIR is a dynamically trained SIR model that uses recent case counts to estimate transmission dynamics, with a thresholding mechanism to mitigate overshooting near outbreak peaks. Baseline models included naive persistence and seasonal averages, providing reference points for model evaluation. We adopted sliding and expanding model-fitting strategies illustrated in Fig. 8. Models leveraging SIR dynamics used the sliding window; other models used expanding windows. Full implementation details, lag selections, hyperparameter choices, and preprocessing strategies are detailed in SI Appendix.

Fig. 8.

Fig. 8.

Schematics of model-fitting techniques. (A) Sliding window model-fitting. (B) Expanding window model-fitting.

Ensemble Systems.

We evaluated three ensembling schemes (SI Appendix, Fig. S4). The EW approach computes a uniform average of all model predictions. The PBW method determines nonnegative weights optimized via least-squares based on recent predictive performance. Last, the WTA strategy selects the single best-performing individual model at each time point based on recent accuracy. All ensembles were recalibrated at each forecast month. Specific ensemble configurations by country and forecast horizon are reported in SI Appendix, Tables S2–S7.

Model Evaluation.

We reported out-of-sample accuracy for each forecast horizon (1-, 2-, and 3-mo ahead) and compared ensemble and individual model performance. Model evaluation focused on the most recent multiyear periods available in each country (SI Appendix, Table S2). To place the ensemble’s accuracy in context, all calculated model errors are errors relative to the performance of the naive persistence model (Fig. 6).

Supplementary Material

Appendix 01 (PDF)

Acknowledgments

Author contributions

S.W., A.G.M., L.C., L.M.S., F.L., R.V., S.M., and M.S. designed research; S.W., A.G.M., L.C., L.M.S., F.L., A.M., R.V., and S.M. performed research; A.G.M., L.C., L.M.S., F.L., A.M., R.V., S.M., and M.S. contributed new reagents/analytic tools; S.W., A.G.M., L.C., L.M.S., F.L., A.M., R.V., and S.M. analyzed data; and S.W., A.G.M., L.C., L.M.S., F.L., and M.S. wrote the paper.

Competing interests

A.M., R.V., and S.M. work for Johnson and Johnson and own equity in Johnson and Johnson. M.S. has received institutional research funds from the Johnson and Johnson foundation and from Janssen global public health.

Footnotes

This article is a PNAS Direct Submission.

PNAS policy is to publish maps as provided by the authors.

Contributor Information

Austin G. Meyer, Email: austin.g.meyer@gmail.com.

Mauricio Santillana, Email: m.santillana@northeastern.edu.

Data, Materials, and Software Availability

Previously published data were used for this work (Data were collected from the websites of the ministries of health in each country. See Data Sources for more details).

Supporting Information

References

  • 1.Yang S., et al. , Advances in using Internet searches to track Dengue. PLoS Comput. Biol. 13, e1005607 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Brady O. J., et al. , Refining the global spatial limits of dengue virus transmission by evidence-based consensus. PLOS Negl. Trop. Dis. 6, e1760 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Kiang M. V., et al. , Incorporating human mobility data improves forecasts of Dengue fever in Thailand. Sci. Rep. 11, 1–12 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Bhatt S., et al. , The global distribution and burden of Dengue. Nature 496, 504–507 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Gluskin R. T., Johansson M. A., Santillana M., Brownstein J. S., Evaluation of Internet-based Dengue query data: Google Dengue trends. PLoS Negl. Trop. Dis. 8, e2713 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Campbell K. M., Lin C., Iamsirithaworn S., Scott T. W., The complex relationship between weather and Dengue virus transmission in Thailand. Am. J. Trop. Med. Hyg. 89, 1066 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.World Health Organization, Dengue and severe dengue. https://www.who.int/news-room/fact-sheets/detail/dengue-and-severe-dengue. Accessed 25 July 2025.
  • 8.Kalayanarooj S., Clinical manifestations and management of Dengue/DHF/DSS. Trop. Med. Health 39, S83–S87 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.C. Kusiak, Real-time Dengue forecasting in Thailand: A comparison of penalized regression approaches using internet search data, Master’s Thesis, University of California, Berkeley, CA (2018).
  • 10.Chen Y., et al. , An ensemble forecast system for tracking dynamics of Dengue outbreaks and its validation in China. PLoS Comput. Biol. 18, e1010218 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Yi C., Cohnstaedt L. W., Scoglio C. M., SEIR-SEI-EnKF: A new model for estimating and forecasting Dengue outbreak dynamics. IEEE Access 9, 156758–156767 (2021). [Google Scholar]
  • 12.Johansson M. A., Hombach J., Cummings D. A., Models of the impact of dengue vaccines: A review of current research and potential approaches. Vaccine 29, 5860–5868 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Reiner R. C. Jr., et al. , A systematic review of mathematical models of mosquito-borne pathogen transmission: 1970–2010. J. R. Soc. Interface. 10, 20120921 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Z. R. D. Omadlao et al. , “Machine learning-based Dengue forecasting system for Irisan, Baguio City, Philippines” in AIP Conference Proceedings (AIP Publishing LLC, 2022), vol. 2472, p. 040019.
  • 15.McGough S. F., Clemente L., Kutz J. N., Santillana M., A dynamic, ensemble learning approach to forecast Dengue fever epidemic years in Brazil using weather and population susceptibility cycles. J. R. Soc. Interface. 18, 20201006 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Stolerman L. M., Maia P. D., Kutz J. N., Forecasting Dengue fever in Brazil: An assessment of climate conditions. PloS One 14, e0220106 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Souza C., Maia P., Stolerman L. M., Rolla V., Velho L., Predicting Dengue outbreaks in Brazil with manifold learning on climate data. Expert Syst. Appl. 192, 116324 (2022). [Google Scholar]
  • 18.Baquero O. S., Santana L. M. R., Chiaravalloti-Neto F., Dengue forecasting in São Paulo city with generalized additive models, artificial neural networks and seasonal autoregressive integrated moving average models. PloS One 13, e0195065 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Johansson M. A., et al. , An open challenge to advance probabilistic forecasting for Dengue epidemics. Proc. Natl. Acad. Sci. U.S.A. 116, 24268–24274 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Othman M., Indawati R., Suleiman A. A., Qomaruddin M. B., Sokkalingam R., Model forecasting development for Dengue fever incidence in Surabaya City using time series analysis. Processes 10, 2454 (2022). [Google Scholar]
  • 21.Thiruchelvam L., Dass S. C., Asirvadam V. S., Daud H., Gill B. S., Determine neighboring region spatial effect on Dengue cases using ensemble ARIMA models. Sci. Rep. 11, 5873 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Attanayake A., Perera S., Liyanage U., Exponential smoothing on forecasting Dengue cases in Colombo. Sri Lanka. J. Sci. 11, 11–22 (2020). [Google Scholar]
  • 23.Li Z., Gurgel H., Xu L., Yang L., Dong J., Improving Dengue forecasts by using geospatial big data analysis in Google Earth engine and the historical Dengue information-aided long short term memory modeling. Biology 11, 169 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Koplewitz G., Lu F., Clemente L., Buckee C., Santillana M., Predicting Dengue incidence leveraging Internet-based data sources. a case study in 20 cities in Brazil. PLoS Negl. Trop. Dis. 16, e0010071 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Rangarajan P., Mody S. K., Marathe M., Forecasting dengue and influenza incidences using a sparse representation of Google trends, electronic health records, and time series data. PLoS Comput. Biol. 15, e1007518 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Guo P., et al. , An ensemble forecast model of Dengue in Guangzhou, China using climate and social media surveillance data. Sci. Total Environ. 647, 752–762 (2019). [DOI] [PubMed] [Google Scholar]
  • 27.Marques-Toledo C. A., et al. , Dengue prediction by the web: Tweets are a useful tool for estimating and forecasting Dengue at country and city level. PLoS Negl. Trop. Dis. 11, e0005729 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Oidtman R. J., et al. , Trade-offs between individual and ensemble forecasts of an emerging infectious disease. Nat. Commun. 12, 5379 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.K. Kempfert et al. , Time series methods and ensemble models to nowcast Dengue at the state level in Brazil. arXiv [Preprint] (2020). https://arxiv.org/abs/2006.02483 (Accessed 1 January 2025).
  • 30.Shashvat K., Basu R., Bhondekar P., Kaur A., An ensemble model for forecasting infectious diseases in India. Trop Biomed 36, 822–832 (2019). [PubMed] [Google Scholar]
  • 31.Buczak A. L., et al. , Ensemble method for Dengue prediction. PLoS One 13, e0189988 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Colón-González F. J., et al. , Probabilistic seasonal dengue forecasting in Vietnam: A modelling study using superensembles. PLOS Med. 18, e1003542 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Chakraborty T., Chattopadhyay S., Ghosh I., Forecasting Dengue epidemics using a hybrid methodology. Phys. A Stat. Mech. Appl. 527, 121266 (2019). [Google Scholar]
  • 34.Mahajan A., et al. , A novel stacking-based deterministic ensemble model for infectious disease prediction. Mathematics 10, 1714 (2022). [Google Scholar]
  • 35.Kerdprasop N., Kerdorasop K., Chuaybamroong P., “A multi-criteria scheme to build model ensemble for Dengue infection case estimation” in 2020 International Conference on Decision Aid Sciences and Application (DASA) (IEEE, 2020), pp. 214–218.
  • 36.Mathis S. M., et al. , Evaluation of flusight influenza forecasting in the 2021–22 and 2022–23 seasons with a new target laboratory-confirmed influenza hospitalizations. Nat. Commun. 15, 6289 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Menkir T. F., et al. , A nowcasting framework for correcting for reporting delays in malaria surveillance. PLoS Comput. Biol. 17, e1009570 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Lu F. S., Hattab M. W., Clemente C. L., Biggerstaff M., Santillana M., Improved state-level influenza nowcasting in the United States leveraging Internet-based data and network approaches. Nat. Commun. 10, 147 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Shea K., et al. , Multiple models for outbreak decision support in the face of uncertainty. Proc. Natl. Acad. Sci. U.S.A. 120, e2207537120 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Fairchild G., et al. , Epidemiological data challenges: Planning for a more robust future through data standards. Front. Public Health 6, 336 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.McGough S. F., Johansson M. A., Lipsitch M., Menzies N. A., Nowcasting by Bayesian Smoothing: A flexible, generalizable model for real-time epidemic tracking. PLoS Comput. Biol. 16, e1007735 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.D. Gamerman, M. O. Prates, T. Paiva, V. D. Mayrink, Building a Platform for Data-Driven Pandemic Prediction: From Data Modelling to Visualisation - the CovidLP Project (CRC Press, 2021).
  • 43.H. Kamarthi, A. Rodríguez, B. A. Prakash, Back2future: Leveraging backfill dynamics for improving real-time predictions in future. arXiv [Preprint] (2021). https://arxiv.org/abs/2106.04420 (Accessed 1 January 2025).
  • 44.Osthus D., Daughton A. R., Priedhorsky R., Even a good influenza forecasting model can benefit from internet-based nowcasts, but those benefits are limited. PLoS Comput. Biol. 15, e1006599 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Yang S., Santillana M., Kou S. C., Accurate estimation of influenza epidemics using Google search data via ARGO. Proc. Natl. Acad. Sci. U.S.A. 112, 14473–14478 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.W. Nicholson, D. Matteson, J. Bien, BigVAR: Tools for modeling sparse high-dimensional multivariate time series. arXiv [Preprint] (2017). http://arxiv.org/abs/1702.07094 (Accessed 1 January 2025).
  • 47.Johansson M. A., Reich N. G., Hota A., Brownstein J. S., Santillana M., Evaluating the performance of infectious disease forecasts: A comparison of climate-driven and seasonal dengue forecasts for Mexico. Sci. Rep. 6, 33707 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Zhao N., et al. , Machine learning and Dengue forecasting: Comparing random forests and artificial neural networks for predicting Dengue burden at national and sub-national scales in Colombia. PLoS Negl. Trop. Dis. 14, e0008056 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Aiken E. L., Nguyen A. T., Viboud C., Santillana M., Toward the use of neural networks for influenza prediction at multiple spatial resolutions. Sci. Adv. 7, eabb1237 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Appendix 01 (PDF)

Data Availability Statement

Previously published data were used for this work (Data were collected from the websites of the ministries of health in each country. See Data Sources for more details).


Articles from Proceedings of the National Academy of Sciences of the United States of America are provided here courtesy of National Academy of Sciences

RESOURCES