Skip to main content
Nature Communications logoLink to Nature Communications
. 2026 May 19;17:6599. doi: 10.1038/s41467-026-72655-7

A real-time early warning system to anticipate respiratory disease outbreaks using transfer learning

Raul Garrido-Garcia 1,2,3, Leonardo Clemente 1,3,4, Austin G Meyer 1,5, George Dewey 1,3, Shihao Yang 1,6, Mauricio Santillana 1,2,3,7,
PMCID: PMC13381708  PMID: 42156378

Abstract

Respiratory disease outbreaks burden American healthcare systems with over one million hospitalizations annually, yet current surveillance systems lag 1–2 weeks behind real-time conditions, preventing timely intervention. We present a machine learning early warning system that combines Google search trends with traditional epidemiological data using ensemble voting algorithms to predict outbreak timing across multiple respiratory pathogens. Unlike prior digital surveillance systems focused on retrospective evaluation or single-pathogen settings, this work presents a unified, prospectively deployed early warning framework that detects both outbreak onsets and peaks across multiple respiratory pathogens at the state level in real time. The system applies anomaly detection and transfer learning to monitor syndromic influenza-like illnesses, and hospitalizations caused by respiratory syncytial virus or influenza, simultaneously, across all 50 states. During operational real-time deployment from August 2024 through the 2024–2025 season, the system detects 98.0% of outbreak onsets and 97.0% of outbreak peaks, with average lead times of approximately 5 and 2 weeks, respectively, and positive predictive values exceeding 82%. This framework transforms reactive public health responses into proactive epidemic preparedness by reducing historical timing uncertainty from 10–20 weeks to consistent 2–6 week prediction windows, providing a scalable approach for monitoring both seasonal outbreaks and emerging respiratory threats.

Subject terms: Epidemiology, Health policy, Population screening


A machine learning early warning system combining Google search trends with epidemiological data detects 98% of respiratory disease outbreak onsets and 97% of peaks across all 50 US states, providing 2-5 weeks of advance warning in real time.

Introduction

Syndromic surveillance of influenza-like illnesses (ILI) in the United States has been conducted for over three decades and serves as a cornerstone of public health preparedness against respiratory disease outbreaks. The World Health Organization defines ILI as an acute respiratory infection characterized by fever ≥38 °C and a cough (or sore throat) with onset within the past 10 days1. This broad definition enables surveillance systems to capture infections caused by influenza, respiratory syncytial virus (RSV), and SARS-CoV-2, among other respiratory pathogens. The public health burden due to ILI is substantial: during the 2024–2025 season alone, the U.S. experienced an estimated 47 million influenza cases with 610,000 hospitalizations, while RSV contributed 190,000 hospitalizations and COVID-19 added 250,000 more24. Accurate early detection of these outbreaks is essential for timely interventions, including vaccination campaigns, allocation of healthcare resources, and public health messaging before healthcare systems become overwhelmed.

Current surveillance systems face critical timing limitations that constrain proactive public health responses. For example, the United States Centers for Disease Control and Prevention (CDC)’s influenza surveillance network provides comprehensive data but often reflects events from 1 to 2 weeks earlier in the season, resulting in reports retrospective rather than predictive5. This delay hampers public health authorities’ efforts to implement preventive measures before outbreaks of respiratory diseases reach their peaks. State-of-the-art efforts to forecast the timing of the onset and peak of respiratory diseases outbreaks in the US, led by the FluSight CDC initiative, have remained experimental in nature. These models define outbreak onset using the standard CDC criterion of sustained activity above a fixed surveillance threshold, whereas alternative definitions based on sustained growth patterns have also been proposed in the literature612. The CDC FluSight models are evaluated for their ability to identify the timing of onsets and peaks only after the events in question have occurred. Even in this retrospective mode, the forecast scores (ranging 0.08–0.49 for the peak week and 0.01–0.41 for onset week) indicate that these models, on average, assign a probability between 1% and 49% that these events occurred in the correct week7. Accordingly, reported performance metrics across studies should be interpreted as contextual benchmarks rather than direct head-to-head comparisons, given differences in event definitions, modeling objectives, and evaluation protocols. This broad range suggests that these forecasts are widely uncertain about the timing and occurrence of the onsets and peaks they aim to identify. Further complicating surveillance efforts are the multiple circulating types of respiratory infections which contribute to annual ILI incidence. Recent work has shown that influenza and RSV outbreaks have distinct timing relationships which vary across states; these differences are especially important in the context of healthcare infrastructure and the deployment of vaccinations and other therapeutics13.

To enhance and refine surveillance efforts, researchers and public health authorities have embraced data from digital traces –information that is left behind by internet users navigating the web– which have emerged as promising indicators of infectious disease activity1420. Indicators such as Google search trends for respiratory symptoms and treatments have been shown to capture population-level changes in health-seeking behavior that often precede clinical case reporting1519,2125. Previous studies have demonstrated the utility of these digital traces for COVID-19 surveillance and their limitations2628, with Kogan et al.29 developing Bayesian indicators that forecast outbreaks weeks in advance, and Stolerman et al.30 creating county-level early warning systems using digital signals to estimate the reproductive number of county-level COVID-19 outbreaks. However, these approaches have primarily focused on single diseases and have not addressed the challenge of monitoring multiple respiratory diseases simultaneously across diverse geographic regions with varying outbreak dynamics.

Building on these advances, we developed a comprehensive machine learning framework that addresses key limitations in current respiratory disease surveillance. Unlike previous single-disease approaches, our early warning system integrates multiple data streams simultaneously to monitor ILI, influenza hospitalizations, and RSV activity across all 50 U.S. states. Rather than relying on complex reproductive number estimations which are prone to reporting delays and underascertainment3133, our approach leverages ensemble voting mechanisms with digital behavioral signals (primarily influenza-related Google search activity) to directly predict outbreak onsets and peaks. We additionally use transfer learning to extend our framework to different respiratory diseases despite the limited availability of historical data. We validated our methodology retrospectively using historical ILI data (2010–2020), tested retrospectively during the 2022–2024 ILI seasons in an out-of-sample way, and deployed prospectively and in real-time for the 2024–2025 ILI season in operational collaboration with the CDC’s Center for Forecasting and Outbreak Analytics, with predictions shared continuously with public health decision-makers beginning in August 2024.

Our work represents a substantial advancement in the field of digital epidemiology and surveillance. While past studies have explored transfer learning for digital epidemiology primarily in the setting of continuous incidence estimation from web-search time series, we instead focus on a different problem setting: prospective early-event detection rather than continuous rate inference. For example, Zou et al. proposed a three-step transfer framework that (i) learns a supervised, regularized regression model (constrained elastic net) to infer ILI rates in a source location with available ground truth, (ii) maps source search queries to target-location queries using a hybrid of semantic similarity (via mono- and cross-lingual word embeddings) and temporal similarity (time-series correlation with allowable seasonal shifts), and (iii) transfers and re-weights the learned query coefficients to estimate ILI rates in a target location without requiring target surveillance data for training34. Similarly, Lampos et al. developed models for tracking COVID-19 dynamics using online search data, combining supervised and unsupervised symptom-based signals and transferring learned relationships across countries by aligning search-query time series and adjusting for domain-specific biases such as media effects, with the goal of reconstructing continuous case trajectories in data-limited settings35. In this study, we construct a real-time early-warning system by aggregating heterogeneous proxy detectors into a calibrated voting signal, where location-specific decision thresholds are learned from prior seasons and then applied prospectively to trigger interpretable onset and peak alarms within a fixed lead-time window.

While our prior work demonstrated the feasibility of search-query-based early warning systems for COVID-19 in retrospective and quasi-prospective settings, the present study addresses a fundamentally broader and more operational problem. Here, we introduce a unified early warning framework that operates prospectively across multiple respiratory pathogens with distinct epidemiological dynamics, supports both onset and peak detection, and is evaluated through real-time deployment at the state level in collaboration with public health authorities. This shift from disease-specific retrospective analysis to multi-pathogen, real-time operational surveillance constitutes a central scientific advance of the present work.

Results

Respiratory disease outbreak timing varies dramatically across U.S. states: our analysis of historical ILI data (2010–2020) shows onset timing varies by an average of 10.7 weeks and peak timing by up to 20 weeks (mean range 14.4 weeks) across states (Supplementary Figs. S1, S2). This unpredictability—with some states beginning ILI season within a one-month window while others exhibit 2–4 months of temporal spread—makes historical patterns unreliable for prospective prediction and limits their utility for public health preparedness.

To address this critical gap, we developed and validated a machine learning early warning system that integrates digital behavioral signals (Google search trends) with traditional epidemiological data through ensemble voting algorithms to anticipate outbreak onsets and peaks 2–6 weeks in advance, as summarized in Table 1. We validated our approach through three phases: retrospective validation (2010–2020), out-of-sample testing (2022–2024), and real-time operational deployment (2024–2025) in collaboration with CDC’s Center for Forecasting and Outbreak Analytics, with prospective predictions shared continuously beginning in August 2024.

Table 1.

Performance of the early warning system (EWS) across different target diseases and validation phases

Target disease Validation phase Onset Peak
Det. (%) Lead (wk) PPV (%) Det. (%) Lead (wk) PPV (%)
ILI Retrospective (2014–2020) 99.3 3.4 93.0 84.2 3.9 69.3
ILI Out-of-sample (2022–2024) 99.0 4.8 96.0 96.1 3.1 82.7
ILI Real-time (2024–2025) 98.0 5.0 83.0 97.0 2.0 82.7
Influenza hosp. Out-of-sample (2024–2025) 96.0 4.0 82.8 60.8 2.6 94.4
RSV activity Out-of-sample (2024–2025) 87.2 1.2 100.0 68.0 2.0 84.6

Det. detection rate, Lead lead time, wk weeks, hosp. hospitalizations, PPV positive predictive value, ILI influenza-like illness, RSV respiratory syncytial virus.

We assessed performance using three operationally relevant metrics: (i) detection sensitivity (proportion of onsets and peaks successfully anticipated), (ii) lead time (average weeks in advance), and (iii) positive predictive value (PPV; alarm reliability). These metrics, evaluated across all validation phases, quantified how well the EWS narrowed wide historical timing variability into consistent, actionable prediction windows. The main contributions are (1) development and real-time validation of the ILI early warning system across all 50 U.S. states, and (2) successful transfer learning extension to influenza hospitalizations and RSV activity with minimal training data.

Early warning system for ILI

We designed our early warning system by inspecting ILI temporal trends and the American public’s influenza-related Google searches at the state and national levels from 2010 to 2020. Specifically, we used time series anomaly detection approaches to characterize the timing of the onsets and peaks in the weekly temporal trends of ILI. To label the onsets of outbreaks of ILI, we identified the beginning of time periods with consistent exponential growth in the number of reported ILI cases29,30 and implemented standard peak detection algorithms36. It is important to note that both of these strategies, when implemented in real-time and for prospective purposes, require future-looking data to detect the timing of the event. For example, generally speaking using standard peak detection packages, such as SciPy in python36 or scorepeak in R37, a peak can only be identified when one observes that the value of the time series, at a given week, is higher than both the previous value of the time series and the next value of the time series. As a result, both of our strategies label an event (either an onset or peak) about 3 weeks after it has occurred. Unlike onset detection, which identifies the beginning of sustained epidemic growth, peak detection requires identifying the maximum of an outbreak trajectory before the full epidemic curve has unfolded. This makes peak detection substantially more challenging in real-time settings, as it must operate under noisy surveillance signals and without access to future incidence.

System development and retrospective validation (2010–2020)

Our early warning system borrows elements from methodologies implemented for real-time onset identification in the context of COVID-1929,30, and consists of identifying historical temporal anomalies in both the patterns of the general public’s disease-related Google searches and in events that may occur in neighboring states that anticipate the emergence of an event in a given state. Intuitively speaking, we hypothesize that a spike in cases of a disease may be detected in a population when sharp increases (exponential growth) in multiple disease-related search queries occur. We identify these queries and the number of queries needed to label an event using an ensemble voting methodology, which combines the results of multiple machine learning models to generate a final prediction. More details are presented in the subsequent sections.

To establish confidence in our methodology, we conducted extensive retrospective evaluations and validation using historical ILI data spanning 10 seasons, starting in the fall of 2010. These analyses were performed for all 50 U.S. states and the national level, with the exception of Florida, where ILINet data were not available. These evaluations allowed us to identify that a training set consisting of a rolling window of the most recent three years to predict subsequent outbreak patterns was optimal, providing a favorable trade-off between sufficient historical data for stable calibration and adaptability to evolving digital search behavior.

With such a choice of training set length, we evaluated the performance of our approach retrospectively during the time period from the fall of 2014 to winter of 2020. In this time period, our EWS detected 99.3% of outbreak onsets, with a mean lead time of 3.4 weeks and a positive predictive value (PPV) of 93.0% (meaning that 93.0% of alarms that were produced anticipated the occurrence of an onset and only 7.0% of alarms were false or too early). To assess the sensitivity of these results to the choice of onset definition, we repeated the retrospective evaluation using the standard CDC ILI onset criterion (three consecutive weeks above a baseline defined as the mean plus two standard deviations of non-seasonal weeks from the previous three seasons) as the target event. Under this alternative definition, the EWS detected 97.3% of CDC-defined onsets, with an average lead time of 3.8 weeks and a positive predictive value (PPV) of 93.4% (Supplementary Fig. S8). These results indicate that the strong onset detection performance observed in the retrospective analysis is robust to the operational definition of onset and is not driven by the use of a growth-based target. Peak detection during this retrospective validation period reached 84.2% sensitivity (the percent of true events that were correctly anticipated), an average of 3.9-week lead time, and 69.3% PPV (Supplementary Figs. S7, S9). This means that among 433 alarms 293 actually anticipated a peak. The fact that in 50 locations during the 6 years of validation we observed 316 peaks, suggests that during a given season there may be locations that may experience more than one peak, making this specific task more challenging than onset detection. These results established the foundational performance characteristics of our ensemble voting approach and validated the correlation-based feature selection methodology. Having established the system’s performance characteristics on historical data, we next evaluated whether these results would hold on previously unseen seasonal patterns.

Out-of-sample validation (2022–2024)

Out-of-sample testing on the 2022–2023 and 2023–2024 seasons was used to further validate the system’s predictive capabilities. All hyperparameters established during development—including the onset and peak definitions, the three-season training window, and the correlation threshold for Google search term selection—were fixed during this evaluation. The only parameter adjusted was the voting system threshold (see Methods for details). The system successfully detected 99.3% of outbreak onsets, with a mean lead time of 4.2 weeks and 96.0% PPV, while peak detection achieved 96.1% sensitivity, an average of 3.1-week lead time and 82.7% PPV. These results demonstrated consistent performance across different seasonal patterns and confirmed the system’s readiness for real-time operational deployment. This high performance suggests that our EWS demonstrates robust generalizability of the digital behavioral signals as predictive variables and ensemble voting mechanisms across evolving epidemiological contexts. The consistent performance across the 2022-2024 seasons confirmed the system was ready for operational real-time deployment with prospective predictions shared with public health authorities before events occurred.

Prospective and real-time operational deployment evaluation (2024–2025)

We next deployed our EWS in real-time during the 2024-2025 ILI season through operational collaboration with the CDC’s Center for Forecasting and Outbreak Analytics. Prospective predictions were shared continuously with CDC researchers and decision-makers beginning in August 2024, providing independent validation of the system’s operational performance. During this time period, our early warning system demonstrated robust performance in detecting respiratory disease outbreaks across all 50 US states. The system successfully anticipated the timing of 98.0% of outbreak onsets, with an average lead time of 5 weeks and achieved a PPV of 83.0% (Fig. 1). Only 14.0% of alarms represented false positives not followed by confirmed outbreaks in the subsequent 6 weeks. The system narrowed the historically wide onset detection window to a consistent 4–6 week range across nearly all states, representing a 2–4 fold reduction in timing uncertainty for states with the highest historical variability.

Fig. 1. Real-time early warning system (EWS) onset detection performance for influenza-like illness (ILI) during the 2024–2025 season, during operational deployment with the U.S.

Fig. 1

Centers for Disease Control and Prevention (CDC) Center for Forecasting and Outbreak Analytics, achieving 98% detection with an average 5-week lead time and a positive predictive value (PPV) of 83%. A Overall classification of onset detections across United States (U.S.) states, categorized by warning type—Early, Synchronous, Late, Soft, and Missed. The donut chart shows that the majority of onsets were detected early, with a smaller portion detected synchronously or late, and a minimal proportion missed entirely. B Distribution of all detection alarms, including false alarms. This breakdown illustrates the relative frequency of true positives, false positives, and the proportion of alarms that did not correspond to an actual onset. C Schematic diagrams illustrating the temporal relationship between predicted and actual onsets across different detection types. Multiple scenarios are visualized, including early, synchronous, and late detections, as well as false alarms. These curves represent hypothetical trajectories of onset activity, with corresponding system predictions overlaid for clarity. These curves are not model outputs and should not be interpreted as shifted, smoothed, or transformed versions of observed incidence curves. D Histogram of detection earliness for onsets. The x-axis shows the number of weeks between the detection and the actual onset, where negative values represent early warnings, zero represents detection during the onset week, and positive values indicate late warnings. The majority of detections occurred 3–5 weeks before the onset, indicating good anticipatory performance. Source data are provided as a Source Data file.

Peak detection performance was similarly strong, and the system accurately identified the timing of epidemic peaks in 97.0% states, with an average lead time of 2 weeks, with PPV of 82.7%. 17.3% of peak alarms were classified as false positives (Fig. 2). This performance significantly outpaced the substantial historical uncertainty in peak timing, where natural variability spans 13–20 weeks in the most unpredictable states. The slight reduction in PPV compared to retrospective validation (83.1% vs. 92.9% for onsets) reflects the inherent challenges of real-time prediction in operational settings, including evolving disease dynamics and potential changes in digital behavior patterns. However, the maintained high sensitivity (97.0 % and 98.0%) demonstrates the system’s reliability for public health decision-making.

Fig. 2. Real-time early warning system (EWS) peak detection performance for influenza-like illness (ILI) during the 2024–2025 season, during operational deployment with the U.S.

Fig. 2

Centers for Disease Control and Prevention (CDC) Center for Forecasting and Outbreak Analytics, achieving 97.0% detection with an average 2-week lead time and a positive predictive value (PPV) of 82.7%. A Overall classification of peak detections across United States (U.S.) states, categorized by warning type—Early, Early-Late, Synchronous, Late, Soft, and Missed. The donut chart shows that the majority of peaks were detected early, followed by a smaller proportion of synchronous, late, and early-late warnings. B Distribution of all detection alarms, including false alarms. This breakdown illustrates the relative frequency of true positives, false positives, and the proportion of alarms that did not correspond to an actual peak. C Schematic diagrams illustrating the temporal relationship between predicted and actual peaks across different detection types. Multiple scenarios are visualized, including early, early-late, synchronous, and late detections, as well as false alarms. These curves represent hypothetical trajectories of peak activity, with corresponding system predictions overlaid for clarity. These curves are not model outputs and should not be interpreted as shifted, smoothed, or transformed versions of observed incidence curves. D Histogram of detection earliness for peaks. The x-axis shows the number of weeks between the detection and the actual peak, where negative values represent early warnings, zero represents detection during the peak week, and positive values indicate late warnings. The majority of detections occurred 3–6 weeks before the peak or the week after the peak preceding the algorithmic detection, indicating good anticipatory performance. Source data are provided as a Source Data file.

To further characterize variability in detection timing beyond mean lead times, we also examined the full distributions of detection earliness (Figs. 1D and 2D). For outbreak onset detection, lead times were tightly concentrated in the early range, with most detections occurring ~3–6 weeks prior to the realized onset and very few detections occurring during or after the onset week. This indicates consistent early detection across locations. In contrast, peak detection exhibited substantially greater variability in detection timing. Although many peaks were detected several weeks in advance, the distribution spanned both early and post-peak detections, with a non-negligible fraction occurring shortly after the realized peak. Consequently, the mean peak detection lead time of ~2 weeks reflects a broad distribution of detection times rather than a uniform advance warning across states.

Transfer learning extension to RSV activity, influenza hospitalizations, and COVID-19 activity

Based on the strong performance of our EWS using both retrospective and prospective data, we next used domain transfer learning to extend our framework to additional targets for respiratory disease. The feasibility of this approach is based on the hypothesis that the search activity related to respiratory disease will anticipate events for other indicators of respiratory diseases, such as hospitalizations for RSV and influenza. In practice, transfer learning was implemented by reusing the feature selection, model structure, and alarm-generation rules established for ILI, while re-calibrating the alarm threshold with the limited available data for each new target disease. This strategy allows models trained on long ILI histories to be rapidly adapted to diseases with only one or two seasons of observations. This transfer learning approach has been shown to be viable despite the limited availability of historical data38. If such a hypothesis is correct, we could use the training approaches and feature selection techniques developed for ILI and rapidly design and train a new surveillance system for a related respiratory disease indicator with minimal training data.

Influenza hospitalizations

For influenza hospitalizations, where only two seasons of training data were available (2022–2024), the system achieved 96.0% onset detection during the 2024–2025 season with 4-week average lead time and 82.8% PPV. Peak detection for influenza hospitalizations reached 60.8% sensitivity with a 2.6 week lead time and 94.4% PPV, with only 5.6% false alarms (Supplementary Fig. S10). The reduced peak detection sensitivity compared to ILI likely reflects differences in hospitalization reporting patterns and the more severe disease threshold, which may exhibit different relationships with early digital behavioral signals.

RSV activity

RSV surveillance presented the greatest transfer learning challenge, with training data limited to just one historical season (2023–2024). Despite this severe constraint, the transfer learning approach successfully detected 87.2% of RSV activity onsets during 2024–2025 with 1.2-week lead time and achieved 100% PPV with zero false alarms. RSV peak detection reached 68.0% sensitivity with 2.0 weeks of lead time and 84.6% PPV (Supplementary Fig. S11).

The perfect PPV for RSV onset detection observed during the 2024–2025 season, despite limited training data, suggests that the digital behavioral signals for RSV may be particularly relevant to RSV activity and that the transfer learning approach offers a promising way to capture the essential predictive signals that arise during the RSV season. The shorter lead times (1.2 weeks vs. 5 weeks for ILI) may reflect more rapid progression of RSV outbreaks or differences in population search behavior patterns.

COVID-19 activity

COVID-19 surveillance represents one of the most challenging transfer learning scenarios considered in this study, as no historical COVID-19 data were available for training at the time of deployment. Consequently, the early warning system was trained exclusively on pre-pandemic ILI data (2017–2019) and applied prospectively to detect the first COVID-19 outbreak signals. Despite this severe mismatch between training and target domains, the system identified outbreak activity in a substantial fraction of cases. Out of 49 observed COVID-19 outbreaks, 25 were detected in advance, while an additional 13 were detected within three days after the epidemiological onset, corresponding to near-synchronous warnings. Eleven outbreaks were not detected, and 16 false alarms were generated during the evaluation period (Supplementary Fig. S12). Across detected events, the system achieved an average lead time of 2.6 weeks. These results highlight that, although absolute accuracy is inherently constrained in such an extreme transfer setting, the primary strength of the proposed framework lies in its rapid adaptability: the ability to provide timely situational awareness for a novel pathogen using only pre-existing respiratory surveillance knowledge, rather than requiring pathogen-specific historical training data.

Comparative performance across respiratory pathogens

Our early warning system demonstrated consistent early detection capabilities across multiple respiratory disease targets during 2024–2025, achieving onset detection rates of 98.0% for ILI, 96.0% for influenza hospitalizations, and 87.2% for RSV activity, with lead times ranging from 1.2 to 5 weeks. Peak detection performance varied by disease target, with ILI showing the strongest performance (97.0% detection) followed by RSV (68.0%) and influenza hospitalizations (60.8%). Positive predictive values remained consistently high across all targets, ranging from 82.7% to 100%, indicating low false alarm rates suitable for operational public health decision-making. The performance gradation from the ILI EWS to those for influenza hospitalizations and RSV likely reflects both the reduced amount of available training data (10+ years vs. 2 seasons vs. 1 season) and inherent differences in disease progression patterns and their relationships to digital behavioral signals. The system’s ability to maintain operational performance across diseases with vastly different training data availability highlights the effectiveness of our transfer learning approach and suggests broad applicability to emerging respiratory threats. As additional training data becomes available, we expect continued performance improvements, particularly for peak detection in newer surveillance targets.

Discussion

Our early warning system achieved 98% onset detection and 97.0% peak detection for ILI during real-time deployment with average lead times of 5 weeks and 2 weeks, respectively. These results directly address the core problem that has limited implementation of early warning systems for public health preparedness: historical ILI onset timing varies by 10–20 weeks across states, while our system provides consistent 4–6 week prediction windows for onsets and 2–4 week windows for peaks. For states with the highest historical variability, this represents a 2–4 fold reduction in timing uncertainty, enabling health departments to plan vaccination campaigns, adjust hospital staffing, and coordinate resource allocation with specific timeframes rather than broad seasonal estimates. Furthermore, we demonstrate that the methods used in our ILI EWS can be applied to data from individual respiratory infections such as time series of influenza hospitalizations or RSV activity.

Beyond performance metrics, the significance of this work lies in how early warning for infectious disease outbreaks is formulated and evaluated in an operational public health context. Unlike prior digital surveillance studies that emphasize retrospective accuracy or disease-specific analyses, our framework is explicitly designed for prospective deployment and real-time decision-making. By unifying onset and peak detection across multiple respiratory pathogens within a single calibrated voting framework, the system moves early warning from a methodological proof-of-concept toward a deployable surveillance capability. This distinction is particularly important for peak detection, which requires anticipating epidemic maxima before the full outbreak trajectory is observed and poses substantially greater challenges in real-time settings.

Importantly, these results were independently validated through our operational collaboration with CDC’s Center for Forecasting and Outbreak Analytics. By sharing prospective predictions continuously with CDC experts beginning in August 2024—before outbreaks occurred—we established transparent, real-time verification of system performance that distinguishes this work from retrospectively evaluated forecasting efforts.

Overcoming fundamental limitations in traditional disease surveillance

The results of this study suggest that our EWS methodology is a distinct improvement over traditional surveillance and forecasting methods29,30. Rather than estimating effective reproductive numbers prone to reporting delays and under-ascertainment31,32, we directly leverage early behavioral signals through an ensemble voting strategy that aggregates Google search trends and neighboring state disease dynamics. By predicting outbreaks using only preceding signals, we avoid the pitfall of extrapolating estimates from delayed surveillance data that often becomes available only after events have occurred, resulting in models experiencing look-ahead bias.

As a result of these improvements, our models performed well compared to existing CDC FLUsight challenge results (see Supplementary Figs. S3S6). While current forecasting models achieve onset detection scores of 0.01–0.41 (assigning 1-41% probability to the correct timing week) and peak detection scores of 0.08–0.497, our system achieved 98% onset detection and 97.0% peak detection with 5-week and 2-week mean lead times, respectively, under the exponential growth-based onset definition used throughout this study. To facilitate comparison with prior work, we additionally evaluated our system using the standard CDC ILI onset definition based on sustained activity above a baseline computed from off-season ILI levels. Under this alternative definition, the system achieved 97.3% onset detection, a comparable mean lead time, and a PPV of 93.4%, demonstrating that our performance is robust to the choice of onset definition. This improvement reflects the advantage of utilizing behavioral signals that anticipate rather than follow epidemiological patterns. When people search for respiratory disease symptoms or treatments, they often do so before seeking clinical care, creating an early warning signal that our ensemble voting approach effectively captures15,16,18,2123. Our Granger causality analysis quantifies this temporal relationship, with 92.7% of selected search terms demonstrating statistical precedence over disease activity, providing empirical support for the intuitive notion that health-seeking behavior manifests first in digital spaces before appearing in clinical settings.

While these results demonstrate the advantages of leveraging early behavioral signals for outbreak anticipation, we emphasize that the effectiveness of the proposed EWS depends on the existence of informative proxy signals that precede clinical surveillance. Performance is therefore expected to degrade in settings where emergent pathogens exhibit minimal symptomatic overlap with historically observed diseases or where no suitable historical proxy data exist for calibration. In such scenarios, additional methodological adaptations—such as alternative proxy selection, dynamic reweighting, or incorporation of non-search-based data streams—would be required prior to deployment. These boundary conditions are intrinsic to any surveillance system relying on behavioral signals and do not diminish the utility of the framework for respiratory pathogens with shared symptomatology, where early digital traces provide actionable lead time that is advantageous to public health strategy.

Transfer learning demonstrates scalability despite data constraints

The transfer learning framework we implemented successfully extended surveillance capabilities of our EWS to outbreaks caused by other respiratory diseases with varying data availability38. The progression from ILI (10+ years training, 98% onset detection) to influenza hospitalizations (2 seasons training, 96% onset detection) to RSV (1 season training, 87.2% onset detection) demonstrates both the framework’s adaptability and the expected relationship between historical data availability and detection performance. While performance declined as a result of the transfer learning process, the impacted performance was expected as a result of the decrease in the quantity of available training data. Despite minimal training data, the RSV implementation achieved 100% positive predictive value for onset detection, indicating that when the system triggered alarms, the alarms consistently preceded actual outbreaks. However, this finding should be interpreted cautiously given the single-season training constraint and relatively small sample size. The lower overall detection rate (87.2%) suggests that while the system produces few false alarms for RSV, predictions for RSV outbreaks may not conform to the limited historical patterns available for training.

The varying performance characteristics across disease targets provide insights into the relationship between clinical presentation and digital behavioral signals, though these interpretations require further validation. RSV’s shorter lead times (1.2 weeks vs. 5 weeks for ILI) may reflect the more rapid progression of RSV outbreaks or differences in population search behaviors, particularly given RSV’s disproportionate impact on vulnerable populations (such as infants and the elderly) who may have caregivers searching on their behalf rather than self-searching. Similarly, the reduced peak detection sensitivity for influenza hospitalizations (60.8%) compared to ILI (97.0%) likely reflects the higher severity threshold for hospitalization, which may have weaker relationships with early behavioral signals that precede milder symptomatic illness.

Taken together, these results suggest potential for rapid deployment of early warning systems to outbreaks of emerging respiratory threats, though with important caveats. The consistent detection of respiratory disease signals across different pathogens indicates that behavioral patterns may share common elements, but the limited training data for newer targets means that initial deployment performance may be constrained. The framework’s value for pandemic preparedness lies primarily in its ability to provide early warning capability quickly, even if initial performance is modest, with the expectation that accuracy would improve as outbreak data accumulates. This capability is precisely what’s needed for pandemic preparedness: the ability to deploy surveillance systems to novel pathogens that share symptoms with existing pathogens for which prior work has shown correlated surveillance signals20,29, within weeks rather than waiting years to accumulate sufficient training data, as demonstrated by our implementation for the first onset of COVID-19 in 2020, using ILI data to train our models, and our RSV implementation, using only a single season of historical RSV observations.

Implementation of such early warning systems will require buy-in from both scientists and public health authorities, in addition to careful coordination with administrators of healthcare facilities and leaders in local government.

Immediate operational applications may transform public health response

The predicitive abilities of these early warning systems enable immediate practical applications that demonstrate the framework’s operational value for evidence-based public health decision-making. For example, current ACIP guidelines recommend uniform October-March administration of the RSV prophylactic nirsevimab that cannot account for substantial state-level variability in RSV timing. Our system’s ability to predict RSV onsets with 1–2 weeks of lead time with high positive predictive value may enable health departments and hospitals to optimize distribution of prophylactics or vaccines by adjusting administration start dates for early seasons, prioritizing shipments based on outbreak predictions, and minimizing both under-protection and wasted coverage from premature dosing13,3941.

Similar applications extend to influenza vaccination campaigns, where 5-week lead time of ILI onsets enables targeting of high-risk populations before community transmission accelerates. Hospital surge planning benefits from 2-week peak detection lead time, allowing for staffing adjustments, supply chain preparation, and patient flow optimization. These applications demonstrate how the framework transforms reactive public health responses into proactive epidemic preparedness, delivering on the scalable monitoring capability promised in our approach.

More broadly, the proposed early warning system can be viewed as a digital epidemic intelligence (EI) framework that integrates indicator-based surveillance signals with event-oriented decision rules to provide early situational awareness. In this sense, our approach is conceptually aligned with existing EI paradigms such as the WHO Epidemic Intelligence from Open Sources (EIOS)42 and long-standing event-based monitoring systems including the Global Public Health Intelligence Network (GPHIN)43, HealthMap44, and the CDC Center for Forecasting and Outbreak Analytics, which emphasize the synthesis of heterogeneous data streams to support timely public health decision-making. Unlike traditional EI systems that rely heavily on qualitative signal triage or continuous incidence forecasts, our framework formalizes the detection of discrete epidemiological events—specifically outbreak onsets and peaks—using location-specific calibration and prospective evaluation.

From an operational perspective, such a system could be integrated into national or global EI workflows as a complementary decision-support layer, flagging regions at elevated risk of imminent outbreak escalation or decline and thereby guiding further investigation, resource allocation, or expert review. Because the framework does not depend on pathogen-specific mechanistic assumptions and instead leverages transferable digital behavioral signals with minimal retraining, it is well suited for rapid adaptation to emerging pathogens or pandemic scenarios characterized by limited historical data. As such, the proposed EWS represents not only a methodological advance in digital surveillance, but also a strategic contribution toward scalable, interpretable, and prospective epidemic intelligence infrastructures.

Current limitations and implications for generalizability

This study has several limitations. First, constrained training data for newer surveillance targets and inherent challenges in peak detection methodology results in performance differences across pathogens, with peak detection ranging from 60.8% for influenza hospitalizations to 97.0% for ILI. These differences reflect both data availability constraints and fundamental differences in how severe disease manifestations relate to early behavioral signals. Short training histories limit generalizability, though performance trends suggest substantial improvement as data accumulate. Second, we selected a 6-week window as it represented a reasonable balance for identifying outbreak onsets for long outbreaks. Importantly, alternative choices of this parameter would have yielded qualitatively similar results across respiratory diseases, as the qualitative patterns remained consistent, see Supplementary Materials section Sensitivity of Onset Detection to Exponential Growth Duration. This consistency arises because the same methodological framework (Supplementary Algorithm 1) was applied uniformly across all time series—ILI, RSV, influenza, and Google searches—so any modification of the parameter would have produced analogous effects across diseases and predictors without altering the overall conclusions. Third, our deliberate choice of binary outputs over probabilistic forecasts prioritizes interpretability for decision-makers while maintaining high positive predictive values (82–100%) across disease targets. We opted against using Bayesian probabilistic approaches given limited training data availability, which would yield poorly calibrated probabilities overly sensitive to modeling assumptions. This design choice reflects the operational reality that public health officials need actionable yes or no decisions rather than complex probabilistic distributions. Lastly, the inherent challenge of peak detection—requiring post-peak data—results in some late warnings, though many soft warnings that approach alarm thresholds still provide valuable confirmatory situational awareness. Additionally, our reliance on digital behavioral signals creates potential vulnerabilities to changes in search behavior patterns or platform algorithms2628. Such effects have been widely documented in prior digital surveillance work, including the ARGO framework, which explicitly notes the sensitivity of search-based signals to behavioral overreaction and changes in search engine design, and emphasizes aggregation and rolling retraining as necessary mitigation strategies21. While our use of proxy aggregation and rolling recalibration improves robustness, residual behavioral drift remains an inherent limitation of behavior-driven surveillance systems rather than a methodological shortcoming unique to this approach.

Broader implications for pandemic preparedness and future directions

In conclusion, our framework’s significance extends beyond seasonal respiratory disease surveillance to pandemic preparedness and emerging threat detection20,33. Collaborative forecasting efforts during COVID-19, while valuable for longer-term projections45, highlighted the continued need for early warning systems that can detect outbreak onsets and peaks prospectively with actionable lead times. The disease-agnostic approach we employed enables rapid deployment to novel pathogens by leveraging behavioral signals that precede clinical reporting, potentially serving as a generalized monitoring system for future respiratory threats. The successful real-time deployment during 2024–2025, achieving detection rates exceeding 95% for onsets across multiple pathogens with 2–5 week lead times, demonstrates that integrating digital behavioral signals with traditional epidemiological data can transform reactive public health responses into proactive epidemic preparedness. Future research directions should focus on expanding the behavioral signal repertoire beyond Google search trends to include internet searches from clinicians and other providers46, social media patterns, mobility data, and healthcare utilization signals. Integration with genomic surveillance data could enhance specificity for emerging variants or novel pathogens. Development of adaptive weighting approaches that dynamically update ensemble parameters within seasons could improve performance for rapidly evolving outbreaks. In parallel, continued refinement of digital behavioral signals—particularly internet search data—represents an important avenue for improving robustness and operational performance. Prior work has shown that careful aggregation, dynamic selection, and regularization of search queries can substantially mitigate behavioral noise and non-stationarity in digital disease surveillance systems47. Building on these insights, future extensions of our framework will explore more structured representations of search behavior, including dynamically clustered queries and adaptive weighting schemes that evolve within and across seasons. As shown in Supplementary Table S1, the inclusion of Google search terms provides consistent gains for peak detection across pathogens and validation settings, while maintaining comparable performance for outbreak onset detection. Notably, as shown in Supplementary Table S2, Google search signals alone already exhibit substantial early-warning capability, particularly for peak detection, achieving multi-week lead times across multiple seasons, which underscores their intrinsic value as a complementary data source. These findings suggest that refined digital signal representations are likely to be especially impactful for peak detection, where behavioral signals provide early indications of rapid epidemic acceleration not yet visible in traditional surveillance data. For outbreak onset detection, where surveillance-based signals already perform strongly, such enhancements are expected to primarily improve lead-time stability and robustness under shifting public attention rather than detection accuracy itself.

The international transferability of this approach presents both opportunities and challenges. While the underlying premise—that health-seeking behaviors precede clinical case reporting—likely holds across healthcare systems, differences in digital infrastructure, search behavior patterns, and healthcare access may require region-specific calibration. Collaborative international implementation could provide valuable comparative insights into digital behavioral signal universality. As digital data streams expand and training datasets accumulate, this approach offers a sustainable pathway for enhancing real-time disease surveillance capabilities at the scale and resolution needed for effective public health action. The demonstrated ability to maintain operational performance across diseases with vastly different training data availability suggests broad applicability to emerging respiratory threats, fulfilling the scalable monitoring promise essential for future pandemic preparedness.

Methods

This study utilized aggregated, de-identified surveillance data that are publicly available. The study was determined to not constitute human subjects research, and institutional review board (IRB) approval was not required.

Our approach combines traditional epidemiological surveillance with digital behavioral signals through an ensemble voting mechanism that predicts outbreak onsets and peaks 2–6 weeks in advance. The system operates by identifying early signals from proxy data sources—Google search trends and neighboring state disease activity—that historically precede confirmed outbreaks in target locations.

Data sources

We integrated multiple data streams to develop and validate our early warning system across different respiratory disease targets. The following datasets provided the foundation for model training, validation, and real-time deployment.

Official influenza-like illness reports

The dataset comprised a weekly time series of the percentage of outpatient visits due to Influenza-like illness (ILI), spanning from October 9, 2010, to April 12, 2025. This data covered each U.S. state and the national level, as provided by the Centers for Disease Control and Prevention (CDC)48. The extensive temporal coverage of ILI data enabled comprehensive model training and served as the primary target for initial system development and validation.

Official respiratory syncytial virus emergency department visits reports

The dataset comprised a daily time series of Respiratory Syncytial Virus (RSV) emergency department visits, spanning from October 8, 2022, to March 9, 2025. Data were available for all U.S. states except Missouri and were obtained from the National Syndromic Surveillance Program (NSSP)49. The limited temporal coverage of RSV data necessitated the use of transfer learning approaches to extend model capabilities to this respiratory disease target.

Official influenza hospitalizations reports

The dataset comprised a weekly time series of influenza hospitalizations, spanning from July 3, 2021, to May 24, 2025. This data covered each U.S. state and the national level, as provided by the Centers for Disease Control and Prevention (CDC) for the FLUsight challenge50. The intermediate temporal coverage of influenza hospitalization data provided an additional target for validating transfer learning capabilities.

Google trends digital behavioral signals

We utilized weekly datasets from the Google Trends API, incorporating search terms related to influenza and respiratory diseases to capture population-level behavioral changes that precede clinical case reporting. Key search terms included specific disease identifiers such as ’influenza b’, ’rsv’, and ’flu virus’. Additionally, terms representing common symptoms of respiratory illnesses, such as ’flu headache’, ’rsv symptoms’, and ’human temperature’, were incorporated to capture symptom-related search behavior. Furthermore, terms related to treatment and prevention, including ’over the counter flu’, ’tamiflu wiki’, and ’flu shot’, were used to enhance the dataset’s comprehensiveness in capturing health-seeking behavior patterns. All Google search queries used in the analysis—including pathogen-specific, symptom-related, and treatment-related terms—are listed in Supplementary Materials Box 1.

Data preprocessing and temporal alignment

Our data streams experience specific availability delays that required temporal adjustments for proper alignment. Google Trends data are available up to 6 days before the current date, necessitating a temporal shift where data reported at time t were shifted to time t + 6 to address the 6-day reporting delay. This adjustment ensures that digital behavioral signals maintain their predictive value without introducing forward-looking bias in real-time applications.

Early warning system methodology

The early warning system integrates epidemiological surveillance data with digital behavioral signals to identify outbreak patterns 2–6 weeks before peak healthcare demand. The methodology consists of three key components: outbreak event detection, predictive signal selection, and ensemble voting for alarm generation.

Outbreak event detection

We developed distinct algorithms to identify outbreak onsets and peaks in both target surveillance data (for model training) and proxy signals (for prediction). For onset detection, we adapted a methodology previously developed by Kogan et al.29 and Stolerman et al.30. Specifically, for each week t of our time series, we assessed exponential growth patterns over the preceding 6 weeks by computing a retrospective linear regression model to determine the multiplicative constant λt that maps case numbers from one week to the next (i.e., the coefficient of a lag-1 autoregressive model without intercept). This parameter λt serves as a proxy for the effective reproductive number Rt, a standard epidemiological measure indicating whether an epidemic is growing (λt > 1) or declining (λt < 1).

The methodologies then diverged based on data type to address temporal bias constraints. For target surveillance data used in model training, we could apply forward-looking validation and confirmed an onset when periods of λt > 1 were followed by six consecutive weeks of exponential growth where λt > 1. For proxy signals like Google Trends used in prediction, we avoided forward-looking bias by identifying onsets as the sixth week of exponential growth windows where λt > 1, ensuring that predictive signals remained strictly historical.

All analyses were implemented in Python (v3.9). Peak identification employs the SciPy library (v1.10)36 peak detection algorithm with three filtering criteria: minimum duration of 2.5 weeks, case counts exceeding the 75th percentile of previous seasons, and temporal separation of at least 20 weeks between peaks (unless a subsequent peak is higher). This approach inherently involves 1–3 week delays since peak confirmation requires observing both the rise and subsequent decline in disease activity.

Predictive signal selection and ensemble voting

For each target location, we perform correlation-based feature selection to identify the most predictive proxy signals from Google search trends and neighboring state surveillance data. Only proxies showing strong historical association (Pearson correlation  > 0.7) with target outbreak patterns during training periods are retained for the ensemble. This threshold was determined during the System Development and Retrospective Validation phase (2010–2020), where we systematically evaluated correlation cutoffs ranging from 0.5 to 0.9 and found 0.7 to yield the best trade-off between sensitivity and specificity (Sensitivity analysis can be found in the Supplementary Materials). To provide a supporting validation of this feature selection, we applied Granger causality tests to the retained proxies. Across rolling three-year windows, 92.7% of the selected terms were Granger-significant at least once (Supplementary Fig. S115), and 56.0% remained significant after adjustment for multiple comparisons using the false discovery rate (FDR51) (Supplementary Fig. S114). This provides a heuristic but meaningful check that the selected features align with temporal predictive structure in the data.

The ensemble voting mechanism aggregates selected proxy signals based on their historical performance in predicting outbreak events. Each proxy is evaluated using standard classification metrics: true positives (proxy events preceding target outbreaks by ≤6 weeks), false positives (proxy events not followed by target outbreaks within 6 weeks), and false negatives (missed target outbreaks). A formal mathematical specification of the event-level labeling, voting threshold calibration, and alarm triggering rules is provided in the Supplementary Materials. The decision threshold τ for triggering early warning alarms is calibrated as the lowest number of correct predictions achieved across the three most recent training outbreaks:

τ=minj=1,2,3(TPj) 1

where TPj represents the number of true positive votes from the three most recent training outbreaks. This conservative approach ensures high confidence in alarm generation while minimizing false alerts. The choice of a three-season training window was motivated by its use in similar early warning and machine learning frameworks21,29,30 and was further validated during the System Development and Retrospective Validation phase (2010–2020), where we empirically tested training horizons ranging from one to five seasons and found three seasons provided the most stable balance between sensitivity and specificity.

At each time point, the early warning system calculates the number of proxy signals that have activated within the previous three weeks. When this count exceeds the threshold τ, the system issues a warning predicting an outbreak event within the following six weeks. Figures 3 and 4 illustrate this process for ILI onset prediction in Illinois and peak prediction in Massachusetts, respectively, showing how individual proxy signals aggregate to trigger early warning alarms.

Fig. 3. Early warning system (EWS) for predicting influenza-like illness (ILI) onset in Illinois.

Fig. 3

The top three panels show proxy signals from Google Trends (e.g., “robitussin'') and neighboring states (Tennessee and Pennsylvania). Blue markers indicate signal activation events. The fourth panel displays the EWS output, with the number of activated proxies (light blue), and a dashed threshold line (11 proxies) above which an early warning is triggered (purple dashed vertical lines). The bottom panel shows the official U.S. Centers for Disease Control and Prevention (CDC) Influenza-like Illness Surveillance Network (ILINet) data for Illinois (gray shaded area) along with training and testing onsets (yellow and orange vertical lines, respectively), and overlaid coronavirus disease 2019 (COVID-19) cases (blue dotted line). The EWS successfully triggers an early warning before the ILI onset for the 2024–2025 season. Source data are provided as a Source Data file.

Fig. 4. Early warning system (EWS) for predicting influenza-like illness (ILI) peak in Massachusetts.

Fig. 4

The top three panels present proxy signals including Google search queries (e.g., “how contagious is the flu'', “how long are you contagious'', “flu cold'') and data from neighboring states. Blue ticks indicate proxy signal activation. The fourth panel shows the number of activated proxies used by the EWS, with early warnings triggered once the activation threshold (44 proxies, dashed line) is exceeded. The bottom panel illustrates U.S. Centers for Disease Control and Prevention (CDC) Influenza-like Illness Surveillance Network (ILINet) data for Massachusetts (gray shaded area) and the timing of training, testing, and early warnings. The system accurately anticipates peak ILI activity during the 2024–2025 season. Source data are provided as a Source Data file.

Transfer learning for influenza hospitalizations and RSV activity

To extend the framework to influenza hospitalizations and RSV activity despite limited historical data, we employed transfer learning by directly applying all model parameters, feature selection criteria, and methodological components from the trained ILI system (2010–2020). A common pool of candidate proxy signals—including Google search terms related to respiratory symptoms and healthcare-seeking behavior—is used across pathogens; however, during training, proxies are filtered and retained based on their historical predictive alignment with each specific target, resulting in an effective signal set that is pathogen-specific. Following this selection step, the alarm threshold (τ) is calibrated using recent historical outbreaks to balance early detection performance against false alarm rates.

Validation strategy

A critical feature of our validation approach is strict temporal integrity: at no point did future data influence past predictions. This mirrors real-world deployment conditions where only historical data informs forecasts. We employed three validation phases for ILI: retrospective validation (2014–2020) for hyperparameter tuning and initial performance assessment, out-of-sample testing (2022–2024) for out-of-sample validation, and real-time deployment (2024–2025) for operational performance evaluation. For influenza hospitalizations and RSV activity we employed a prospective testing (2024–2025) for out-of-sample validation.

The training strategy uses rolling three-season windows to capture evolving disease dynamics while maintaining temporal relevance. For each prediction target, the system trains on the three most recent historical seasons and predicts the subsequent season’s outbreak patterns. The 2024–2025 season data was completely withheld during development to provide an unbiased assessment of real-world performance.

Performance metrics

We categorized system performance using temporal alignment between predicted and observed outbreak events:

  • Early Warnings: alarms triggered 1–6 weeks before confirmed outbreak events

  • Synchronous Warnings: alarms triggered during the same week as confirmed outbreak events

  • Late Warnings: alarms triggered up to 3 weeks after confirmed outbreak events

  • Early-Late Warnings: alerts issued after the true outbreak onset but prior to the observed epidemic peak, such that the system still provides actionable advance notice before peak disease activity. Early-late warnings are counted as successful peak detections but not as successful onset detections

  • Soft Warnings: Sub-threshold alert signals in which at least 75% of the required proxy detectors activate within a six-week window following outbreak onset, but the aggregated voting signal does not exceed the calibrated alarm threshold. Soft warnings indicate elevated outbreak risk without triggering a formal alert

  • Missed Outbreaks: confirmed outbreak events with no corresponding alarm within the evaluation window

We distinguished between two types of false alarms to better understand system behavior:

  • True False Alarms: alarms not followed by any detectable increase in disease activity within 6 weeks

  • Warning Signals: alarms not followed by confirmed outbreaks but associated with measurable increases in disease activity, indicating elevated transmission risk

Performance evaluation emphasized positive predictive value (PPV) to quantify alarm reliability and detection sensitivity to assess outbreak capture rates. This dual focus ensured the system balanced early warning capabilities with practical utility for public health decision-making. Rather than traditional ROC curves, our evaluation focuses on aggregated detection rates and positive predictive values. This choice reflects our system’s architecture: dynamically retrained, location-specific classifiers that update as new data becomes available, rather than a single static global model. ROC curves for each classifier (with multiple models per location and time period) would not provide an interpretable summary of overall system performance. Our reported metrics—detection sensitivity, PPV, and lead times—directly quantify the system’s operational utility for public health decision-making.

Reporting summary

Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.

Supplementary information

Reporting Summary (2.2MB, pdf)

Source data

Source Data 1 (10.2KB, xlsx)
Source Data 2 (10.5KB, xlsx)
Source Data 3 (43.2KB, xlsx)
Source Data 4 (39.1KB, xlsx)

Acknowledgements

We thank Mr. Xi Chen for his valuable contributions to the implementation of the Granger causality analysis and the development of corresponding visualizations. This manuscript was made possible in part by cooperative agreement CDC-RFA-FT-23-0069 from the CDC’s Center for Forecasting and Outbreak Analytics (to R.G.G., L.C., G.D., and M.S.). Its contents are solely the responsibility of the authors and do not necessarily represent the official views of the Centers for Disease Control and Prevention. Source data are provided with this paper.

Author contributions

R.G.G., L.C., and M.S. designed the study; R.G.G. and L.C. performed the research; R.G.G. developed and implemented the analytic tools; R.G.G., L.C., and M.S. analyzed the data; A.G.M. contributed to data interpretation and clinical contextualization; G.D. contributed to epidemiological analysis and interpretation; S.Y. contributed to data interpretation and statistical analysis; R.G.G. and M.S. wrote the first draft of the manuscript; all authors reviewed, contributed to, and approved the final version of the manuscript; M.S. supervised the research.

Peer review

Peer review information

Nature Communications thanks the anonymous reviewers for their contribution to the peer review of this work.

Data availability

All data used in this study are publicly available. Weekly influenza-like illness (ILI) surveillance data were obtained from the CDC FluView Interactive dashboard (https://gis.cdc.gov/grasp/fluview/fluportaldashboard.html). Respiratory syncytial virus (RSV) emergency department visit data were obtained from the CDC National Syndromic Surveillance Program (https://data.cdc.gov). Influenza hospitalization data were obtained from the CDC FluSight challenge (https://data.cdc.gov). Google Trends digital behavioral signal data were obtained through the Google Trends API (https://trends.google.com). Code and sample data to reproduce the results presented in this article are available at https://github.com/MIGHTE-lab/Respiratory-EWS-Sample. Source data are provided with this paper.

Code availability

A sample of data and code to reproduce results in this article is available at https://github.com/MIGHTE-lab/Respiratory-EWS-Sample and archived on Zenodo at 10.5281/zenodo.1937148252.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Supplementary information

The online version contains supplementary material available at 10.1038/s41467-026-72655-7.

References

  • 1.Organization, W.H.: Global epidemiological surveillance standards for influenza: Case definitions. Technical report, World Health Organization. https://cdn.who.int/media/docs/default-source/influenza/who_ili_sari_case_definitions_2014.pdf (2014).
  • 2.Centers for Disease Control and Prevention: Preliminary Estimated Flu Burden: 2024-2025. U.S. Department of Health & Human Services. https://www.cdc.gov/flu-burden/php/data-vis/2024-2025.html (2024).
  • 3.Centers for Disease Control and Prevention: RSV Disease Burden Estimates. U.S. Department of Health & Human Services. https://www.cdc.gov/rsv/php/surveillance/burden-estimates.html (2024).
  • 4.Centers for Disease Control and Prevention: COVID-19 Disease Burden Estimates. U.S. Department of Health & Human Services. https://www.cdc.gov/covid/php/surveillance/burden-estimates.html (2024).
  • 5.Centers for Disease Control and Prevention: About Flu Forecasting. https://www.cdc.gov/flu-forecasting/about/index.html (2024).
  • 6.Reich, N. G. et al. Accuracy of real-time multi-model ensemble forecasts for seasonal influenza in the us. PLoS Comput. Biol.15, 1007486 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Reich, N. G. et al. A collaborative multiyear, multimodel assessment of seasonal influenza forecasting in the united states. Proc. Natl. Acad. Sci. USA116, 3146–3154 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Mathis, S. M. et al. Evaluation of flusight influenza forecasting in the 2021–22 and 2022–23 seasons with a new target laboratory-confirmed influenza hospitalizations. Nat. Commun.15, 6289 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Viboud, C. & Vespignani, A. The future of influenza forecasts. Proc. Natl. Acad. Sci. USA116, 2802–2804 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Reich, N. G. et al. A collaborative multiyear, multimodel assessment of seasonal influenza forecasting in the united states. PLOS Comput. Biol.17, 1008994 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Kim, T. H., Chinthaginjala, R., Srinivasulu, A., Tera, S. P. & Rab, S. O. Covid-19 health data prediction: a critical evaluation of cnn-based approaches. Sci. Rep.15, 9121 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Dishar, H. K. & Muhammed, L. A. A review of the overfitting problem in convolution neural network and remedy approaches. J. Al-Qadisiyah Comput. Sci. Math.15, 155 (2023). [Google Scholar]
  • 13.Dewey, G., Meyer, A.G., Garrido Garcia, R. & Santillana, M. Uncovering the post-pandemic timing of influenza, rsv, and covid-19 driving seasonal influenza-like illness in the united states. medRxiv10.1101/2025.08.21.25333432 (2025). [DOI] [PMC free article] [PubMed]
  • 14.McGough, S. F., Brownstein, J. S., Hawkins, J. B. & Santillana, M. Forecasting zika incidence in the 2016 latin america outbreak combining traditional disease surveillance with search, social media, and news report data. PLoS Negl. Trop. Dis.11, 0005295 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Santillana, M. et al. Combining search, social media, and traditional data sources to improve influenza surveillance. PLoS Comput. Biol.11, 1004513 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Dugas, A. F. et al. Influenza forecasting with google flu trends. PloS one8, 56176 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Lee, K., Agrawal, A., Choudhary, A. Forecasting influenza levels using real-time social media streams. In: 2017 IEEE International Conference on Healthcare Informatics (ICHI), pp. 409–414 (IEEE, 2017).
  • 18.Aiken, E. L. et al. Real-time estimation of disease activity in emerging outbreaks using internet search information. PLoS Comput. Biol.16, 1008117 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Lu, F. S. et al. Accurate influenza monitoring and forecasting using novel internet data streams: a case study in the boston metropolis. JMIR Public Health Surveill.4, 8950 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Lu, F. S. et al. Estimating the cumulative incidence of covid-19 in the united states using influenza surveillance, virologic testing, and mortality data: Four complementary approaches. PLOS Comput. Biol.17, 1008994 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Yang, S., Santillana, M. & Kou, S. C. Accurate estimation of influenza epidemics using google search data via argo. Proc. Natl. Acad. Sci. USA112, 14473–14478 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Yuan, Q. et al. Monitoring influenza epidemics in china with search query from baidu. PloS One8, 64323 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Althouse, B. M., Ng, Y. Y. & Cummings, D. A. Prediction of dengue incidence using search query surveillance. PLoS Negl. Trop. Dis.5, 1258 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Ginsberg, J. et al. Detecting influenza epidemics using search engine query data. Nature457, 1012–1014 (2009). [DOI] [PubMed] [Google Scholar]
  • 25.Lu, F. S., Hattab, M. W., Clemente, C. L., Biggerstaff, M. & Santillana, M. Improved state-level influenza nowcasting in the united states leveraging internet-based data and network approaches. Nat. Commun.10, 147 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Gluskin, R. T., Johansson, M. A., Santillana, M. & Brownstein, J. S. Evaluation of internet-based dengue query data: Google dengue trends. PLoS Negl. Trop. Dis.8, 2713 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Santillana, M., Zhang, D. W., Althouse, B. M. & Ayers, J. W. What can digital disease detection learn from (an external revision to) google flu trends? Am. J. Prevent. Med.47, 341–347 (2014). [DOI] [PubMed] [Google Scholar]
  • 28.Lazer, D. et al. The parable of google flu: traps in big data analysis. Science343, 1203–1205 (2014). [DOI] [PubMed] [Google Scholar]
  • 29.Kogan, N. E. et al. An early warning approach to monitor covid-19 activity with multiple digital traces in near real time. Sci. Adv.7, 6989 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Stolerman, L. M. et al. Using digital traces to build prospective and real-time county-level early warning systems to anticipate covid-19 outbreaks in the united states. Sci. Adv.9, 0199 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Nash, R. K., Nouvellet, P. & Cori, A. Real-time estimation of the epidemic reproduction number: Scoping review of the applications and challenges. PLOS Digital Health1, 0000052 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Gostic, K. M. et al. Practical considerations for measuring the effective reproductive number, r t. PLoS Comput. Biol.16, 1008409 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Gallien, Y. et al. Using the near real-time effective reproduction number rt as an early-warning tool for seasonal bronchiolitis and influenza-like illness epidemics. Am. J. Epidemiol.194, 1332–1340 (2024). [DOI] [PubMed]
  • 34.Zou, B., Lampos, V. & Cox, I. Transfer learning for unsupervised influenza-like illness models from online search data. In Proc. World Wide Web Conference 2505–2516 (ACM, 2019).
  • 35.Lampos, V. et al. Tracking covid-19 using online search. NPJ Digital Med.4, 17 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Virtanen, P. et al. Scipy 1.0: fundamental algorithms for scientific computing in Python. Nat. Methods17, 261–272 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Palshikar, G. Simple algorithms for peak detection in time-series. In Proc. 1st Int. Conf. Advanced Data Analysis, Business Analytics and Intelligence, Vol. 122 (IIMA, 2009).
  • 38.Meyer, A.G., Lu, F., Clemente, L. & Santillana, M. A prospective real-time transfer learning approach to estimate influenza hospitalizations with limited data. Epidemics50, 100816 (2025). [DOI] [PMC free article] [PubMed]
  • 39.Hanage, W. P. & Schaffner, W. Burden of acute respiratory infections caused by influenza virus, respiratory syncytial virus, and sars-cov-2 with consideration of older adults: A narrative review. Infect. Dis. Ther.14, 5–37 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Moline, H.L.: Early estimate of nirsevimab effectiveness for prevention of respiratory syncytial virus–associated hospitalization among infants entering their first respiratory syncytial virus season—-new vaccine surveillance network, october 2023–february 2024. MMWR. Morbidity and mortality weekly report 73 (2024). [DOI] [PMC free article] [PubMed]
  • 41.Xu, H. et al. Estimated effectiveness of nirsevimab against respiratory syncytial virus. JAMA Netw. open8, 250380–250380 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.World Health Organization: Epidemic intelligence from open sources (EIOS) (2026). https://www.who.int/initiatives/eios.
  • 43.Mawudeku, A., Blench, M.: Global public health intelligence network (gphin). In: Proceedings of Machine Translation Summit X: Invited Papers (2005).
  • 44.Brownstein, J. S., Freifeld, C. C. & Madoff, L. C. Digital disease detection—harnessing the web for public health surveillance. N. Engl. J. Med.360, 2153 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Cramer, E. Y. et al. Evaluation of individual and ensemble probabilistic forecasts of covid-19 mortality in the united states. Proc. Natl. Acad. Sci. USA119, 2113561119 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Santillana, M., Nsoesie, E. O., Mekaru, S. R., Scales, D. & Brownstein, J. S. Using clinicians’ search query data to monitor influenza epidemics. Clin. Infect. Dis.59, 1446–1450 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Djorno, C., Santillana, M. & Yang, S. Restoring the forecasting power of Google Trends with statistical preprocessing. Int. J. Forecast. 10.1016/j.ijforecast.2026.03.001 (2026).
  • 48.Centers for Disease Control and Prevention: FluView: Influenza Surveillance Dashboard. https://gis.cdc.gov/grasp/fluview/fluportaldashboard.html (2025).
  • 49.Centers for Disease Control and Prevention: RSV-NET - Respiratory Syncytial Virus Infection (RSV). https://www.cdc.gov/rsv/php/surveillance/rsv-net.html (2025).
  • 50.Centers for Disease Control and Prevention: Weekly U.S. Influenza Surveillance Report. https://www.cdc.gov/flu/weekly/weeklyarchives2023-2024/week33.htm (2023).
  • 51.Benjamini, Y. & Hochberg, Y. Controlling the false discovery rate: a practical and powerful approach to multiple testing. J. R. Stat. Soc.57, 289–300 (1995). [Google Scholar]
  • 52.Garrido-Garcia, R. et al. A real-time early warning system to anticipate respiratory disease outbreaks using transfer learning. Zenodo10.5281/zenodo.19371482 (2025). [DOI] [PMC free article] [PubMed]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Reporting Summary (2.2MB, pdf)
Source Data 1 (10.2KB, xlsx)
Source Data 2 (10.5KB, xlsx)
Source Data 3 (43.2KB, xlsx)
Source Data 4 (39.1KB, xlsx)

Data Availability Statement

All data used in this study are publicly available. Weekly influenza-like illness (ILI) surveillance data were obtained from the CDC FluView Interactive dashboard (https://gis.cdc.gov/grasp/fluview/fluportaldashboard.html). Respiratory syncytial virus (RSV) emergency department visit data were obtained from the CDC National Syndromic Surveillance Program (https://data.cdc.gov). Influenza hospitalization data were obtained from the CDC FluSight challenge (https://data.cdc.gov). Google Trends digital behavioral signal data were obtained through the Google Trends API (https://trends.google.com). Code and sample data to reproduce the results presented in this article are available at https://github.com/MIGHTE-lab/Respiratory-EWS-Sample. Source data are provided with this paper.

A sample of data and code to reproduce results in this article is available at https://github.com/MIGHTE-lab/Respiratory-EWS-Sample and archived on Zenodo at 10.5281/zenodo.1937148252.


Articles from Nature Communications are provided here courtesy of Nature Publishing Group

RESOURCES