Google Says Its AI-Built Flu Forecast Led 39 Models in CDC Season Evaluation
The forecasts predict weekly hospital demand, not individual diagnoses.
Loading page…
The forecasts predict weekly hospital demand, not individual diagnoses.
Listen to this story
Google’s forecasting model ranked first among 39 eligible entries in the CDC’s retrospective evaluation of U.S. flu-hospitalization forecasts for the 2025–26 season. Google developed it with ERA, a system that uses an LLM and tree search to build scientific software against a defined quality measure. The result offers a real-world test of AI-assisted software development in public health, but Google disclosed no winning margin, and one season’s ranking does not establish performance across diseases or future seasons.
FluSight accepts weekly forecasts from October through May, covering U.S. hospital admissions for the current week and three weeks ahead.
The CDC’s season-end comparison measures forecasts against observed admissions; it is separate from the agency’s weekly combined forecast for medical-service planning.
A Nature paper published May 19 describes ERA as combining a large language model with tree search to optimize a specified quality measure.
Google says an AI-built forecasting model came closest to the flu hospitalizations that actually occurred across the United States during the 2025–26 season. In a September 30 announcement, the company said its model led 39 eligible entries in the CDC’s season-end FluSight evaluation. The task has a concrete public-health purpose: the CDC uses combined forecasts to communicate anticipated demand for medical services at the state level.
Google says it developed the flu forecasts with Empirical Research Assistance, or ERA. The tool generates optimization algorithms for scientific fields. AI helped build the forecasting model; it did not turn a hospitalization forecast into a conversational answer.
A paper published in Nature on May 19 describes ERA as a system for creating scientific software against a specified quality measure. It combines a large language model with tree search, a method for exploring possible solutions. The system works to improve that measure as it searches, giving the software-building process a defined target rather than simply asking AI to produce code.
The researchers say ERA explores and integrates complex research ideas from external sources. They evaluated it on scientific-software tasks, including public health and time-series forecasting, or predicting how a quantity changes over time.
FluSight brings together weekly submissions from government, industry and academic teams from October through May. Each submission predicts U.S. hospital admissions for the current week and three weeks into the future. The current-week estimate and the forward-looking predictions belong to the same forecasting exercise; the result is about admission counts, not diagnosing whether a particular person has flu.
Google says the CDC published its end-of-season analysis this week, comparing eligible models with the hospital admissions observed during the season. That is a retrospective assessment of forecasts submitted through the season, distinct from the CDC’s weekly use of a combined forecast to describe expected medical-service demand.
Google says its model ranked first among 39 eligible models in the CDC’s 2025–26 FluSight season analysis.
The ranking remains Google’s account of the CDC evaluation. Readers can identify the claimed winner from that disclosure, but cannot tell whether it won narrowly or by a substantial margin.
Google interprets the performance as support for combining AI with human ingenuity to improve disease forecasting worldwide. That is a broader expectation than the result it describes: one model’s performance for U.S. flu-related hospital admissions during one season.
Loading discussion...
Join the conversation
Explain what would make you trust it—or want more evidence.
Be the first to share a perspective or an experience.
Reader comments
Newest comments first. Replies stay oldest first.