IA 360
Medio Ambiente

NOAA adds AI models to global forecasting without removing physics

AIGFS, AIGEFS, and the hybrid HGEFS ensemble entered operations as new guidance for forecasters. Comparing speed, skill, uncertainty, and event-specific failures explains the advance better than a single accuracy figure.

5 min read AI-generated Leer en español
NOAA adds AI models to global forecasting without removing physics

On December 17, 2025, the U.S. National Oceanic and Atmospheric Administration put three artificial-intelligence-supported global forecasting systems into operation: AIGFS, AIGEFS, and HGEFS. A National Weather Service change notice set the start at the 1200 UTC cycle and published the data locations. AI was entering production as forecasting guidance; it was not replacing forecasters or retiring physics-based models.

The story teaches a method for reading any prediction announcement: keep cost, latency, skill, calibration, and decision value separate. A model can run much faster, improve a cyclone's track, and degrade its intensity forecast at the same time. Summarizing that as “more accurate” removes the information a user needs.

Three products, three questions

AIGFS is deterministic guidance. It starts from an atmospheric analysis and produces one evolution. NOAA built it on Google DeepMind's GraphCast and fine-tuned it with analyses from its Global Data Assimilation System. An NCEP technical note describes training on GDAS data and the process used to validate and compare the system with the operational Global Forecast System.

AIGEFS is an AI ensemble. Instead of one trajectory, it runs 31 members to represent variations in initial conditions and the model. The value of an ensemble is not to vote for a winning map. It is to approximate a distribution: how many members support an outcome, how widely they spread, and how the signal changes with lead time.

HGEFS combines those 31 AI members with 31 from the physics-based GEFS. The operational document defines it as a 62-member hybrid ensemble. The design recognizes that two error families can complement one another. If every member shares similar assumptions and data, adding more does not guarantee useful diversity; mixing approaches can keep one failure class from dominating the suite.

Operational does not mean infallible

At a weather agency, “operational” has a specific meaning: the system runs on a schedule with defined formats and distribution, and its products enter working flows. The notice specifies GRIB2 files, a quarter-degree grid, and forecast horizons. The word does not certify that every variable beats the previous system or turn an output into an automatic public warning.

Forecasters combine this guidance with observations, radar, satellites, other models, and local knowledge. A global prediction does not resolve an urban thunderstorm, a mountain effect, or river flow by itself. Nor does it make an evacuation decision. Interpretation, regional products, thresholds, procedures, and accountable officials stand between model and action.

Operational use does raise the monitoring standard. An experiment can focus on selected periods; a service must watch degradation, input failures, delays, seasonal changes, and out-of-distribution events. Versioning the model and preserving parallel comparisons helps distinguish a genuine upgrade from an easier evaluation period.

Speed and compute: the efficiency axis

NOAA's announcement says a 16-day AIGFS forecast finishes in about 40 minutes and uses 0.3% of the computing resources consumed by operational GFS. For AIGEFS, the agency reports 9% of GEFS compute. These are compute comparisons for particular configurations, not percentages of agency-wide energy use or final cost.

Lower latency can create value even if skill is unchanged. A center can receive guidance sooner, run more variants inside a fixed window, or reserve compute for higher-resolution models. A fair savings comparison still needs a boundary: training or inference alone, preprocessing, postprocessing, storage, hardware, and resolution. The announcement compares operational runs; it does not provide full life-cycle accounting.

Speed must not be confused with useful reach. Producing 16 days in 40 minutes does not mean day sixteen has the quality of day one. “Forecast horizon produced” and “horizon with useful skill” belong in different columns.

Skill always has a variable and lead time

NOAA reported that AIGFS improved large-scale synoptic patterns and long-lead tropical-cyclone tracks. The same document acknowledged degraded cyclone intensity forecasts in the initial version. There is no contradiction: track and intensity are different targets. A storm can be positioned well while its wind or pressure is wrong.

A superiority claim therefore needs at least five qualifiers: variable, region, lead time, metric, and baseline. Track error can be measured as distance; intensity can use maximum wind or pressure; extreme rainfall requires thresholds and location. Compressing those into a general score can hide the failure that matters for emergency management.

NOAA's announcement contains two formulations about AIGEFS that require restraint. Its summary attributes better performance and an additional 18 to 24 hours of forecast skill to early results; the detailed section calls performance comparable to GEFS. Without metric-level tables, this article does not turn that interval into a universal gain. The careful reading is that NOAA saw improvements in evaluations while describing overall performance as comparable and continuing to improve ensemble spread.

An ensemble must be calibrated, not merely average correctly

Suppose 20 of 31 members place heavy rain over a watershed. The raw fraction is not automatically a calibrated probability. Interpretation requires many past cases: how often did the event occur when the ensemble issued a similar signal? If it occurred six times in ten, a 60% forecast is calibrated; if it occurred three, the ensemble is overconfident.

Spread matters too. A narrow ensemble can create false certainty if all members share a bias. A very wide one may contain the outcome but offer little discrimination. Reliability diagrams, rank information, and probabilistic scores reveal properties that the ensemble mean cannot.

HGEFS seeks diversity by joining AI and physics. NOAA says initial tests outperformed either component across major verification metrics. The statement should remain attached to those tests and metrics, not become a law. Operational monitoring will show which seasons, variables, and events benefit from the mixture and where one family weakens the other.

Using a forecast without obeying it

A decision starts by defining potential harm and the time required to act. A utility may stage crews at one wind threshold; an evacuation needs another threshold and additional evidence. The costs of a false alarm and inaction differ. A model informs that tradeoff, but it does not decide which risk society accepts.

Users should ask which model and version produced a map, at what time, from which initial observations, with what spread, which independent systems agree, and which known limitation affects the event. They should also retain the previous update: movement between cycles can be as informative as the newest value.

Evaluating progress over months requires scorecards by variable, region, lead time, and season, together with availability and latency. Failure cases should be published alongside averages. Trust grows when users can see when a system should not be the dominant guidance.

NOAA did not choose between physics and AI. It placed an AI deterministic model, an AI ensemble, and a hybrid combination beside its existing infrastructure. The transferable skill is to read progress on separate axes: how long it takes, what it costs, what it predicts better, where it fails, how it represents uncertainty, and who converts guidance into a decision.

This article was produced with artificial intelligence under human editorial oversight.

Share this article

This website uses cookies to improve the browsing experience. Cookie policy.

↑↓ navigate ↵ open esc close