seasonmap

seasonmap field guide

A field guide to the weather models

seasonmap carries several forecast systems because each covers a different range, resolution, or purpose. This page describes what each one is, the range it's built for, and where the AI models differ from the physics-based ones.

A model is not a forecast

A weather model is a simulation: given the atmosphere's measured state at one moment, it computes forward what physics (or a trained network) says happens next. Its output is one simulated future, unedited — no person has checked it, and every model has documented tendencies that push its answers in characteristic directions.

A forecast — the NWS point forecast served on this site — is the official product: forecasters at your local Weather Forecast Office weigh the models against each other and against local knowledge, and their gridded forecast is published as the National Digital Forecast Database (NDFD). The boundary is worth stating precisely: the NBM, though statistically calibrated against observations, is model guidance — an input the forecasters consult, not the official forecast itself. That is why the NWS chip is the default on the forecast page and every other chip, the blend included, is labeled model output.

Raw model output is still worth reading — that is most of what this site shows — but for what it is: a scenario. Use models for timing trends, for how the pattern is evolving run to run, and for the spread between models. Use the official forecast for decisions.

Physics-based models

ModelWho / whatBest range
GFSNOAA's global model, ~13 km, four runs a day out to 16 daysFull-globe coverage, refreshed every 6 hours
ECMWF ("Euro")The European global model, ~9 km — the standard skill benchmarkDays 3–10, where its skill advantage is most consistent
HRRRNOAA's 3-km CONUS model, new run every hourSame-day, hour-by-hour: it resolves individual storms and wind shifts that global models only parameterize
NBMThe National Blend of Models — dozens of inputs, statistically calibrated to observationsPoint forecasts where a single calibrated number is needed
GEFSThe GFS run 31 times from deliberately perturbed starting statesUncertainty estimation — the mean smooths noise, the spread quantifies confidence
CFSNOAA's coupled seasonal system — a single run reaches months outLarge-scale pattern only; at that range it gives regime probabilities on the 500 mb field, not point forecasts

The deterministic/ensemble distinction matters at longer lead times. A deterministic run is one plausible evolution rendered in full detail; an ensemble is many evolutions rendered more coarsely. Past about day 4, a single deterministic run showing a specific storm is one member of a wider distribution, and the ensemble spread indicates how many other members agree with it.

AI-based models

Since 2023, machine-learning models trained on decades of reanalysis data have matched or exceeded physics-based models on several standard verification metrics, at a fraction of the compute cost (Bi et al. 2023; Lam et al. 2023). seasonmap carries three lineages:

These models produce accurate large-scale patterns and jet placement at low compute cost, but they tend toward smoother fields than the physics models: fine-scale extremes are underrepresented, so a 985 mb low may be forecast at 992 mb in the correct location. In verification this shows up as strong MAE with less separation on the largest-error cases that RMSE weights.

In-house runs

Consensus is a weighted multi-model blend, with member weights set by lead time and categorical fields decided by weighted vote rather than averaged. It is scored on the public board alongside every other model.

CAESAR — the in-house AI program (Pangu-Weather rollouts with adaptive recalibration and a perturbed-member ensemble) — is paused pending dedicated compute; its verification history is retained and it returns to the board when it can run cleanly. Consensus, the in-house multi-model blend, remains live and scored.

Known tendencies — and how to check them

Every model has documented habits. The short version, for the systems carried here:

These are prior expectations, not fixed laws — and this site is built to let you check them instead of trusting them. The scoreboard verifies every model here against the same RTMA analysis on the same schedule, including a bias column: whether a model has been running warm or cold, and by how much, over the window you select. When a tendency claim and the live verification disagree, believe the verification.

Ensembles: forecasting the uncertainty

A single model run cannot tell you how confident to be — it renders one future in full detail, right or wrong. An ensemble reruns the same model from slightly perturbed starting states (and often perturbed physics), turning the forecast into a distribution. Two rules for reading one:

On this site: the GEFS layers publish the 31-member mean and spread, plus individual-member products (spaghetti, probabilities, plumes); EPS members stream the same way.

Choosing a model by task

Model choice depends on the forecast question. Hour-by-hour detail today: HRRR. A single calibrated point number: NBM. The day 4–8 pattern: Euro and the AI models, cross-checked against the GEFS spread. When models disagree, check the scoreboard: every model on this site, including the in-house runs, is verified against the same analysis on the same schedule (see how forecast verification works).

Sources
Bi, K., et al. (2023). Accurate medium-range global weather forecasting with 3D neural networks. Nature, 619, 533–538.
Lam, R., et al. (2023). Learning skillful medium-range global weather forecasting. Science, 382, 1416–1421.
Leith, C. E. (1974). Theoretical skill of Monte Carlo forecasts. Mon. Wea. Rev., 102, 409–418.
WMO Lead Centre for Deterministic NWP Verification — exchanged headline scores (hosted at ECMWF).
NWS Meteorological Development Laboratory — National Blend of Models documentation.
Compare them live on the map →