seasonmap field guide
A field guide to the weather models
seasonmap carries several forecast systems because each covers a different range, resolution, or purpose. This page describes what each one is, the range it's built for, and where the AI models differ from the physics-based ones.
A model is not a forecast
A weather model is a simulation: given the atmosphere's measured state at one moment, it computes forward what physics (or a trained network) says happens next. Its output is one simulated future, unedited — no person has checked it, and every model has documented tendencies that push its answers in characteristic directions.
A forecast — the NWS point forecast served on this site — is the official product: forecasters at your local Weather Forecast Office weigh the models against each other and against local knowledge, and their gridded forecast is published as the National Digital Forecast Database (NDFD). The boundary is worth stating precisely: the NBM, though statistically calibrated against observations, is model guidance — an input the forecasters consult, not the official forecast itself. That is why the NWS chip is the default on the forecast page and every other chip, the blend included, is labeled model output.
Raw model output is still worth reading — that is most of what this site shows — but for what it is: a scenario. Use models for timing trends, for how the pattern is evolving run to run, and for the spread between models. Use the official forecast for decisions.
Physics-based models
| Model | Who / what | Best range |
|---|---|---|
| GFS | NOAA's global model, ~13 km, four runs a day out to 16 days | Full-globe coverage, refreshed every 6 hours |
| ECMWF ("Euro") | The European global model, ~9 km — the standard skill benchmark | Days 3–10, where its skill advantage is most consistent |
| HRRR | NOAA's 3-km CONUS model, new run every hour | Same-day, hour-by-hour: it resolves individual storms and wind shifts that global models only parameterize |
| NBM | The National Blend of Models — dozens of inputs, statistically calibrated to observations | Point forecasts where a single calibrated number is needed |
| GEFS | The GFS run 31 times from deliberately perturbed starting states | Uncertainty estimation — the mean smooths noise, the spread quantifies confidence |
| CFS | NOAA's coupled seasonal system — a single run reaches months out | Large-scale pattern only; at that range it gives regime probabilities on the 500 mb field, not point forecasts |
The deterministic/ensemble distinction matters at longer lead times. A deterministic run is one plausible evolution rendered in full detail; an ensemble is many evolutions rendered more coarsely. Past about day 4, a single deterministic run showing a specific storm is one member of a wider distribution, and the ensemble spread indicates how many other members agree with it.
AI-based models
Since 2023, machine-learning models trained on decades of reanalysis data have matched or exceeded physics-based models on several standard verification metrics, at a fraction of the compute cost (Bi et al. 2023; Lam et al. 2023). seasonmap carries three lineages:
- AIFS — ECMWF's own AI forecast system, run operationally alongside the physics-based IFS.
- AI-GFS — NOAA's experimental AI run (GraphCast lineage), initialized from GFS analyses.
- Pangu-Weather — the openly published model behind seasonmap's paused in-house CAESAR-OP program.
These models produce accurate large-scale patterns and jet placement at low compute cost, but they tend toward smoother fields than the physics models: fine-scale extremes are underrepresented, so a 985 mb low may be forecast at 992 mb in the correct location. In verification this shows up as strong MAE with less separation on the largest-error cases that RMSE weights.
In-house runs
Consensus is a weighted multi-model blend, with member weights set by lead time and categorical fields decided by weighted vote rather than averaged. It is scored on the public board alongside every other model.
CAESAR — the in-house AI program (Pangu-Weather rollouts with adaptive recalibration and a perturbed-member ensemble) — is paused pending dedicated compute; its verification history is retained and it returns to the board when it can run cleanly. Consensus, the in-house multi-model blend, remains live and scored.
Known tendencies — and how to check them
Every model has documented habits. The short version, for the systems carried here:
- GFS — a long-documented progressive tendency at medium range: mid-latitude troughs and storm systems often move through faster than observed. Reduced across upgrades, not eliminated.
- ECMWF (IFS) — the consistent leader on hemispheric verification scores (WMO Lead Centre rankings), which is an average statement, not a guarantee for any particular storm. At ~9 km its convection is still parameterized: individual thunderstorms are represented, not simulated.
- HRRR — convection-allowing: the storms on its maps are explicitly simulated cells. Read them as "storm mode and environment," never as "a storm will be at this pixel at this hour" — placement and initiation timing legitimately jump run to run, and with a run every hour you will see that jumpiness.
- NBM — statistically calibrated against observations, so its average point error is hard to beat. The cost of blending is the tails: record heat, localized heavy precipitation bands, and sharp gradients are systematically muted, because averaging many inputs trims exactly those extremes. Use the blend for the most likely number; look at individual models for how bad the reasonable worst case is.
- AI models (AIFS, AI-GFS, Pangu) — smooth fields and underrepresented extremes, as covered below: right location, damped amplitude.
- ICON / GEM — independent dynamical cores from DWD and ECCC. Their independent errors are the point: when they agree with GFS and ECMWF, confidence rises; when they split, the split is information.
These are prior expectations, not fixed laws — and this site is built to let you check them instead of trusting them. The scoreboard verifies every model here against the same RTMA analysis on the same schedule, including a bias column: whether a model has been running warm or cold, and by how much, over the window you select. When a tendency claim and the live verification disagree, believe the verification.
Ensembles: forecasting the uncertainty
A single model run cannot tell you how confident to be — it renders one future in full detail, right or wrong. An ensemble reruns the same model from slightly perturbed starting states (and often perturbed physics), turning the forecast into a distribution. Two rules for reading one:
- The mean is smooth by construction. Averaging members mutes extremes the same way the NBM's blending does — an ensemble mean will almost never show the deepest plausible low or the heaviest plausible band. Its skill advantage over a single run at longer leads is a statistical property of the averaging (Leith 1974), not evidence that the smooth solution is the likely one.
- The spread is the message. Small spread = the members agree; large spread = genuinely multiple futures. A deterministic run showing a major storm at day 7 with wide ensemble spread is one member of a distribution, not a prediction.
On this site: the GEFS layers publish the 31-member mean and spread, plus individual-member products (spaghetti, probabilities, plumes); EPS members stream the same way.
Choosing a model by task
Model choice depends on the forecast question. Hour-by-hour detail today: HRRR. A single calibrated point number: NBM. The day 4–8 pattern: Euro and the AI models, cross-checked against the GEFS spread. When models disagree, check the scoreboard: every model on this site, including the in-house runs, is verified against the same analysis on the same schedule (see how forecast verification works).
Bi, K., et al. (2023). Accurate medium-range global weather forecasting with 3D neural networks. Nature, 619, 533–538.
Lam, R., et al. (2023). Learning skillful medium-range global weather forecasting. Science, 382, 1416–1421.
Leith, C. E. (1974). Theoretical skill of Monte Carlo forecasts. Mon. Wea. Rev., 102, 409–418.
WMO Lead Centre for Deterministic NWP Verification — exchanged headline scores (hosted at ECMWF).
NWS Meteorological Development Laboratory — National Blend of Models documentation.