seasonmap field guide
A field guide to the weather models
Every model on seasonmap exists because it answers a different question. Here's what each one actually is, where it shines, where it lies to you — and what the new AI models change.
The physics workhorses
| Model | Who / what | Reach for it when |
|---|---|---|
| GFS | NOAA's global model, ~13 km, four runs a day out to 16 days | You want the full-globe, full-range picture, refreshed every 6 hours |
| ECMWF (“Euro”) | The European global model, ~9 km — the long-standing skill benchmark | The medium range, days 3–10, where its lead is most consistent |
| HRRR | NOAA's 3-km CONUS model, new run every hour | Today and tonight: individual storms, snow bands, wind shifts — it resolves the convection the globals only parameterize |
| NBM | The National Blend of Models — dozens of inputs, statistically calibrated to observations | You just want the most accurate number for a point forecast |
| GEFS | The GFS run 31 times from deliberately perturbed starting states | You want to know how confident to be — the mean smooths noise, the spread maps uncertainty |
The deterministic/ensemble split is the important one. A single deterministic run is one plausible future rendered in high detail; an ensemble is many futures rendered coarser. Past about day 4, a deterministic snowstorm at hour 168 is a hypothesis — the ensemble spread tells you whether it's a lonely one.
The AI generation
Since 2023, machine-learning models trained on decades of reanalysis have matched or beaten the physics models on many headline scores — while running in minutes instead of supercomputer-hours. seasonmap carries three lineages:
- AIFS — ECMWF's own AI forecast system, run operationally alongside the physics IFS.
- AI-GFS — NOAA's experimental AI run (GraphCast lineage), initialized from GFS analyses.
- Pangu-Weather — the openly published model that powers our in-house CAESAR-OP runs.
Their character is distinctive: excellent large-scale patterns and jet placement, cheap to run, but smooth — they hedge fine-scale extremes, so a 985 mb bomb cyclone may verify as 992 mb of the right storm in the right place. Great for the pattern, cautious on the peaks. That trade shows up in verification: strong MAE, less separation on the events RMSE punishes.
The in-house runs
Consensus is our weighted multi-model blend — member weights set by lead time, categorical fields decided by weighted vote rather than smearing averages. Blends win on average by cancelling independent errors; ours is scored publicly like everything else.
CAESAR-OP is our operational AI rollout: a Pangu-Weather integration we run ourselves from ECMWF open-data initial conditions, published with full soundings, adaptively bias-corrected against our own verification history, and verified on the same board as the majors.
CAESAR-ENS is the ensemble arm: the same rollouts run from an array of perturbed initial states — published as an ensemble mean and spread, exactly like the GEFS layers, so you can see where our own model is confident and where it isn't.
So which do you trust?
Wrong question — the right one is for which job. Hour-by-hour today: HRRR. The cleanest point number: NBM. The day 4–8 pattern: Euro and the AI models, cross-checked against the GEFS spread. And when models disagree, don't average vibes: check the scoreboard — we verify every model on this site, including our own, against the same analysis, every hour.
Compare them live on the map →