seasonmap

seasonmap field guide

A field guide to the weather models

Every model on seasonmap exists because it answers a different question. Here's what each one actually is, where it shines, where it lies to you — and what the new AI models change.

The physics workhorses

ModelWho / whatReach for it when
GFSNOAA's global model, ~13 km, four runs a day out to 16 daysYou want the full-globe, full-range picture, refreshed every 6 hours
ECMWF (“Euro”)The European global model, ~9 km — the long-standing skill benchmarkThe medium range, days 3–10, where its lead is most consistent
HRRRNOAA's 3-km CONUS model, new run every hourToday and tonight: individual storms, snow bands, wind shifts — it resolves the convection the globals only parameterize
NBMThe National Blend of Models — dozens of inputs, statistically calibrated to observationsYou just want the most accurate number for a point forecast
GEFSThe GFS run 31 times from deliberately perturbed starting statesYou want to know how confident to be — the mean smooths noise, the spread maps uncertainty

The deterministic/ensemble split is the important one. A single deterministic run is one plausible future rendered in high detail; an ensemble is many futures rendered coarser. Past about day 4, a deterministic snowstorm at hour 168 is a hypothesis — the ensemble spread tells you whether it's a lonely one.

The AI generation

Since 2023, machine-learning models trained on decades of reanalysis have matched or beaten the physics models on many headline scores — while running in minutes instead of supercomputer-hours. seasonmap carries three lineages:

Their character is distinctive: excellent large-scale patterns and jet placement, cheap to run, but smooth — they hedge fine-scale extremes, so a 985 mb bomb cyclone may verify as 992 mb of the right storm in the right place. Great for the pattern, cautious on the peaks. That trade shows up in verification: strong MAE, less separation on the events RMSE punishes.

The in-house runs

Consensus is our weighted multi-model blend — member weights set by lead time, categorical fields decided by weighted vote rather than smearing averages. Blends win on average by cancelling independent errors; ours is scored publicly like everything else.

CAESAR-OP is our operational AI rollout: a Pangu-Weather integration we run ourselves from ECMWF open-data initial conditions, published with full soundings, adaptively bias-corrected against our own verification history, and verified on the same board as the majors.

CAESAR-ENS is the ensemble arm: the same rollouts run from an array of perturbed initial states — published as an ensemble mean and spread, exactly like the GEFS layers, so you can see where our own model is confident and where it isn't.

So which do you trust?

Wrong question — the right one is for which job. Hour-by-hour today: HRRR. The cleanest point number: NBM. The day 4–8 pattern: Euro and the AI models, cross-checked against the GEFS spread. And when models disagree, don't average vibes: check the scoreboard — we verify every model on this site, including our own, against the same analysis, every hour.

Compare them live on the map →