seasonmap field guide
How forecast verification works
"Which model is best?" is an empirical question, and the weather community has spent seventy years building the machinery to answer it. Here is what the scores on a verification board actually measure — and the two traps that make scoreboards lie.
Step one: agree on what "truth" is
You can't score a forecast without an answer key. Station observations feel like the obvious truth, but they're points — comparing a 13-km model grid to a thermometer behind an airport hangar mixes up model error with representativeness error. So operational verification usually scores forecasts against an analysis: a quality-controlled best estimate of what actually happened, on a grid, everywhere.
seasonmap verifies 2-m temperature against NOAA's RTMA (Real-Time Mesoscale Analysis), which blends thousands of surface observations onto a ~2.5-km grid every hour. Every model — including our own — is bilinearly sampled to the same grid and scored against the same hours. Same answer key for everyone, or the comparison is meaningless.
The three scores worth knowing
| Score | What it is | What it tells you |
|---|---|---|
| MAE | mean absolute error — the average miss, in °F | Typical day-to-day accuracy. The single most honest headline number. |
| RMSE | root-mean-square error — misses squared, then averaged | Punishes big busts. A model with fine MAE but bloated RMSE nails quiet days and whiffs the events that matter. |
| Bias | mean signed error | Systematic lean. +1.5°F bias means the model runs warm as a habit — correctable if you know it, dangerous if you don't. |
Read them together. MAE is the headline, RMSE-minus-MAE is the "bust tax," and bias is the correction you can apply for free.
Lead time changes everything
A 12-hour forecast and a 5-day forecast are different sports. Skill decays roughly exponentially with lead — small initial-condition errors double every couple of days, which is the practical face of chaos theory. So an honest board buckets by lead: day 1, 24–48 h, beyond 48 h. A model can win the short range and fade at range (common for high-resolution models like the HRRR, built for the first day or two) while a global model ages more gracefully.
Why the blend usually wins — and why that's not cheating
On most temperature boards, statistically calibrated blends like NOAA's NBM sit on top, ahead of any single raw model. Averaging across models cancels their independent errors; calibration then irons out each one's known biases. It's the same reason ensemble means beat single runs. Raw model scores are still the interesting ones for understanding why a forecast went wrong — but if you just want the number for tomorrow, the blend is usually it.
The two traps
Sample size. A ranking built on a handful of scored hours is noise wearing a suit. One quiet ridge day can flatter a mediocre model; one busted front can slander a good one. Watch the n column, and trust separations only when they persist across weeks and regimes. This is also why period filters matter — a model that handled last week's heat dome may rank differently over the storm before it.
The single-number fallacy. "Best at 2-m temperature over CONUS this month" is not "best at everything." Verification is per-variable, per-region, per-regime. The same model can lead at temperature and trail at precipitation placement.