seasonmap field guide
How forecast verification works
Verification scores each model's forecasts against what actually happened and reports the miss in standard units. This page describes the scores used on the public scoreboard — MAE, RMSE, and bias — what analysis they're checked against, and the two ways a ranking can mislead if read without its sample size and lead-time bucket.
The reference: analysis, not station data
Scoring a forecast requires an answer key. Station observations are points, and comparing a 13-km model grid directly to a thermometer behind an airport hangar mixes model error with representativeness error. Operational verification instead scores forecasts against an analysis: a quality-controlled estimate of the true state on a grid, everywhere, built by blending observations with a short model background (Jolliffe & Stephenson 2012).
seasonmap verifies 2-m temperature against NOAA's RTMA (Real-Time Mesoscale Analysis), which blends thousands of surface observations onto a ~2.5-km grid every hour. Every model, including the in-house runs, is bilinearly sampled to that same grid and scored against the same hours, so the comparison uses one answer key for every entry on the board.
Three scores
| Score | What it is | What it tells you |
|---|---|---|
| MAE | mean absolute error — the average miss, in °F | Typical day-to-day accuracy across the scored period. |
| RMSE | root-mean-square error — misses squared, then averaged, then square-rooted | Weights large errors more heavily than MAE. A model with low MAE but high RMSE is accurate on ordinary days and inaccurate on the days with the largest misses. |
| Bias | mean signed error | Systematic lean. A +1.5°F bias means the model runs warm on average, independent of its scatter. |
MAE gives the typical miss, RMSE minus MAE indicates how much of the error comes from a small number of large misses rather than a consistent small one, and bias identifies a correctable systematic offset (Murphy 1993).
Lead time
Forecast error grows with lead time as small initial-condition errors amplify, roughly doubling every one to three days depending on the flow regime. A board that reports one aggregate score across all lead times obscures this growth, so seasonmap buckets scores by lead: day 1, 24–48 h, beyond 48 h. A high-resolution model like the HRRR, built to resolve the first day or two, can lead its bucket and still trail a global model at longer range — the two are not competing for the same lead-time window.
Why blended products score well
On most temperature boards, statistically calibrated blends such as NOAA's NBM outscore any single raw model. Averaging across models cancels their independent errors, and calibration removes each input's known bias — the same mechanism by which an ensemble mean outperforms a single ensemble member. Raw single-model scores remain useful for diagnosing why a particular forecast missed; the blend is the better number for a single point estimate.
Sample size and scope
A ranking built on a small number of scored hours is not statistically reliable: one quiet period can flatter a mediocre model, and one missed event can penalize a good one disproportionately. Check the n (hours scored) on any board, and treat a separation between models as meaningful only once it persists across multiple weeks and weather regimes.
Verification is also specific to variable, region, and regime. A ranking for 2-m temperature over CONUS this month does not generalize to precipitation placement, a different region, or a different season; the same model can lead on one and trail on another.
Murphy, A. H. (1993). What is a good forecast? An essay on the nature of goodness in weather forecasting. Wea. Forecasting, 8, 281–293.
Jolliffe, I. T., & Stephenson, D. B. (2012). Forecast Verification: A Practitioner's Guide, 2nd ed. Wiley.
The synoptic standards: 500-mb height, MSLP, 850-mb temperature, 250-mb wind, PWAT
Surface temperature is what you feel, but the numbers forecasters and modeling centers quote against each other are upper-air: root-mean-square error of the 500-mb geopotential height field at days 3, 5, and 7, and mean-sea-level pressure error alongside it. Z500 measures whether a model has the large-scale pattern — the troughs and ridges that steer everything else — in the right place. A model can win day-1 surface temperature on calibration alone; it cannot fake the day-5 pattern.
Two more upper-air fields sit beside them on the board: 850-mb temperature, the standard low-level thermal field (airmass boundaries, warm-air advection, the rain–snow problem in winter), and 250-mb wind speed, the jet stream itself — jet-streak placement organizes surface cyclogenesis, so day-5 jet errors foreshadow day-6 surface errors. Precipitable water joins them — column moisture errors foreshadow QPF errors. All are scored as RMSE and bias; the anomaly correlation stays a Z500-only column because ACC requires a per-field climatology and we currently carry ERA5's for Z500 alone.
The surface board verifies more than temperature: a variable toggle scores 2-m dewpoint and 10-m wind speed against the same RTMA analysis, with the same day buckets, time windows, and issue-cycle splits. Dewpoint is the harder surface test — moisture errors compound into convection and heat-index errors — and wind verification carries an honest caveat: RTMA's wind analysis rests on a sparser observation set than its temperature field, so treat small ranking gaps there with more skepticism.
seasonmap scores every model that publishes these fields — including our in-house CAESAR runs — at fixed leads from 24 to 168 hours, latitude-weighted (a degree of longitude shrinks toward the pole; cosφ weighting keeps northern rows from dominating), against a common reference: the GFS analysis at each 00/06/12/18Z valid time. One honest caveat comes with that choice: an observation-only analysis with upper-air fields does not exist on open feeds, so GFS short-lead scores partially verify the model against itself — read its day-1 row as a floor rather than a victory. By day 3 and beyond, analysis differences are small relative to forecast error and the comparison is fair.
Z500 cells also carry the anomaly correlation coefficient (ACC) — the classic headline score. ACC asks a sharper question than RMSE: after removing the climatological pattern for that calendar day (here the ERA5 1990–2019 climatology, via WeatherBench2), how well do the forecast anomalies correlate with the analysis anomalies? A model can score decent RMSE by hugging climatology; ACC exposes that. By convention, ACC above 0.6 marks the edge of a synoptically useful forecast — watch where each model's day-5 and day-7 columns sit against that line.