Daily verification · methodology v0.3.1
How far off was each weather model?
No group in this view yet has enough scored days to rank (n < 30). Every number below is still published, greyed out, with its sample size.
Scope: raw model output on the native 0.25° grid. This view scores the sampled daily maximum against the maximum of the four *observed* samples on the same day. Errors are computed in °C and shown in °F (methodology v0.3.1).
Leader, lead day 1
—
no ranked group with an interval yet
n < 30 everywhere
Persistence baseline
2.83°F
[2.48, 3.22]
yesterday's observation — the bar every model has to clear
Scored days
30
23 stations · 12 systems
in the last 30 days, lead day 1
Systems ranked
0 of 12 ranked yetn < 30
next update
below n = 30: published, greyed, unranked
Mean absolute error with 95 % confidence intervals
Lead day 1 · the last 30 days · 23 stations pooled · 12Z · bilinear · sampled daily maximum. Whiskers are moving-block bootstrap intervals; models whose whiskers overlap are not distinguishable at this sample size. Baselines are drawn in grey and never ranked.
Source: CastCheck 0.1.0, methodology v0.3.1 · ECMWF open data, NOAA/NCEP GFS, NOAA/CIRA AIWP, NWS ASOS · data through 2026-08-30 · CC BY 4.0. Hover a bar for its value.
Lead day 1target = init + 1 d
permalinkThe last 30 days, all 23 stations pooled, 12Z initialization, bilinear interpolation, sampled daily maximum against the maximum of the four *observed* samples on the same day. Lower MAE is better; bias is positive when the model is too warm. Every number links to its permanent page.
Skill is 1 − MAE ÷ MAE(persistence) computed on the days both have a value — the small print in that column is that denominator and the size of the intersection, so the number can be checked against the persistence row rather than contradicting it. The persistence row's own n is its whole record, marked all days. Skill, debiased (out-of-sample) repeats the score after removing a bias estimated on the 30 scored days before each day and applied forward, never on the day itself. An interval reads — when the bootstrap could not run: fewer than 28 scored days or fewer than 4 blocks. Hit-rate intervals are Wilson score intervals, not bootstraps. — in the rank column means fewer than 30 scored days, so the group is published but not ranked.
| Rank | Model | MAE °F | Bias °F | ±3 °F | Skill | Skill, debiased (out-of-sample) | n | vs leader |
|---|---|---|---|---|---|---|---|---|
| — | ECMWF AIFS Singleaifs_single | 2.04[1.98, 2.11] | −0.38[−0.49, −0.27] | 78% | +0.27[+0.18, +0.34]vs 2.80 (n=29) | +0.48n=14 | 2922.7 stns | |
| — | ECMWF IFS HRESifs_hres | 2.19[2.07, 2.31] | −0.52[−0.64, −0.40] | 75% | +0.22[+0.13, +0.28]vs 2.80 (n=29) | +0.34n=14 | 2922.7 stns | |
| — | Pangu-Weather (IFS init)pangu_ifs | 2.46— | −0.42— | 67% | ——vs 3.48 (n=2) | —n=0 | 219 stns | |
| — | FourCastNet v2 (GFS init)fourcastnet_gfs | 2.55— | −1.09— | 77% | ——vs 3.48 (n=2) | —n=0 | 219 stns | |
| — | FourCastNet v2 (IFS init)fourcastnet_ifs | 2.60— | −1.31— | 57% | ——vs 3.48 (n=2) | —n=0 | 219 stns | |
| — | Persistence (baseline)persistence · baseline | 2.83[2.48, 3.22] | +0.01[−0.22, 0.26] | 63% | — | — | 30all days | |
| — | Pangu-Weather (GFS init)pangu_gfs | 2.93— | −0.58— | 56% | ——vs 3.48 (n=2) | —n=0 | 219 stns | |
| — | GraphCast (IFS init)graphcast_ifs | 3.00[2.84, 3.14] | −2.53[−2.73, −2.32] | 54% | −0.08[−0.21, +0.05]vs 2.79 (n=28) | +0.41n=28 | 2822.6 stns | |
| — | NCEP GFSgfs | 3.19[2.97, 3.37] | +1.88[1.57, 2.20] | 55% | −0.14[−0.31, −0.03]vs 2.80 (n=29) | +0.30n=14 | 2922.7 stns | |
| — | Aurora (IFS init)aurora_ifs | 3.46— | −3.30— | 43% | ——vs 3.48 (n=2) | —n=0 | 219 stns | |
| — | GraphCast (GFS init)graphcast_gfs | 3.72— | −3.37— | 46% | ——vs 3.48 (n=2) | —n=2 | 219 stns | |
| — | Aurora (GFS init)aurora_gfs | 4.24— | −4.11— | 37% | ——vs 3.48 (n=2) | —n=0 | 219 stns |
+ model too warm − model too cold bias interval includes zero ★ lowest MAE = not distinguishable from the leader ▼ worse than the leader ▲ better than the leader — all three Holm-corrected within this table, over the family of comparisons against the leader; the uncorrected verdict for every pair is published in the pairwise table of each permanent link and in pairwise_latest.csv.gzn < 30 greyed and unranked CI is a 95 % moving-block bootstrap interval on the group's own days, and reads — when it could not be computed skill reads — when fewer than 10 days are common to the model and the baseline: the ratio of two means over a handful of shared days is not a number worth printing
Lead day 3target = init + 3 d
permalinkSame sampling as lead day 1. Skill is 1 − MAE ÷ MAE(persistence) on the days both have a value.
| Rank | Model | MAE °F | Bias °F | Skill | n | vs leader |
|---|---|---|---|---|---|---|
| — | ECMWF AIFS Singleaifs_single | 2.20— | −0.52— | +0.43—vs 3.84 (n=27) | 2722.6 stns | |
| — | ECMWF IFS HRESifs_hres | 2.41— | −0.19— | +0.37—vs 3.84 (n=27) | 2722.6 stns | |
| — | GraphCast (IFS init)graphcast_ifs | 3.22[2.97, 3.44] | −2.60[−2.89, −2.33] | +0.17[+0.05, +0.28]vs 3.89 (n=28) | 2822.9 stns | |
| — | NCEP GFSgfs | 3.50— | +1.97— | +0.09—vs 3.84 (n=27) | 2722.6 stns | |
| — | Persistence (baseline)persistence · baseline | 3.90[3.60, 4.19] | −0.01[−0.52, 0.59] | — | 30all days |
Lead day 5target = init + 5 d
permalinkSame sampling as lead day 1. Skill is 1 − MAE ÷ MAE(persistence) on the days both have a value.
| Rank | Model | MAE °F | Bias °F | Skill | n | vs leader |
|---|---|---|---|---|---|---|
| 1 | GraphCast (IFS init)graphcast_ifs | 3.54[3.22, 3.88] | −2.64[−2.91, −2.37] | +0.12[−0.07, +0.28]vs 4.03 (n=30) | 3022.6 stns | ★ lowest MAE in this view |
| — | ECMWF AIFS Singleaifs_single | 2.49— | −0.56— | +0.39—vs 4.05 (n=25) | 2522.6 stns | |
| — | ECMWF IFS HRESifs_hres | 2.85— | +0.08— | +0.30—vs 4.05 (n=25) | 2522.6 stns | |
| — | NCEP GFSgfs | 3.66— | +1.62— | +0.10—vs 4.05 (n=25) | 2522.6 stns | |
| — | Persistence (baseline)persistence · baseline | 4.03[3.60, 4.48] | +0.03[−0.68, 0.80] | — | 30all days |
Lead day 7target = init + 7 d
permalinkSame sampling as lead day 1. Skill is 1 − MAE ÷ MAE(persistence) on the days both have a value.
| Rank | Model | MAE °F | Bias °F | Skill | n | vs leader |
|---|---|---|---|---|---|---|
| 1 | GraphCast (IFS init)graphcast_ifs | 3.84[3.39, 4.34] | −2.78[−3.27, −2.32] | +0.11[−0.06, +0.25]vs 4.29 (n=30) | 3022.6 stns | ★ lowest MAE in this view |
| — | ECMWF AIFS Singleaifs_single | 2.86— | −1.09— | +0.33—vs 4.26 (n=23) | 2322.6 stns | |
| — | ECMWF IFS HRESifs_hres | 3.11— | −0.28— | +0.27—vs 4.26 (n=23) | 2322.6 stns | |
| — | NCEP GFSgfs | 4.05— | +1.54— | +0.05—vs 4.26 (n=23) | 2322.6 stns | |
| — | Persistence (baseline)persistence · baseline | 4.29[4.00, 4.58] | +0.04[−0.75, 0.80] | — | 30all days |
The same four samples as a daily maximum and minimum
Lead day 1, the last 30 days, 12Z, bilinear. Here the forecast's max/min of the four samples is scored against the observation's max/min of the same four instants. Like for like: whatever the four-sample definition misses, it misses on both sides. The comparison against the true NWS daily extremes is a different question and lives on the station and permanent-link pages.
Sampled daily maximumtmax_s
| Rank | Model | MAE °F | Bias °F | Skill | n | vs leader |
|---|---|---|---|---|---|---|
| — | ECMWF AIFS Singleaifs_single | 2.04[1.98, 2.11] | −0.38[−0.49, −0.27] | +0.27vs 2.80 (n=29) | 29 | |
| — | ECMWF IFS HRESifs_hres | 2.19[2.07, 2.31] | −0.52[−0.64, −0.40] | +0.22vs 2.80 (n=29) | 29 | |
| — | Pangu-Weather (IFS init)pangu_ifs | 2.46— | −0.42— | —vs 3.48 (n=2) | 2 | |
| — | FourCastNet v2 (GFS init)fourcastnet_gfs | 2.55— | −1.09— | —vs 3.48 (n=2) | 2 | |
| — | FourCastNet v2 (IFS init)fourcastnet_ifs | 2.60— | −1.31— | —vs 3.48 (n=2) | 2 | |
| — | Persistence (baseline)persistence | 2.83[2.48, 3.22] | +0.01[−0.22, 0.26] | — | 30all days | |
| — | Pangu-Weather (GFS init)pangu_gfs | 2.93— | −0.58— | —vs 3.48 (n=2) | 2 | |
| — | GraphCast (IFS init)graphcast_ifs | 3.00[2.84, 3.14] | −2.53[−2.73, −2.32] | −0.08vs 2.79 (n=28) | 28 | |
| — | NCEP GFSgfs | 3.19[2.97, 3.37] | +1.88[1.57, 2.20] | −0.14vs 2.80 (n=29) | 29 | |
| — | Aurora (IFS init)aurora_ifs | 3.46— | −3.30— | —vs 3.48 (n=2) | 2 | |
| — | GraphCast (GFS init)graphcast_gfs | 3.72— | −3.37— | —vs 3.48 (n=2) | 2 | |
| — | Aurora (GFS init)aurora_gfs | 4.24— | −4.11— | —vs 3.48 (n=2) | 2 |
Sampled daily minimumtmin_s
| Rank | Model | MAE °F | Bias °F | Skill | n | vs leader |
|---|---|---|---|---|---|---|
| — | FourCastNet v2 (IFS init)fourcastnet_ifs | 2.03— | −0.02— | —vs 2.54 (n=2) | 2 | |
| — | FourCastNet v2 (GFS init)fourcastnet_gfs | 2.11— | +0.33— | —vs 2.54 (n=2) | 2 | |
| — | Aurora (GFS init)aurora_gfs | 2.14— | −0.20— | —vs 2.54 (n=2) | 2 | |
| — | ECMWF IFS HRESifs_hres | 2.18[2.02, 2.35] | −0.55[−0.75, −0.35] | +0.07vs 2.34 (n=29) | 29 | |
| — | GraphCast (GFS init)graphcast_gfs | 2.19— | +0.18— | —vs 2.54 (n=2) | 2 | |
| — | Pangu-Weather (GFS init)pangu_gfs | 2.20— | +0.15— | —vs 2.54 (n=2) | 2 | |
| — | Aurora (IFS init)aurora_ifs | 2.22— | −0.85— | —vs 2.54 (n=2) | 2 | |
| — | GraphCast (IFS init)graphcast_ifs | 2.26[2.14, 2.39] | −1.21[−1.39, −1.03] | +0.04vs 2.34 (n=28) | 28 | |
| — | ECMWF AIFS Singleaifs_single | 2.28[2.09, 2.44] | −1.33[−1.54, −1.14] | +0.02vs 2.34 (n=29) | 29 | |
| — | Pangu-Weather (IFS init)pangu_ifs | 2.30— | −0.99— | —vs 2.54 (n=2) | 2 | |
| — | Persistence (baseline)persistence | 2.34[2.15, 2.51] | +0.02[−0.18, 0.21] | — | 30all days | |
| — | NCEP GFSgfs | 2.40[2.21, 2.61] | +0.17[−0.09, 0.45] | −0.03vs 2.34 (n=29) | 29 |
Every model × every lead day
MAE in °F with the bias and n underneath, the last 30 days, 12Z, bilinear, sampled daily maximum. The sparkline is the same model's MAE across lead days 1–9 on a shared vertical scale.
| Model | d1 | d2 | d3 | d4 | d5 | d6 | d7 | d8 | d9 | lead 1–9 |
|---|---|---|---|---|---|---|---|---|---|---|
| ECMWF AIFS Singleaifs_single | 2.04−0.38 · n=29 | 2.12−0.54 · n=28 | 2.20−0.52 · n=27 | 2.30−0.46 · n=26 | 2.49−0.56 · n=25 | 2.83−0.75 · n=24 | 2.86−1.09 · n=23 | 3.21−1.53 · n=22 | 3.73−2.05 · n=21 | |
| Aurora (GFS init)aurora_gfs | 4.24−4.11 · n=2 | 5.06−5.05 · n=1 | — | — | — | — | — | — | — | |
| Aurora (IFS init)aurora_ifs | 3.46−3.30 · n=2 | 3.81−3.50 · n=1 | — | — | — | — | — | — | — | |
| FourCastNet v2 (GFS init)fourcastnet_gfs | 2.55−1.09 · n=2 | 2.47−1.83 · n=1 | — | — | — | — | — | — | — | |
| FourCastNet v2 (IFS init)fourcastnet_ifs | 2.60−1.31 · n=2 | 2.60−1.67 · n=1 | — | — | — | — | — | — | — | |
| NCEP GFSgfs | 3.19+1.88 · n=29 | 3.24+1.86 · n=28 | 3.50+1.97 · n=27 | 3.54+1.75 · n=26 | 3.66+1.62 · n=25 | 3.92+1.31 · n=24 | 4.05+1.54 · n=23 | 4.38+1.36 · n=22 | 4.80+1.21 · n=21 | |
| GraphCast (GFS init)graphcast_gfs | 3.72−3.37 · n=2 | 4.79−4.69 · n=1 | — | — | — | — | — | — | — | |
| GraphCast (IFS init)graphcast_ifs | 3.00−2.53 · n=28 | 3.12−2.60 · n=28 | 3.22−2.60 · n=28 | 3.34−2.51 · n=29 | 3.54−2.64 · n=30 | 3.49−2.54 · n=30 | 3.84−2.78 · n=30 | 4.32−3.11 · n=30 | 4.63−3.22 · n=30 | |
| ECMWF IFS HRESifs_hres | 2.19−0.52 · n=29 | 2.30−0.24 · n=28 | 2.41−0.19 · n=27 | 2.54+0.00 · n=26 | 2.85+0.08 · n=25 | 3.13+0.25 · n=24 | 3.11−0.28 · n=23 | 3.42−0.18 · n=22 | 3.80−0.63 · n=21 | |
| Pangu-Weather (GFS init)pangu_gfs | 2.93−0.58 · n=2 | 2.80−1.06 · n=1 | — | — | — | — | — | — | — | |
| Pangu-Weather (IFS init)pangu_ifs | 2.46−0.42 · n=2 | 2.70−0.76 · n=1 | — | — | — | — | — | — | — | |
| Persistence (baseline)persistence | 2.83+0.01 · n=30 | 3.46+0.09 · n=30 | 3.90−0.01 · n=30 | 3.93+0.04 · n=30 | 4.03+0.03 · n=30 | 4.19+0.06 · n=30 | 4.29+0.04 · n=30 | 4.30−0.04 · n=30 | 4.50−0.10 · n=30 |
Where the errors are
Both maps show a fixed quantity per station. Neither shows which model wins where: on samples this short the per-station winner is mostly noise, and a map of winners would invite a comparison the data cannot carry.
ECMWF IFS HRES bias
Mean bias of the fixed reference model ECMWF IFS HRES (lead day 1, sampled daily maximum, 30d window, 12Z, bilinear). One model everywhere, so the colours compare stations, not models. Dot area grows with the number of scored days; fill is the mean bias (warm = model too warm, cool = too cold, grey = interval includes zero). The frame is a latitude/longitude graticule on an Albers projection.
Source: CastCheck · station coordinates NWS/NOAA · Albers conic projection · stations.csv. Hover a dot for its numbers.
model too cold interval includes zero model too warm dot area ∝ scored days
ECMWF IFS HRES bias as a table
| Station | Model | MAE °F | Bias °F | n |
|---|---|---|---|---|
| KATL Atlanta Hartsfield | ECMWF IFS HRES | 1.79 | +0.29 | 29 |
| KAUS Austin Bergstrom | ECMWF IFS HRES | 2.11 | −1.83 | 29 |
| KBOS Boston Logan | ECMWF IFS HRES | 2.48 | +0.70 | 26 |
| KDCA Washington Reagan | ECMWF IFS HRES | 1.95 | +0.21 | 29 |
| KDEN Denver Intl | ECMWF IFS HRES | 2.31 | +1.46 | 28 |
| KDFW Dallas-Fort Worth | ECMWF IFS HRES | 1.51 | +1.13 | 29 |
| KEWR Newark Liberty | ECMWF IFS HRES | 3.37 | −2.17 | 29 |
| KIAH Houston Bush | ECMWF IFS HRES | 1.63 | −0.94 | 29 |
| KLAS Las Vegas Harry Reid | ECMWF IFS HRES | 2.79 | −2.41 | 28 |
| KLAX Los Angeles Intl | ECMWF IFS HRES | 1.37 | +0.38 | 28 |
| KMIA Miami Intl | ECMWF IFS HRES | 2.73 | −1.26 | 29 |
| KMSP Minneapolis-St Paul | ECMWF IFS HRES | 2.29 | −0.70 | 29 |
| KMSY New Orleans Intl | ECMWF IFS HRES | 1.92 | −1.04 | 29 |
| KNYC New York Central Park | ECMWF IFS HRES | 2.50 | +1.05 | 29 |
| KOKC Oklahoma City | ECMWF IFS HRES | 2.03 | −0.47 | 29 |
| KORD Chicago O'Hare | ECMWF IFS HRES | 1.55 | −0.38 | 29 |
| KPHL Philadelphia Intl | ECMWF IFS HRES | 2.24 | −1.22 | 29 |
| KPHX Phoenix Sky Harbor | ECMWF IFS HRES | 2.08 | −1.06 | 28 |
| KSAN San Diego Lindbergh | ECMWF IFS HRES | 3.31 | −3.31 | 28 |
| KSAT San Antonio Intl | ECMWF IFS HRES | 1.21 | −0.59 | 29 |
| KSEA Seattle-Tacoma | ECMWF IFS HRES | 2.16 | +0.81 | 28 |
| KSFO San Francisco Intl | ECMWF IFS HRES | 2.59 | −2.18 | 28 |
| KTTN Trenton Mercer | ECMWF IFS HRES | 2.36 | +1.56 | 29 |
All-model mean bias
Mean bias averaged over every scored model at each station (lead day 1, sampled daily maximum, 30d window, 12Z, bilinear). A station that is cold for all of them is a station property — elevation, coastline, a grid cell that is partly sea — not a model ranking. Dot area grows with the number of scored days; fill is the mean bias (warm = model too warm, cool = too cold). The frame is a latitude/longitude graticule on an Albers projection.
Source: CastCheck · station coordinates NWS/NOAA · Albers conic projection · stations.csv. Hover a dot for its numbers.
model too cold interval includes zero model too warm dot area ∝ scored days
All-model mean bias as a table
| Station | Models | MAE °F | Bias °F | n |
|---|---|---|---|---|
| KATL Atlanta Hartsfield | 11 models | 2.29 | −0.61 | 29 |
| KAUS Austin Bergstrom | 11 models | 3.99 | −3.51 | 29 |
| KBOS Boston Logan | 11 models | 2.24 | +0.63 | 26 |
| KDCA Washington Reagan | 11 models | 2.33 | +0.60 | 29 |
| KDEN Denver Intl | 11 models | 1.64 | −0.50 | 28 |
| KDFW Dallas-Fort Worth | 11 models | 2.77 | −1.68 | 29 |
| KEWR Newark Liberty | 11 models | 2.38 | −0.61 | 29 |
| KIAH Houston Bush | 11 models | 1.75 | −0.52 | 29 |
| KLAS Las Vegas Harry Reid | 11 models | 4.56 | −4.36 | 28 |
| KLAX Los Angeles Intl | 11 models | 1.89 | −0.57 | 28 |
| KMIA Miami Intl | 11 models | 3.33 | −2.59 | 29 |
| KMSP Minneapolis-St Paul | 11 models | 3.96 | −2.81 | 29 |
| KMSY New Orleans Intl | 11 models | 2.98 | −2.77 | 29 |
| KNYC New York Central Park | 11 models | 2.38 | +1.63 | 29 |
| KOKC Oklahoma City | 11 models | 3.20 | −1.95 | 29 |
| KORD Chicago O'Hare | 11 models | 1.56 | −0.57 | 29 |
| KPHL Philadelphia Intl | 11 models | 2.27 | −0.39 | 29 |
| KPHX Phoenix Sky Harbor | 11 models | 5.48 | −5.19 | 28 |
| KSAN San Diego Lindbergh | 11 models | 5.51 | −5.44 | 28 |
| KSAT San Antonio Intl | 11 models | 3.38 | −2.93 | 29 |
| KSEA Seattle-Tacoma | 11 models | 3.57 | −2.41 | 28 |
| KSFO San Francisco Intl | 11 models | 2.50 | −1.80 | 28 |
| KTTN Trenton Mercer | 11 models | 3.16 | +2.94 | 29 |
Data availability
Each model is scored only over its own available period (2026-01-01 → 2026-08-30); the windows above are therefore not identical across models. Pairwise comparisons on the permanent-link pages use common days only.
| Model | Period | Scored days | Coverage |
|---|---|---|---|
| ECMWF AIFS Single | 2026-08-02 → 2026-08-30 | 29 | 2026-01-012026-08-30 |
| Aurora (GFS init) | 2026-08-29 → 2026-08-30 | 2 | 2026-01-012026-08-30 |
| Aurora (IFS init) | 2026-08-29 → 2026-08-30 | 2 | 2026-01-012026-08-30 |
| FourCastNet v2 (GFS init) | 2026-08-29 → 2026-08-30 | 2 | 2026-01-012026-08-30 |
| FourCastNet v2 (IFS init) | 2026-08-29 → 2026-08-30 | 2 | 2026-01-012026-08-30 |
| NCEP GFS | 2026-08-02 → 2026-08-30 | 29 | 2026-01-012026-08-30 |
| GraphCast (GFS init) | 2026-01-02 → 2026-08-30 | 36 | 2026-01-012026-08-30 |
| GraphCast (IFS init) | 2026-01-02 → 2026-08-30 | 197 | 2026-01-012026-08-30 |
| ECMWF IFS HRES | 2026-08-02 → 2026-08-30 | 29 | 2026-01-012026-08-30 |
| Pangu-Weather (GFS init) | 2026-08-29 → 2026-08-30 | 2 | 2026-01-012026-08-30 |
| Pangu-Weather (IFS init) | 2026-08-29 → 2026-08-30 | 2 | 2026-01-012026-08-30 |
| Persistence (baseline) | 2026-01-01 → 2026-08-30 | 215 | 2026-01-012026-08-30 |
Stations23
full station table- KNYC — New York Central Park
- KEWR — Newark Liberty
- KPHL — Philadelphia Intl
- KTTN — Trenton Mercer
- KBOS — Boston Logan
- KDCA — Washington Reagan
- KATL — Atlanta Hartsfield
- KMIA — Miami Intl
- KORD — Chicago O'Hare
- KMSP — Minneapolis-St Paul
- KDFW — Dallas-Fort Worth
- KIAH — Houston Bush
- KAUS — Austin Bergstrom
- KSAT — San Antonio Intl
- KMSY — New Orleans Intl
- KOKC — Oklahoma City
- KDEN — Denver Intl
- KPHX — Phoenix Sky Harbor
- KLAS — Las Vegas Harry Reid
- KLAX — Los Angeles Intl
- KSAN — San Diego Lindbergh
- KSFO — San Francisco Intl
- KSEA — Seattle-Tacoma