Raw model output on the native 0.25° grid — no MOS, no bias correction, no post-processing. Scores understate operational forecast quality.
Read the fairness statement →
CastCheckmethodology v0.2 · data through 2026-08-30
How far off was each weather model?
Over all available history, across 23 U.S. stations, the most accurate
raw daily maximum temperature forecast 1 day ahead
is GraphCast (IFS init): mean absolute error
4.87 °F[4.71, 5.05],
bias -4.79 °F, on n = 33 scored days.
Scope: raw model output on the native 0.25° grid, daily maximum/minimum taken as the
max/min of the four common 6-hourly samples, scored against the first final NWS Daily Climate
Report, errors computed in °C and shown in °F
(methodology v0.2).
window all available history
init 00Z
interpolation bilinear
data through 2026-08-30
updated 2026-08-31T03:52:06+00:00
next update 2026-08-31T11:00:00+00:00
Lead day 1 — daily maximum temperature
all available history, all stations pooled, 00Z initialization,
bilinear interpolation. Lower MAE is better; bias is positive when the model is too warm.
Every number links to its permanent page. Skill is measured against persistence;
skill, debiased repeats it after removing each station's constant offset, which is why a
station with a steady warm bias can look skill-less in one column and skilful in the other.
Under n: the mean number of stations behind each scored day, and how many of those days
carry a QC flag on the observation.
+ model too warm ·
− model too cold ·
bias interval includes zero ·
★ lowest MAE · = not distinguishable from the leader · ▼ significantly worse ·
n < 30 greyed and unranked · CI is a 95 % moving-block bootstrap interval.
Lead day 3 — daily maximum temperature
all available history, all stations pooled, 00Z initialization,
bilinear interpolation. Lower MAE is better; bias is positive when the model is too warm.
Every number links to its permanent page. Skill is measured against persistence;
skill, debiased repeats it after removing each station's constant offset, which is why a
station with a steady warm bias can look skill-less in one column and skilful in the other.
Under n: the mean number of stations behind each scored day, and how many of those days
carry a QC flag on the observation.
all available history, all stations pooled, 00Z initialization,
bilinear interpolation. Lower MAE is better; bias is positive when the model is too warm.
Every number links to its permanent page. Skill is measured against persistence;
skill, debiased repeats it after removing each station's constant offset, which is why a
station with a steady warm bias can look skill-less in one column and skilful in the other.
Under n: the mean number of stations behind each scored day, and how many of those days
carry a QC flag on the observation.
all available history, all stations pooled, 00Z initialization,
bilinear interpolation. Lower MAE is better; bias is positive when the model is too warm.
Every number links to its permanent page. Skill is measured against persistence;
skill, debiased repeats it after removing each station's constant offset, which is why a
station with a steady warm bias can look skill-less in one column and skilful in the other.
Under n: the mean number of stations behind each scored day, and how many of those days
carry a QC flag on the observation.
— in the rank column means fewer than 30 scored
days, so the group is published but not ranked.
MAE in °F with the bias underneath, all available history, 00Z,
bilinear, daily maximum. The sparkline is the same model's MAE across lead days
1–9 on a shared vertical scale.
Each station's best model at lead day 1 (all available history, 00Z,
bilinear, daily maximum): dot area grows with the number of scored days, fill is the
mean bias (warm = model too warm, cool = too cold, grey = interval includes zero). There is no
coastline because CastCheck ships no third-party boundary file; the frame is a latitude/longitude
graticule. Hover a dot for its numbers.The same data as a table
Each model is scored only over its own available period (2024-01-02 →
2026-08-30); the windows above are therefore not identical across models. Pairwise
comparisons on the permanent-link pages use common days only.