Raw model output on the native 0.25° grid — no MOS, no bias correction, no post-processing. Scores understate operational forecast quality.
Read the fairness statement →
CastCheckmethodology v0.2 · data through 2026-08-30
How far off was each weather model?
No group in this view yet has enough scored days to rank
(n < 30). Every number below is still published, greyed out,
with its sample size.
Scope: raw model output on the native 0.25° grid, daily maximum/minimum taken as the
max/min of the four common 6-hourly samples, scored against the first final NWS Daily Climate
Report, errors computed in °C and shown in °F
(methodology v0.2).
window the last 90 days
init 12Z
interpolation nearest
data through 2026-08-30
updated 2026-08-31T03:52:06+00:00
next update 2026-08-31T11:00:00+00:00
Lead day 1 — daily maximum temperature
the last 90 days, all stations pooled, 12Z initialization,
nearest interpolation. Lower MAE is better; bias is positive when the model is too warm.
Every number links to its permanent page. Skill is measured against persistence;
skill, debiased repeats it after removing each station's constant offset, which is why a
station with a steady warm bias can look skill-less in one column and skilful in the other.
Under n: the mean number of stations behind each scored day, and how many of those days
carry a QC flag on the observation.
— in the rank column means fewer than 30 scored
days, so the group is published but not ranked.
+ model too warm ·
− model too cold ·
bias interval includes zero ·
★ lowest MAE · = not distinguishable from the leader · ▼ significantly worse ·
n < 30 greyed and unranked · CI is a 95 % moving-block bootstrap interval.
Lead day 3 — daily maximum temperature
the last 90 days, all stations pooled, 12Z initialization,
nearest interpolation. Lower MAE is better; bias is positive when the model is too warm.
Every number links to its permanent page. Skill is measured against persistence;
skill, debiased repeats it after removing each station's constant offset, which is why a
station with a steady warm bias can look skill-less in one column and skilful in the other.
Under n: the mean number of stations behind each scored day, and how many of those days
carry a QC flag on the observation.
— in the rank column means fewer than 30 scored
days, so the group is published but not ranked.
the last 90 days, all stations pooled, 12Z initialization,
nearest interpolation. Lower MAE is better; bias is positive when the model is too warm.
Every number links to its permanent page. Skill is measured against persistence;
skill, debiased repeats it after removing each station's constant offset, which is why a
station with a steady warm bias can look skill-less in one column and skilful in the other.
Under n: the mean number of stations behind each scored day, and how many of those days
carry a QC flag on the observation.
— in the rank column means fewer than 30 scored
days, so the group is published but not ranked.
the last 90 days, all stations pooled, 12Z initialization,
nearest interpolation. Lower MAE is better; bias is positive when the model is too warm.
Every number links to its permanent page. Skill is measured against persistence;
skill, debiased repeats it after removing each station's constant offset, which is why a
station with a steady warm bias can look skill-less in one column and skilful in the other.
Under n: the mean number of stations behind each scored day, and how many of those days
carry a QC flag on the observation.
— in the rank column means fewer than 30 scored
days, so the group is published but not ranked.
MAE in °F with the bias underneath, the last 90 days, 12Z,
nearest, daily maximum. The sparkline is the same model's MAE across lead days
1–9 on a shared vertical scale.
Each station's best model at lead day 1 (the last 90 days, 12Z,
nearest, daily maximum): dot area grows with the number of scored days, fill is the
mean bias (warm = model too warm, cool = too cold, grey = interval includes zero). There is no
coastline because CastCheck ships no third-party boundary file; the frame is a latitude/longitude
graticule. Hover a dot for its numbers.The same data as a table
Each model is scored only over its own available period (2024-01-02 →
2026-08-30); the windows above are therefore not identical across models. Pairwise
comparisons on the permanent-link pages use common days only.