The honesty board
Model vs. raw forecast
Every forecast we publish gets logged, then graded against what Mumbai's airport actually observed — no cherry-picking. The calibrated model only ships when it beats both the raw forecast and climatology on a time-ordered holdout. If it can't, it doesn't ship.
Calibrated model — under pressure
7780 / 200 labelled
Logistic champion trained 2026-09-21 06:35:23 UTC. Holdout recomputed at build on the latest diary (same 80/20 time split as train.py).
Holdout skill · model vs baselines
Time-ordered last 20% of graded hours (1,556 test · 6,224 train). Lower Brier is better. Raw rain call = forecast ≥ 0.3 mm·h⁻¹(same as the Now page). Logistic rain call = p ≥ 0.5.
- Model Brier
- 0.1438
- Raw Brier
- 0.3702
- Climatology
- 0.1423
- Skill vs raw
- +61.2%
degraded
hard ≥0.3 mm call
train base rate
Brier skill score
- Beats raw?yesrequired for promotion
- Beats clim?norequired for promotion
- Skill vs clim-1.0%positive = better than always-guess base rate
- Stored at promote0.1352model.json snapshot from last promotion
On the holdout, logistic p≥0.5: 25 hits · 43 false alarms · 229 misses · 1259 correct dry.
At the decision threshold (p≥0.5)
What the numbers above mean if you actually act on them. CSI ignores the many correct-dry hours, so unlike accuracy it can't be flattered by a dry month.
- POD
- 9.8%
- FAR
- 63.2%
- CSI
- 8.4%
- Bias
- 0.27
of rainy hours we called
of rain calls that were wrong
threat score, dry hours excluded
1.00 = we call rain as often as it rains
Reliability — when we say X%, does it rain X%?
Brier alone hides miscalibration. A well-calibrated model hasobserved close to forecast on every row; that is the diagonal a reliability diagram draws. Rows with no holdout hours are left blank rather than shown as zero.
| Forecast band | Hours | Mean forecast | Observed |
|---|---|---|---|
| 0–20% | 1153 | 9.2% | 14.7% |
| 20–40% | 265 | 28.3% | 19.2% |
| 40–60% | 111 | 48.4% | 16.2% |
| 60–80% | 18 | 68.4% | 55.6% |
| 80–100% | 9 | 88.2% | 66.7% |
What we've collected
- Snapshots logged
- 16,224
- Graded rows
- 7,780
- Observed rain hours
- 1,742
- Awaiting a grade
- 8,444
across 676 hourly runs
observed by METAR
of 7,780 graded
future hours, not yet observed
The raw forecast's record
On all 7,780 graded hours (not just the holdout), counting a rain call whenever the raw forecast read ≥ 0.3 mm·h⁻¹— the same line the Now page draws.
| Observed rain | Observed dry | |
|---|---|---|
| Forecast rain≥ 0.3 mm | 1272hits | 2593false alarms |
| Forecast dry< 0.3 mm | 470misses | 3445correct dry |
- Base rate22.4%of graded hours actually had rain
- Detection · POD73.0%of 1742 rain hours, 1272 were forecast
- False-alarm ratio67.1%of 3865 rain calls, 2593 were wrong
- Raw agreement60.6%forecast matched observation
The raw forecast cries wolf: 3865 rain calls, only 1272 right — and it still missed 470 of 1742 real rain hours. Cutting those false alarms is the whole job of the calibrated model.
Promotion gate
Daily retrain on GitHub Actions. A candidate ships only when its holdout Brier is strictly no worse than raw, climatology, and the current champion. A worse model can never replace a better one; a champion that loses to raw or clim is demoted to raw passthrough.
- 1 Beat the raw forecast (≥0.3 mm hard call).
- 2 Beat climatology (train-set base rate).
- 3 Beat the current champion model.
- 4 Otherwise — rejected. The champion stays (or demotes to raw).