The honesty board
Model vs. raw forecast
Every forecast we publish gets logged, then graded against what Mumbai's airport actually observed — no cherry-picking. The calibrated model only ships when it beats both the raw forecast and climatology on a time-ordered holdout. If it can't, it doesn't ship.
Calibrated model — beating baselines
7710 / 200 labelled
Logistic champion trained 2026-09-21 06:35:23 UTC. Holdout recomputed at build on the latest diary (same 80/20 time split as train.py).
Holdout skill · model vs baselines
Time-ordered last 20% of graded hours (1,542 test · 6,168 train). Lower Brier is better. Raw rain call = forecast ≥ 0.3 mm·h⁻¹(same as the Now page). Logistic rain call = p ≥ 0.5.
- Model Brier
- 0.1352
- Raw Brier
- 0.3632
- Climatology
- 0.1366
- Skill vs raw
- +62.8%
live champion
hard ≥0.3 mm call
train base rate
Brier skill score
- Beats raw?yesrequired for promotion
- Beats clim?yesrequired for promotion
- Skill vs clim+1.0%positive = better than always-guess base rate
- Stored at promote0.1352model.json snapshot from last promotion
On the holdout, logistic p≥0.5: 25 hits · 38 false alarms · 209 misses · 1270 correct dry.
At the decision threshold (p≥0.5)
What the numbers above mean if you actually act on them. CSI ignores the many correct-dry hours, so unlike accuracy it can't be flattered by a dry month.
- POD
- 10.7%
- FAR
- 60.3%
- CSI
- 9.2%
- Bias
- 0.27
of rainy hours we called
of rain calls that were wrong
threat score, dry hours excluded
1.00 = we call rain as often as it rains
Reliability — when we say X%, does it rain X%?
Brier alone hides miscalibration. A well-calibrated model hasobserved close to forecast on every row; that is the diagonal a reliability diagram draws. Rows with no holdout hours are left blank rather than shown as zero.
| Forecast band | Hours | Mean forecast | Observed |
|---|---|---|---|
| 0–20% | 1168 | 9.1% | 13.7% |
| 20–40% | 243 | 28.4% | 16.5% |
| 40–60% | 107 | 48.4% | 16.8% |
| 60–80% | 15 | 69.6% | 66.7% |
| 80–100% | 9 | 88.2% | 66.7% |
What we've collected
- Snapshots logged
- 16,128
- Graded rows
- 7,710
- Observed rain hours
- 1,719
- Awaiting a grade
- 8,418
across 672 hourly runs
observed by METAR
of 7,710 graded
future hours, not yet observed
The raw forecast's record
On all 7,710 graded hours (not just the holdout), counting a rain call whenever the raw forecast read ≥ 0.3 mm·h⁻¹— the same line the Now page draws.
| Observed rain | Observed dry | |
|---|---|---|
| Forecast rain≥ 0.3 mm | 1269hits | 2578false alarms |
| Forecast dry< 0.3 mm | 450misses | 3413correct dry |
- Base rate22.3%of graded hours actually had rain
- Detection · POD73.8%of 1719 rain hours, 1269 were forecast
- False-alarm ratio67.0%of 3847 rain calls, 2578 were wrong
- Raw agreement60.7%forecast matched observation
The raw forecast cries wolf: 3847 rain calls, only 1269 right — and it still missed 450 of 1719 real rain hours. Cutting those false alarms is the whole job of the calibrated model.
Promotion gate
Daily retrain on GitHub Actions. A candidate ships only when its holdout Brier is strictly no worse than raw, climatology, and the current champion. A worse model can never replace a better one; a champion that loses to raw or clim is demoted to raw passthrough.
- 1 Beat the raw forecast (≥0.3 mm hard call).
- 2 Beat climatology (train-set base rate).
- 3 Beat the current champion model.
- 4 Otherwise — rejected. The champion stays (or demotes to raw).