पाऊस

The honesty board

Model vs. raw forecast

Every forecast we publish gets logged, then graded against what Mumbai's airport actually observed — no cherry-picking. The calibrated model only ships when it beats both the raw forecast and climatology on a time-ordered holdout. If it can't, it doesn't ship.

Calibrated model — beating baselines

7710 / 200 labelled

Logistic champion trained 2026-09-21 06:35:23 UTC. Holdout recomputed at build on the latest diary (same 80/20 time split as train.py).

Holdout skill · model vs baselines

Time-ordered last 20% of graded hours (1,542 test · 6,168 train). Lower Brier is better. Raw rain call = forecast ≥ 0.3 mm·h⁻¹(same as the Now page). Logistic rain call = p ≥ 0.5.

Model Brier
0.1352

live champion

Raw Brier
0.3632

hard ≥0.3 mm call

Climatology
0.1366

train base rate

Skill vs raw
+62.8%

Brier skill score

On the holdout, logistic p≥0.5: 25 hits · 38 false alarms · 209 misses · 1270 correct dry.

At the decision threshold (p≥0.5)

What the numbers above mean if you actually act on them. CSI ignores the many correct-dry hours, so unlike accuracy it can't be flattered by a dry month.

POD
10.7%

of rainy hours we called

FAR
60.3%

of rain calls that were wrong

CSI
9.2%

threat score, dry hours excluded

Bias
0.27

1.00 = we call rain as often as it rains

Reliability — when we say X%, does it rain X%?

Brier alone hides miscalibration. A well-calibrated model hasobserved close to forecast on every row; that is the diagonal a reliability diagram draws. Rows with no holdout hours are left blank rather than shown as zero.

Forecast bandHoursMean forecastObserved
0–20%11689.1%13.7%
20–40%24328.4%16.5%
40–60%10748.4%16.8%
60–80%1569.6%66.7%
80–100%988.2%66.7%

What we've collected

Snapshots logged
16,128

across 672 hourly runs

Graded rows
7,710

observed by METAR

Observed rain hours
1,719

of 7,710 graded

Awaiting a grade
8,418

future hours, not yet observed

The raw forecast's record

On all 7,710 graded hours (not just the holdout), counting a rain call whenever the raw forecast read ≥ 0.3 mm·h⁻¹— the same line the Now page draws.

Raw forecast outcomes versus observed rain, on 7710 graded hours.
Observed
rain
Observed
dry
Forecast rain≥ 0.3 mm1269hits2578false alarms
Forecast dry< 0.3 mm450misses3413correct dry

The raw forecast cries wolf: 3847 rain calls, only 1269 right — and it still missed 450 of 1719 real rain hours. Cutting those false alarms is the whole job of the calibrated model.

Promotion gate

Daily retrain on GitHub Actions. A candidate ships only when its holdout Brier is strictly no worse than raw, climatology, and the current champion. A worse model can never replace a better one; a champion that loses to raw or clim is demoted to raw passthrough.

  1. 1 Beat the raw forecast (≥0.3 mm hard call).
  2. 2 Beat climatology (train-set base rate).
  3. 3 Beat the current champion model.
  4. 4 Otherwise — rejected. The champion stays (or demotes to raw).