पाऊस

The honesty board

Model vs. raw forecast

Every forecast we publish gets logged, then graded against what Mumbai's airport actually observed — no cherry-picking. The calibrated model only ships when it beats both the raw forecast and climatology on a time-ordered holdout. If it can't, it doesn't ship.

Calibrated model — under pressure

7780 / 200 labelled

Logistic champion trained 2026-09-21 06:35:23 UTC. Holdout recomputed at build on the latest diary (same 80/20 time split as train.py).

Holdout skill · model vs baselines

Time-ordered last 20% of graded hours (1,556 test · 6,224 train). Lower Brier is better. Raw rain call = forecast ≥ 0.3 mm·h⁻¹(same as the Now page). Logistic rain call = p ≥ 0.5.

Model Brier
0.1438

degraded

Raw Brier
0.3702

hard ≥0.3 mm call

Climatology
0.1423

train base rate

Skill vs raw
+61.2%

Brier skill score

On the holdout, logistic p≥0.5: 25 hits · 43 false alarms · 229 misses · 1259 correct dry.

At the decision threshold (p≥0.5)

What the numbers above mean if you actually act on them. CSI ignores the many correct-dry hours, so unlike accuracy it can't be flattered by a dry month.

POD
9.8%

of rainy hours we called

FAR
63.2%

of rain calls that were wrong

CSI
8.4%

threat score, dry hours excluded

Bias
0.27

1.00 = we call rain as often as it rains

Reliability — when we say X%, does it rain X%?

Brier alone hides miscalibration. A well-calibrated model hasobserved close to forecast on every row; that is the diagonal a reliability diagram draws. Rows with no holdout hours are left blank rather than shown as zero.

Forecast bandHoursMean forecastObserved
0–20%11539.2%14.7%
20–40%26528.3%19.2%
40–60%11148.4%16.2%
60–80%1868.4%55.6%
80–100%988.2%66.7%

What we've collected

Snapshots logged
16,224

across 676 hourly runs

Graded rows
7,780

observed by METAR

Observed rain hours
1,742

of 7,780 graded

Awaiting a grade
8,444

future hours, not yet observed

The raw forecast's record

On all 7,780 graded hours (not just the holdout), counting a rain call whenever the raw forecast read ≥ 0.3 mm·h⁻¹— the same line the Now page draws.

Raw forecast outcomes versus observed rain, on 7780 graded hours.
Observed
rain
Observed
dry
Forecast rain≥ 0.3 mm1272hits2593false alarms
Forecast dry< 0.3 mm470misses3445correct dry

The raw forecast cries wolf: 3865 rain calls, only 1272 right — and it still missed 470 of 1742 real rain hours. Cutting those false alarms is the whole job of the calibrated model.

Promotion gate

Daily retrain on GitHub Actions. A candidate ships only when its holdout Brier is strictly no worse than raw, climatology, and the current champion. A worse model can never replace a better one; a champion that loses to raw or clim is demoted to raw passthrough.

  1. 1 Beat the raw forecast (≥0.3 mm hard call).
  2. 2 Beat climatology (train-set base rate).
  3. 3 Beat the current champion model.
  4. 4 Otherwise — rejected. The champion stays (or demotes to raw).