Three different concepts
Calibration: among forecasts near p, events occur near frequency p. Resolution: separating cases into groups with differing event frequencies. Sharpness: concentration of predictive distributions, independent of outcomes. A constant base-rate predictor can be calibrated but uninformative about differences between cases.
For groups k with counts n_k, forecast value f_k, observed frequency o_k and overall frequency o_bar: REL=sum_k (n_k/N)(f_k-o_k)^2; RES=sum_k (n_k/N)(o_k-o_bar)^2; UNC=o_bar*(1-o_bar). For exact groups of identical forecast probabilities, BS=REL-RES+UNC under one-component binary scoring.
Binning trap
When forecasts vary inside a bin, replacing them with the bin mean changes the forecast system. The three-term expression then exactly equals the Brier score of the coarsened forecasts, not generally the original raw Brier.
The helper reports raw_brier, coarsened_brier, reliability, resolution, uncertainty and binning_residual=raw_brier-coarsened_brier. Exact grouping should make the residual zero up to numerical error.
Bins are useful descriptively but should not be selected after seeing outcomes just to make calibration look better. Include every nonempty bin’s count. Empty bins are not evidence of perfect calibration.
Resolved records only
Calibration is defined against adjudicated outcomes, so only resolved questions enter the table. Supply unresolved records with a null outcome: the helper excludes them and reports count_excluded and excluded_unresolved, exactly as the Brier helper does. Never score an unresolved question as a non-event; that silently manufactures observed frequencies and biases reliability toward whatever the open questions were predicted to be.
Coverage matters when questions resolve on different schedules. A ledger whose slow-resolving questions are still open is not a random sample of the ledger, so report the excluded count alongside the diagnosis rather than only the scored n.
Sampling and dependence
Small groups have noisy observed frequencies. A bin containing one 70% forecast that resolves yes does not prove 30 percentage points of stable bias. Confidence bands based on independent Bernoulli trials can be misleading when events share causes.
If intervals are needed, state the method and unit of resampling. Resample entire questions or event clusters, not individual repeated updates. Provide a sensitivity analysis for bins and time periods. The included helper intentionally does not manufacture intervals from sparse or dependent data.
Recalibration
A common candidate is sigmoid(a+b*logit(p)). a adjusts overall bias; b changes probability spread. This is a candidate statistical model, not the same as one-parameter zero-intercept extremization in every context.
Fit on a designated calibration set after model training. Test on a separate holdout. Avoid applying a mapping learned on elections to software delivery without transfer evidence. With limited data, favor reporting uncertainty over a complex corrective model.
Helper
python3 scripts/calibration.py --input records.json --groups exact
or --groups bins --bins 10.
Use one frozen binary forecast per question in the supplied JSON list. It rejects duplicate question IDs, invalid probabilities and unscored rows. Resolve relative paths from the skill directory.