Diagnose calibration
Calibration asks whether events assigned probability near p occur near frequency p. Resolution asks whether forecasts distinguish groups with different outcome frequencies. Sharpness describes how concentrated the predictions are without reference to outcomes.
A constant base-rate predictor can be calibrated and still fail to distinguish individual cases. A sharp forecaster can be confidently wrong. Do not use either sharpness or ensemble agreement as a substitute for outcome-based evaluation.
Reliability tables
calibration::diagnose accepts one frozen binary forecast per question. Question IDs must be nonempty and unique. Unresolved outcomes are excluded and counted. It groups probabilities either exactly or into an explicitly chosen number of equal-width bins.
Each nonempty bin reports its count, mean forecast, and observed frequency. Empty bins are not evidence of perfect calibration. Choose the grouping before inspecting results, then report reasonable sensitivity to bin choices where appropriate.
extern crate supercast;
use supercast::{Probability as P, calibration::{Record, Grouping, diagnose}};
let rows = vec![
Record { question_id: "a".into(), probability: P::new(0.2)?, outcome: Some(false) },
Record { question_id: "b".into(), probability: P::new(0.8)?, outcome: Some(true) },
Record { question_id: "c".into(), probability: P::new(0.7)?, outcome: None },
];
let report = diagnose(&rows, Grouping::Exact)?;
assert_eq!(report.resolved, 2);
assert_eq!(report.excluded_unresolved, 1);
assert!(report.binning_residual.abs() < 1e-12);
Ok::<(), supercast::Error>(())
The decomposition and its trap
For bin k, let f be its mean forecast, o its outcome frequency, and w its fraction of all scored questions. With overall outcome frequency o-bar:
reliability = sum(w * (f-o)^2)
resolution = sum(w * (o-o_bar)^2)
uncertainty = o_bar * (1-o_bar)
For exact groups of identical probabilities, Brier equals reliability minus resolution plus uncertainty. For wider bins, that expression corresponds to the coarsened forecast system, where each original probability is replaced by its bin mean. It need not equal the raw Brier score.
Supercast therefore reports raw Brier, coarsened Brier, all three terms, and the binning residual raw - coarsened. Treating a nonzero residual as a bug in the forecast system would confuse a grouping transformation with the original predictions.
Sampling limits
A single 70% forecast that resolves yes is not proof of a stable 30-point bias. Small groups have noisy frequencies. Related questions can share outcomes or common causes, so an independent-Bernoulli interval may overstate precision.
The diagnostic does not fit a correction or manufacture intervals. Use cluster-level validation for a paired score comparison, and separate calibration training from evaluation. Coverage also matters: slow-resolving open questions may differ systematically from the resolved subset.