Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Score predictions properly

A proper scoring rule rewards an honest predictive distribution in expectation. A realized score evaluates a particular prediction/outcome pair; it does not reveal a hidden “true probability” from one outcome.

Use a scoring convention before seeing the outcomes, keep it consistent across comparisons, and report which unresolved or void questions were excluded.

Binary Brier and log loss

Binary Brier is (p-y)^2, where y is zero or one. Supercast uses the one-component zero-to-one convention. Lower is better.

extern crate supercast;
use supercast::{Probability as P, score};

let p = P::new(0.3)?;
assert!((score::brier(p, true) - 0.49).abs() < 1e-12);
assert!(score::log_loss(P::ZERO, true).is_infinite());
assert!(score::brier_skill(0.2, 0.0).is_err());
Ok::<(), supercast::Error>(())

Log loss is -ln(p) for yes and -ln(1-p) for no. A confidently wrong endpoint incurs positive infinity. JSON cannot encode infinity as a number, so binary_score returns a tagged value: positive_infinity, or finite with a numeric value.

Brier skill is 1 - model_score / reference_score. A zero reference score makes the ratio undefined; an unrepresentable ratio is also an error. Absolute scores and their difference can still be reported when relative skill is unavailable.

Categorical and ordered outcomes

A Categorical distribution has at least two entries and sums to one within tolerance. Categorical Brier sums squared errors across categories and ranges from zero to two. A two-category vector therefore produces twice the one-component binary score.

The ranked probability score instead compares cumulative probabilities over ordered categories. Supercast returns the unnormalized sum over K-1 thresholds, on a zero-to-K-1 scale. Use it for ordered categories such as low/medium/high, not unrelated labels.

Continuous outcomes

The empirical continuous ranked probability score, CRPS, evaluates a sample-based predictive distribution in the outcome’s units:

CRPS = mean(|sample - outcome|) - 0.5 * mean(|sample_i - sample_j|)

The implementation sorts the caller’s sample slice and integrates squared empirical-CDF error over adjacent gaps in O(n log n), an equivalent nonnegative formula that avoids a quadratic distance matrix and cancellation between large expectations. It preserves finite answers when some unscaled intermediate distances would overflow; truly unrepresentable results are errors.

extern crate supercast;
use supercast::score::empirical_crps;

let mut samples = [1.0, 2.0, 3.0];
let score = empirical_crps(&mut samples, 2.0)?;
assert!((score - 2.0 / 9.0).abs() < 1e-12);
Ok::<(), supercast::Error>(())

Repeated forecasts

Do not average every post as if each were an independent question. A frequent poster would otherwise receive different weighting. grid_brier carries the latest available forecast forward over a prespecified grid and averages within one question. Average those question scores according to the frozen cohort policy.

The grid must be strictly increasing, and a forecast must exist at its first time. The function does not fill a missing start with 0.5. Align grid times with the event window and evaluation protocol before calling it.