Score predictions properly
A proper scoring rule rewards an honest predictive distribution in expectation. A realized score evaluates a particular prediction/outcome pair; it does not reveal a hidden “true probability” from one outcome.
Use a scoring convention before seeing the outcomes, keep it consistent across comparisons, and report which unresolved or void questions were excluded.
Binary Brier and log loss
Binary Brier is (p-y)^2, where y is zero or one. Supercast uses the one-component zero-to-one convention. Lower is better.
extern crate supercast;
use supercast::{Probability as P, score};
let p = P::new(0.3)?;
assert!((score::brier(p, true) - 0.49).abs() < 1e-12);
assert!(score::log_loss(P::ZERO, true).is_infinite());
assert!(score::brier_skill(0.2, 0.0).is_err());
Ok::<(), supercast::Error>(())
Log loss is -ln(p) for yes and -ln(1-p) for no. A confidently wrong endpoint incurs positive infinity. JSON cannot encode infinity as a number, so binary_score returns a tagged value: positive_infinity, or finite with a numeric value.
Brier skill is 1 - model_score / reference_score. A zero reference score makes the ratio undefined; an unrepresentable ratio is also an error. Absolute scores and their difference can still be reported when relative skill is unavailable.
Categorical and ordered outcomes
A Categorical distribution has at least two entries and sums to one within tolerance. Categorical Brier sums squared errors across categories and ranges from zero to two. A two-category vector therefore produces twice the one-component binary score.
The ranked probability score instead compares cumulative probabilities over ordered categories. Supercast returns the unnormalized sum over K-1 thresholds, on a zero-to-K-1 scale. Use it for ordered categories such as low/medium/high, not unrelated labels.
Continuous outcomes
The empirical continuous ranked probability score, CRPS, evaluates a sample-based predictive distribution in the outcome’s units:
CRPS = mean(|sample - outcome|) - 0.5 * mean(|sample_i - sample_j|)
The implementation sorts the caller’s sample slice and integrates squared empirical-CDF error over adjacent gaps in O(n log n), an equivalent nonnegative formula that avoids a quadratic distance matrix and cancellation between large expectations. It preserves finite answers when some unscaled intermediate distances would overflow; truly unrepresentable results are errors.
extern crate supercast;
use supercast::score::empirical_crps;
let mut samples = [1.0, 2.0, 3.0];
let score = empirical_crps(&mut samples, 2.0)?;
assert!((score - 2.0 / 9.0).abs() < 1e-12);
Ok::<(), supercast::Error>(())
Repeated forecasts
Do not average every post as if each were an independent question. A frequent poster would otherwise receive different weighting. grid_brier carries the latest available forecast forward over a prespecified grid and averages within one question. Average those question scores according to the frozen cohort policy.
The grid must be strictly increasing, and a forecast must exist at its first time. The function does not fill a missing start with 0.5. Align grid times with the event window and evaluation protocol before calling it.