Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Validate forecasting skill

A winning score on a selected sample is not sufficient evidence of persistent skill. Many equally capable forecasters can produce an impressive winner by chance. Select a method on one set of outcomes and evaluate it on new, appropriately separated questions.

Freeze the comparison

Declare eligible questions, lead times or snapshots, resolution rules, baselines, weights, and exclusion rules before evaluation. Record attempted and missing questions as well as scored ones. Selective abstention can make a leaderboard misleading if coverage is hidden.

evaluation::Case pairs a model forecast and baseline on the same question, time, outcome, and event cluster. compare reports matched means and their difference. It rejects repeated question IDs and forecasts at or after resolution.

validate_holdout also requires all training labels by the cutoff, all holdout forecasts after it, and no event cluster shared across the split. The supplied IDs and timestamps must reflect the real information structure. The function cannot detect undisclosed pretraining or retrieval contamination.

Uncertainty with event clusters

cluster_bootstrap resamples entire declared event clusters with replacement, using the same draw for model and baseline. Questions inside each selected cluster remain together. The statistic remains the question-weighted paired mean difference.

It returns a percentile interval, bootstrap standard error, counts, confidence level, seed, replicate count, and a degeneracy flag. The seed makes an analysis reproducible. At least two clusters are computationally required; substantially more may be needed for useful inference.

extern crate supercast;
use supercast::{Probability as P, evaluation::{Case, cluster_bootstrap}};

let cases = [
    Case { question_id: "a".into(), cluster_id: "a".into(), forecast_at: 1,
        resolved_at: 5, model: P::ONE, baseline: P::HALF, outcome: true },
    Case { question_id: "b".into(), cluster_id: "b".into(), forecast_at: 1,
        resolved_at: 5, model: P::ZERO, baseline: P::HALF, outcome: true },
];
let report = cluster_bootstrap(&cases, 1000, 0.95, 7)?;
assert_eq!(report.clusters, 2);
assert_eq!(report.difference, 0.25); // positive means the model did worse
Ok::<(), supercast::Error>(())

This interval is about a performance statistic under sampling assumptions, not about the true probability of one event. A degenerate interval does not prove an absence of model uncertainty. Few independent clusters, biased selection, and temporal change can make percentile coverage unreliable.

Select a calibration candidate

select_calibration takes a prespecified grid of intercept/slope pairs. It picks the lowest training Brier, resolving exact ties by candidate order, then reports raw, calibrated, and baseline holdout scores. The chosen candidate remains chosen even if it performs worse on the holdout.

Changing the grid after inspecting that result consumes the holdout. Use a new evaluation period for a subsequent selection cycle. The library limits candidate and simulation workloads, but a computational limit is not a statistical adequacy rule.

Database evaluation

The evaluate JSON command reads a declared cohort at common forecast and resolution cutoffs. It reports which cases had no forecast, were unresolved, were void, or fell outside the allowed snapshot window. With no scored cases it returns a null summary rather than a fabricated zero score.

That command uses declared timestamps. Imported historical forecasts are not made prospective by assigning them earlier dates. A surrounding prospective process must freeze the plan and preserve real recording times. Claims of elite skill need that external evidence, not merely a passed software gate.