Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Evaluation contracts

Supercast provides distinct evaluation tools for different tasks. Do not interpret one tool’s checks as evidence that another task’s assumptions have been established.

ToolWhat it establishes from supplied inputs
comparePaired score means on unique, pre-resolution question records
validate_holdoutQuestion uniqueness, event-cluster separation and temporal order
grid_scoreTime-grid Brier within one question using carry-forward forecasts
cluster_bootstrapA resampled paired performance interval under declared clusters
select_calibrationTraining-only candidate choice with subsequent holdout scoring
evaluateA database snapshot evaluation with explicit exclusions

Case-level evaluation

A Case holds question ID, cluster ID, forecast time, resolution time, model probability, baseline probability and binary outcome. The model/baseline pair shares one outcome and timestamp, preventing accidental comparison of unrelated rows within the record format.

validate_holdout requires nonempty training and holdout sets. Training labels are available by the cutoff; held-out forecasts occur after it. Every forecast precedes its own resolution. Related questions must use the same cluster ID so that cluster overlap can be rejected.

The baseline’s historical availability is not independently verified by the Case structure. The caller must freeze a legitimate benchmark. Do not use retrospective prevalence as if it were an operational ex-ante forecast without saying so.

Bootstrap implementation

Clusters are sorted by ID; rows inside them are sorted by question ID before their score differences are summed. Each replicate draws the same number of clusters with replacement and divides the total paired score difference by the total selected question count. Unequal cluster sizes therefore retain question weighting.

SplitMix64 provides seeded deterministic draws, with rejection sampling to remove modulo bias. At most 128 rejection attempts are allowed for an index; failure is explicit. Sampling permits 100–100,000 replicates and at most 10 million cluster draws. The computation allocates cluster summaries and one vector of replicate statistics, not a copied question matrix per replicate.

Percentile limits use type-7 interpolation. Standard error is the sample standard deviation of the replicate statistics. A constant distribution sets degenerate: true. These are percentile bootstrap intervals, not BCa intervals or posterior credible intervals.

Calibration selection

The candidate grid contains 1–1,000 finite intercept/nonnegative-slope pairs. At most 10 million training candidate/question evaluations are permitted. Each candidate transforms interior probabilities through a sigmoid of affine log odds and is scored on training outcomes.

Exact score ties use input candidate order. The chosen parameters are then applied to the holdout once; the operation returns all training scores and the selected method’s raw, calibrated and baseline holdout scores. It does not select a new candidate because the first performed poorly on held-out data.

The caller must prevent repeated adaptive use of the same holdout. A function cannot infer prior experiments that were not supplied.

Database snapshots

evaluate accepts 1–10,000 explicit keys, a common forecast cutoff, a resolution cutoff, and a fixed baseline with its declared as-of time. Duplicate question IDs are rejected even if the version differs. A missing key fails the request rather than shrinking the cohort.

It reads a consistent database transaction, selects the most recent forecast available by the cutoff, and selects the latest resolution available by the outcome cutoff. Cases outside the question’s forecast window, missing forecasts, unresolved outcomes and voids are reported as exclusions. No scored rows produces a null summary.

This operation labels its output declared_timestamp_snapshot. Supplied historical timestamps and baseline dates do not prove prospective participation. Commit-time enforcement, cohort preregistration, and independently observed future outcomes remain necessary for a prospective performance claim.