Definition and propriety
For binary forecast p and outcome y in {0,1}, BS=(p-y)^2. Average across N evaluated questions. Lower is better; this convention ranges from 0 to 1.
For a categorical vector p_1,…,p_K and one-hot outcome y_1,…,y_K, BS=sum_k (p_k-y_k)^2. This convention ranges from 0 to 2. A binary vector [p,1-p] therefore scores twice the one-component binary convention.
If the actual event probability is q, expected one-component score is (p-q)^2+q(1-q), minimized at p=q. This is why sincere probability reporting is optimal under the score in isolation. Tournament prizes and strategic participation can alter incentives.
A single realized score is not an estimate of “true probability.” An event assigned 10% can occur without demonstrating miscalibration.
Comparison design
Freeze the scoring dataset and horizon. Compare forecasters on matched questions and equivalent cutoffs. A perfect score on one obvious near-resolution question is not comparable with a broad record of difficult long-horizon questions.
For a simple benchmark, use an ex ante base-rate forecast or other prespecified predictor. Brier skill score is 1-BS_model/BS_reference when BS_reference>0. A zero reference score makes the ratio undefined. State how the reference was obtained; a retrospectively estimated prevalence is descriptive, not an operational ex ante forecast.
Weighting must be explicit and ideally fixed independently of outcomes. If importance weights depend on which outcomes happened, strict propriety may be lost. Do not remove inconvenient resolved questions after seeing their scores.
Repeated updates
For time-weighted evaluation, carry the last available forecast forward over a prespecified grid, score each grid time, average within a question, then average questions. Multiple posts between grid points do not create additional independent observations. Declare behavior before the first forecast and after closure.
Do not average every submitted update equally: that rewards submission frequency and changes the score’s meaning.
Helper
python3 scripts/score.py --input records.json
Supply a JSON list. Binary records require question_id, p and outcome. Categorical records use probabilities and an integer outcome index with --mode categorical. A null outcome is excluded and counted; malformed or out-of-range values are rejected. Use already adjudicated, aligned records; the helper does not verify real-world timestamps or evidence.
Scope
Nominal Brier ignores distances between categories. For ordered outcomes or continuous quantities, choose an appropriate proper rule such as a ranked probability score or CRPS and document that it is a different scoring implementation. This helper does not compute those rules.