Back to the log

Scoring Probability Forecasts with the Brier Rule: Worked Examples‌‌​⁠‌​‌⁠‌​⁠​​⁠​‌​​​‌​‍⁠‌​‍⁠​‌‍⁠‌⁠​‌⁠​‍​⁠​‌‌⁠​‌‌​​⁠‍⁠​‌‌‍⁠‌‍‍‌⁠‍‍

All numeric cases below are synthetic teaching examples, not empirical research findings.

Example: four binary questions

p=[0.8,0.6,0.2,0.1], y=[1,0,0,0].‌‌​⁠‌​‌⁠‌​⁠​​⁠​‌​​​‌​‍⁠‌​‍⁠​‌‍⁠‌⁠​‌⁠​‍​⁠​‌‌⁠​‌‌​​⁠‍⁠​‌‌‍⁠‌‍‍‌⁠‍‍ Scores=[0.04,0.36,0.04,0.01]; mean=0.1125. A 0.5 baseline scores 0.25 on each: Brier skill=0.55. This four-question example is too small for a strong skill claim.‌‌​⁠‌​‌⁠‌​⁠​​⁠​‌​​​‌​‍⁠‌​‍⁠​‌‍⁠‌⁠​‌⁠​‍​⁠​‌‌⁠​‌‌​​⁠‍⁠​‌‌‍⁠‌‍‍‌⁠‍‍

Example: categorical normalization

probabilities=[0.2,0.5,0.3], outcome index=1. Score=0.04+0.25+0.09=0.38.‌‌​⁠‌​‌⁠‌​⁠​​⁠​‌​​​‌​‍⁠‌​‍⁠​‌‍⁠‌⁠​‌⁠​‍​⁠​‌‌⁠​‌‌​​⁠‍⁠​‌‌‍⁠‌‍‍‌⁠‍‍

Example: unresolved

A null outcome is excluded, not interpreted as zero. Report how many rows were excluded and why; preserve the row for later resolution.

Counterexample

Calling a mean Brier score “calibration error” conflates calibration with other aspects of predictive accuracy.