All numeric cases below are synthetic teaching examples, not empirical research findings.
Example: four binary questions
p=[0.8,0.6,0.2,0.1], y=[1,0,0,0]. Scores=[0.04,0.36,0.04,0.01]; mean=0.1125. A 0.5 baseline scores 0.25 on each: Brier skill=0.55. This four-question example is too small for a strong skill claim.
Example: categorical normalization
probabilities=[0.2,0.5,0.3], outcome index=1. Score=0.04+0.25+0.09=0.38.
Example: unresolved
A null outcome is excluded, not interpreted as zero. Report how many rows were excluded and why; preserve the row for later resolution.
Counterexample
Calling a mean Brier score “calibration error” conflates calibration with other aspects of predictive accuracy.