Back to the log

Diagnosing Calibration, Resolution and Sampling Limits: Worked Examples‌‌​⁠‌​‌⁠‌​⁠​​⁠​‌​​​‌⁠⁠​​‌‍‌​⁠‌⁠⁠‌‍‌‍‌⁠⁠⁠⁠‍‌‌​​‍‌​​​​⁠‍​‌​‌​​‌‌‌​

All numeric cases below are synthetic teaching examples, not empirical research findings.

Example: calibrated but uninformative

100 forecasts all at 0.3 with 30 events: REL=0, RES=0 and BS=UNC=0.21. The forecaster is calibrated on this sample but does not distinguish individual cases.‌‌​⁠‌​‌⁠‌​⁠​​⁠​‌​​​‌⁠⁠​​‌‍‌​⁠‌⁠⁠‌‍‌‍‌⁠⁠⁠⁠‍‌‌​​‍‌​​​​⁠‍​‌​‌​​‌‌‌​

Example: broad bins change the identity

Pairs: (0.1,0),(0.2,1),(0.8,1),(0.9,0). Raw BS=0.375. Use two bins: each has outcome frequency 0.5, mean predictions 0.15 and 0.85.‌‌​⁠‌​‌⁠‌​⁠​​⁠​‌​​​‌⁠⁠​​‌‍‌​⁠‌⁠⁠‌‍‌‍‌⁠⁠⁠⁠‍‌‌​​‍‌​​​​⁠‍​‌​‌​​‌‌‌​ REL=0.1225, RES=0, UNC=0.25. Their combination is 0.3725: the coarsened score. The raw-score residual is 0.0025; do not report an exact raw decomposition.‌‌​⁠‌​‌⁠‌​⁠​​⁠​‌​​​‌⁠⁠​​‌‍‌​⁠‌⁠⁠‌‍‌‍‌⁠⁠⁠⁠‍‌‌​​‍‌​​​​⁠‍​‌​‌​​‌‌‌​

Counterexample

“Eight of ten 80% predictions happened, so the model is proven calibrated” overstates a small-sample result and says nothing about other probability ranges.