All numeric cases below are synthetic teaching examples, not empirical research findings.
Example: insufficient evidence
Forecaster A scores 0.08 on ten easy questions; B scores 0.12 on one hundred different questions. Do not rank their underlying ability from those means alone. Request matched items/cutoffs or report the comparison as confounded.
Example: genuine holdout
Select a team on year-one data. Freeze the rule and assess on year-two questions without retroactive exclusions. Report year-two skill relative to prespecified competitors and its uncertainty.
Example: model leakage
A 2026 model answers questions that resolved in 2024. Hiding the result in the prompt does not prove the model never learned it. Label the exercise retrospective and potentially contaminated.
Counterexample
“Top 2% in this tiny synthetic exercise” does not establish professional Superforecaster status.