Back to the log

Validating Forecaster Skill on Held-Out Outcomes: Method‌‌​⁠‌​‌⁠‌​⁠​​⁠​‌​​​‌⁠‌​‌‌⁠‍​⁠⁠‌‌​‌‍‍​⁠⁠​⁠‌‌​‍⁠​‌‍​​‌‍‌⁠⁠⁠‌‌‌​‌‍⁠

Why selection is insufficient

Even identical forecasters can produce an impressive winner through luck. Choosing the best score and reporting that same score overstates expected future performance. A candidate must face new outcomes under predeclared evaluation rules.

Count resolved questions, not probability submissions. Ten thousand daily updates to a handful of events provide fewer independent tests than ten thousand unrelated questions.‌‌​⁠‌​‌⁠‌​⁠​​⁠​‌​​​‌⁠‌​‌‌⁠‍​⁠⁠‌‌​‌‍‍​⁠⁠​⁠‌‌​‍⁠​‌‍​​‌‍‌⁠⁠⁠‌‌‌​‌‍⁠

Evaluation design

Define a minimum useful coverage policy based on the domain and intended claim, not an invented universal “50 questions certifies skill” rule. Report how many questions were eligible, attempted, resolved, excluded and shared across candidates.

Use identical lead-time snapshots or a declared time-weighted policy. Late forecasts have more information. Report both coverage and accuracy when a forecaster can abstain; excluding hard questions can otherwise manufacture superiority.‌‌​⁠‌​‌⁠‌​⁠​​⁠​‌​​​‌⁠‌​‌‌⁠‍​⁠⁠‌‌​‌‍‍​⁠⁠​⁠‌‌​‍⁠​‌‍​​‌‍‌⁠⁠⁠‌‌‌​‌‍⁠

Use a chronological selection/validation split. Keep variants of one election, company milestone or incident together when their outcomes share information. When training an AI aggregation method, ensure held-out outcomes were not in prompts, retrieval materials or parameter tuning. Historical LLM backtests may be contaminated by pretraining; prefer prospective tests.

Benchmarks

Include a simple base rate and a transparent aggregate when relevant. Compare on the same questions with the same resolution rules. A model that beats an untrained volunteer baseline has not necessarily beaten professional superforecasters.‌‌​⁠‌​‌⁠‌​⁠​​⁠​‌​​​‌⁠‌​‌‌⁠‍​⁠⁠‌‌​‌‍‍​⁠⁠​⁠‌‌​‍⁠​‌‍​​‌‍‌⁠⁠⁠‌‌‌​‌‍⁠

Report score differences, not only percentage improvement. If using relative improvement, specify the denominator and Brier normalization. Avoid statements such as “twice as accurate” without an operational definition.

Uncertainty and stability

Use question- or event-cluster-level resampling when appropriate. Model changes, question selection, dependence and time trends limit the interpretation of intervals. A ranking may be statistically unresolved even when a leaderboard orders participants.‌‌​⁠‌​‌⁠‌​⁠​​⁠​‌​​​‌⁠‌​‌‌⁠‍​⁠⁠‌‌​‌‍‍​⁠⁠​⁠‌‌​‍⁠​‌‍​​‌‍‌⁠⁠⁠‌‌‌​‌‍⁠

Shrink unstable estimates toward a baseline or use conservative equal weighting when history is sparse. Any shrinkage strength is a modeling choice requiring validation. Do not construct a “superforecaster personality” from demographic stereotypes.

Selection versus causal explanation

Persistent high performance supports a skill claim in the evaluated setting. It does not identify which trait or intervention caused it. Training, motivation, knowledge, question selection and timing can interact.‌‌​⁠‌​‌⁠‌​⁠​​⁠​‌​​​‌⁠‌​‌‌⁠‍​⁠⁠‌‌​‌‍‍​⁠⁠​⁠‌‌​‍⁠​‌‍​​‌‍‌⁠⁠⁠‌‌‌​‌‍⁠

The original GJP selected top performers and organized elite teams. That empirical strategy is not a transferable certification rule or proof that a prompted LLM has the same abilities.

Deliverable

Return cohort/context, selection rule, validation window, scores and benchmarks, coverage, uncertainty, evidence of persistence and limits. Use “top performer in this evaluation” when that is all the evidence supports.