Scope and input contract
Ask what event is being forecast, by when, from what information cutoff, and for what decision. If one missing detail materially changes the event, ask a focused question. Otherwise state provisional assumptions. Do not make the user choose every method before beginning.
For live requests, verify current status before forecasting: an event may already have occurred. For historical cutoffs, document possible model-memory leakage. For high-stakes use, present forecasting as one input to a decision rather than a replacement for qualified judgment.
1. Operationalize
Write an independently resolvable proposition. Define entity, threshold, event deadline, time zone, observation source, data revisions and void/unresolved rules. Preserve the version for later scoring. Separate “happens by” from “is announced by.”
2. Establish the outside view
Define comparable historical cases before choosing favorable examples. Record numerator, denominator, time window, missingness and transfer limitations. Use actual counts where available. If data are absent, label the prior judgmental and identify what would improve it.
Do not treat zero observed cases as proof of impossibility. Distinguish uncertainty in a rate from differences between past and future regimes.
3. Build a small model
| Use P(A and B)=P(A)P(B | A), or a disjoint exhaustive scenario mixture sum P(S)P(Y | S). Avoid multiplying marginal estimates without independence. Model alternative pathways. Compare to a direct outside-view estimate and investigate large disagreement rather than mechanically averaging two correlated approaches. |
4. Update on evidence
Trace new claims to their origin. Separate source reliability from likelihood. Where defensible, use: p_new=pa/(pa+(1-p)*b), with a=P(E|H), b=P(E|not H).
For subsequent evidence, condition likelihoods on what is already known. Duplicated reports do not create independent multipliers. With no justified numerical likelihood, use clearly labeled judgment or a scenario sensitivity, not fabricated statistical precision.
5. Challenge the conclusion
Construct a plausible way the favored outcome fails. Identify a diagnostic observation and estimate how the forecast changes under competing assumptions. Do not equate counterargument generation with evidence. Preserve asymmetry when one side is much better supported.
6. Aggregate only when needed
One analyst can produce a forecast without inventing a panel. If multiple estimates exist, align question versions and cutoffs, note shared information, and calculate a transparent mean.
A candidate logit pool is sigmoid(alpha*mean(logit(p_i))). Use alpha=1 without a justified alternative. A hypothetical alpha comparison is allowed when labeled. Do not apply a fixed multiplier because historical research found extremization useful in a different dataset.
7. Check uncertainty and communicate
Return the best-supported event probability and main reasons. An illustrative assumption envelope is a sensitivity range. A credible interval requires a specified probability model; a confidence interval requires a stated statistical procedure. Evidence quality is a separate assessment, not a second probability that the first probability is correct.
Keep the output proportional to the task. Do not create an elaborate report for a simple supplied-input calculation. State what facts are known, what was inferred and what was assumed.
8. Preserve and learn
Keep the original probability, timestamp, evidence cutoff and source references. Append later revisions rather than overwriting. Define future scoring before resolution: one-component binary Brier is (p-y)^2, lower is better.
Use matched lead times and question sets for comparisons. A low realized score on a few events is not proof of calibration or superforecaster status. Offer update triggers; do not claim recurring monitoring has started unless a scheduling capability was authorized and completed.
Research-to-agent transfer
This operational sequence is a synthesis. The human tournament literature supplies evidence about particular interventions in particular settings, not an empirically validated specification for this exact AI workflow. The mathematical identities can be tested exactly; improved real-world judgment requires prospective outcomes.