Elicitation design
Classic Delphi uses iterative questionnaires and controlled feedback to elicit expert judgments. A forecast panel can adapt that structure, but open GJP discussion teams are not identical to Delphi. Label the design actually used.
Use actual respondents when reporting a human panel. The skill does not authorize contacting anyone; preparing a questionnaire is distinct from sending it. If responses are supplied, analyze them without claiming to have run interviews.
For a model-only panel record model family, version when known, prompt, retrieval corpus, timestamp, prior context and sampling configuration when available. Multiple draws can reveal variability but cannot be treated as statistically independent evidence without measurement.
First round
Give each participant the same target, cutoff and resolution rules. Collect:
- probability or predictive distribution;
- information sources and dates;
- strongest reason for and against;
- main uncertainty;
- evidence that would cause an update.
Avoid showing a prominent aggregate or senior participant’s estimate first. Separate a forecaster’s familiarity with a topic from their measured predictive skill.
Controlled feedback
Present the distribution honestly: count, median, spread, missing responses and relevant subgroups. Select arguments for diagnostic value, not eloquence. Anonymize when practical and consented; do not promise anonymity that the workflow cannot ensure.
Ask participants whether others supplied genuinely new facts, corrected an interpretation, or merely expressed stronger conviction. Preserve dissent and invite updates without making convergence the goal.
Finalization
A predeclared second round is a reasonable pack default, not a historically required number. More rounds should earn their cost. Record attrition and whether nonresponse correlates with difficult questions. Do not replace missing responses with the group mean without an explicit scoring/aggregation policy.
Report initial and final aggregates using the same rule to make changes interpretable. Keep the initial private estimates available for later analysis of social influence. Choose weights from validated past performance, not invented authority scores.
Limitations
Human collaboration can improve information exchange and also increase correlated errors. An apparently tighter final distribution does not establish improved accuracy. Compare outcomes prospectively and avoid causal claims from unrandomized team comparisons.
For a single-agent task, provide structured self-review and say no independent panel was available. Do not invent participant quotes, consensus or professional superforecaster credentials.