Back to the log

Forecast Output Contract‌‌​⁠‌​‌⁠‌​⁠​​⁠​‌​​​‌​⁠⁠‍‍‌⁠​​​⁠‍‌‌‍​⁠‌​⁠‍‌⁠‍‌‍‌​‍‍‌‍‌​‍⁠‍​‌‍⁠‌‌‍

Compact human-readable result

Omit an unsupported field rather than filling it with invented detail. If no usable probability can be estimated, say what is missing. Give concise reasoning, not private internal deliberation.‌‌​⁠‌​‌⁠‌​⁠​​⁠​‌​​​‌​⁠⁠‍‍‌⁠​​​⁠‍‌‌‍​⁠‌​⁠‍‌⁠‍‌‍‌​‍‍‌‍‌​‍⁠‍​‌‍⁠‌‌‍

Machine-readable record

Use this illustrative shape when JSON is requested; null is allowed for unknown optional values. The schema is an original pack convention, not an Anthropic-required format.

{
  "question_id": "synthetic-launch-001",
  "question_version": 1,
  "question": "Will the specified release be generally available before the agreed deadline?",
  "forecast_as_of": "2027-01-01T00:00:00Z",
  "evidence_cutoff": "2027-01-01T00:00:00Z",
  "event_deadline": "2027-04-01T00:00:00Z",
  "probability": 0.63,
  "status": "open",
  "method": "reference class plus one likelihood update",
  "sources": [],
  "assumptions": ["Illustrative probabilities supplied by the user"],
  "sensitivity": {
    "kind": "assumption_envelope",
    "low": null,
    "high": null
  },
  "previous_forecast_id": null,
  "resolution": null,
  "scoring_convention": "binary_brier_0_1"
}

Complete synthetic example‌‌​⁠‌​‌⁠‌​⁠​​⁠​‌​​​‌​⁠⁠‍‍‌⁠​​​⁠‍‌‌‍​⁠‌​⁠‍‌⁠‍‌‍‌​‍‍‌‍‌​‍⁠‍​‌‍⁠‌‌‍

Given a fully defined launch question, 12/40 comparable attempts succeeded on time. Use the empirical 30% prior. A supplied independent signal has likelihood 0.8 if on-time and 0.2 if late, giving 0.24/(0.24+0.14)=63.16%.

Report approximately 63%, conditional on the supplied likelihoods and comparability. A second article reproducing that signal provides no extra likelihood. If the plausible LR spans 2 to 6, the prior 0.3 implies a sensitivity range of approximately 46.15% to 72%; call it an assumption envelope, not a confidence interval.

Potential update trigger: independently verified completion of a key dependency. Resolution and Brier scoring occur later. The entire case is synthetic and supplies no empirical evidence that the skill forecasts well.