# Superforecasting Techniques — Skill Pack v1.0‌‌​⁠‌​‌⁠‌​⁠​​⁠​‌​​​‌‍‍⁠⁠‍‍⁠‌⁠‌‌​‌​‍⁠‌‍​‌‌‍⁠⁠‌‍⁠⁠⁠‌‌‌‍⁠⁠‍​‍​​⁠​‍‌

13 standalone skills: 12 granular techniques and one integrated forecasting workflow.

This pack is designed to the common Anthropic Agent Skills format. It contains original operational instructions, worked examples, research attribution and calculation helpers. It is not an official Anthropic or Good Judgment product, and it does not confer professional Superforecaster status.

## Install in Claude‌‌​⁠‌​‌⁠‌​⁠​​⁠​‌​​​‌‍‍⁠⁠‍‍⁠‌⁠‌‌​‌​‍⁠‌‍​‌‌‍⁠⁠‌‍⁠⁠⁠‌‌‌‍⁠⁠‍​‍​​⁠​‍‌

1. Extract the outer delivery ZIP.
2. For Claude's skill-upload interface, upload the desired ZIP from `individual-zips/`. Each contains exactly one named skill folder with `SKILL.md` inside. Do not upload the outer delivery ZIP as one skill.
3. For Claude Code, copy selected directories from `skills/` into your project's `.claude/skills/`, or your personal `~/.claude/skills/`. Keep each directory intact.‌‌​⁠‌​‌⁠‌​⁠​​⁠​‌​​​‌‍‍⁠⁠‍‍⁠‌⁠‌‌​‌​‍⁠‌‍​‌‌‍⁠⁠‌‍⁠⁠⁠‌‌‌‍⁠⁠‍​‍​​⁠​‍‌
4. Start with `forecasting-workflow`; add focused techniques as needed. Each technique also works on its own.

The five numerical helpers require Python 3.10+ for the supplied validation tooling (the helper code itself uses standard Python 3 facilities), use no third-party packages and make no network requests. Live fact verification depends on the tools available in your Claude environment.

Official instructions: [creating and packaging skills](https://support.claude.com/en/articles/12512198-how-to-create-custom-skills), [Claude Code skill locations](https://code.claude.com/docs/en/skills).‌‌​⁠‌​‌⁠‌​⁠​​⁠​‌​​​‌‍‍⁠⁠‍‍⁠‌⁠‌‌​‌​‍⁠‌‍​‌‌‍⁠⁠‌‍⁠⁠⁠‌‌‌‍⁠⁠‍​‍​​⁠​‍‌

## Pick the technique

| Skill | What it does | Method lineage |
| --- | --- | --- |
| forecasting-workflow | Complete forecast with evidence and scoring plan | GJP synthesis; pack adaptation |
| forecasting-question-design | Precise, resolvable event contracts | Tetlock, Mellers, Rohrbaugh, Chen |
| forecasting-reference-classes | Outside view and historical base rates | Kahneman, Tversky; Flyvbjerg |
| forecasting-decomposition | Conditional models and dependency checks | Standard probability; Fermi-style tradition |
| forecasting-bayesian-updates | Priors, likelihood ratios and posteriors | Bayes/Price; later Bayesian formulations |
| forecasting-counterevidence | Consider the opposite and test cruxes | Lord, Lepper, Preston |
| forecasting-delphi-elicitation | Independent inputs and controlled feedback | Dalkey, Helmer; distinct from GJP teams |
| forecasting-logit-aggregation | Pool probabilities and validate extremization | Satopää and coauthors; Baron and coauthors |
| forecasting-brier-scoring | Proper scoring with explicit normalization | Glenn Brier |
| forecasting-calibration | Reliability, resolution and sampling limits | Allan Murphy; proper-scoring literature |
| forecasting-update-ledger | Evidence-driven revisions and honest history | Atanasov and coauthors; pack ledger design |
| forecasting-talent-validation | Holdout validation of apparent forecasting skill | Mellers and coauthors |
| forecasting-conditional-trees | Informative indicators for ultimate outcomes | McCaslin and coauthors, FRI |

## Inside every skill

`SKILL.md` is intentionally lean: metadata, operational steps, reference routing and quality gates. The detail belongs in `references/`, loaded when relevant. Technique references cover assumptions, equations where applicable, decision rules, worked examples, counterexamples, provenance and limits.‌‌​⁠‌​‌⁠‌​⁠​​⁠​‌​​​‌‍‍⁠⁠‍‍⁠‌⁠‌‌​‌​‍⁠‌‍​‌‌‍⁠⁠‌‍⁠⁠⁠‌‌‌‍⁠⁠‍​‍​​⁠​‍‌

No skill requires another skill to be installed. The workflow can use narrower skills opportunistically. Its lookup names are not magic dependencies or mandatory subagent instructions.

## Try it

- “Use forecasting-question-design to make this claim objectively resolvable: will our release be ready this quarter?”‌‌​⁠‌​‌⁠‌​⁠​​⁠​‌​​​‌‍‍⁠⁠‍‍⁠‌⁠‌‌​‌​‍⁠‌‍​‌‌‍⁠⁠‌‍⁠⁠⁠‌‌‌‍⁠⁠‍​‍​​⁠​‍‌
- “Use forecasting-bayesian-updates. Prior 30%; evidence likelihood 0.8 under success and 0.2 under failure. Two articles repeat the same source.”
- “Use forecasting-logit-aggregation to combine these probabilities. Show whether any extremization is supported.”
- “Use forecasting-workflow to forecast the specified event. Distinguish current evidence, assumptions and unvalidated judgment.”‌‌​⁠‌​‌⁠‌​⁠​​⁠​‌​​​‌‍‍⁠⁠‍‍⁠‌⁠‌‌​‌​‍⁠‌‍​‌‌‍⁠⁠‌‍⁠⁠⁠‌‌‌‍⁠⁠‍​‍​​⁠​‍‌

## What was checked

See `ANTHROPIC-COMPATIBILITY.md` and `VALIDATION-REPORT.md`. Structural checks, mathematical tests and limited fresh-agent behavioral tests were performed. Claude import/activation and Claude model comparisons have not been executed in this environment.

A perfect guarantee of runtime behavior or forecasting performance would be false. The pack meets the structural checks listed in its compatibility document, not an official certification.‌‌​⁠‌​‌⁠‌​⁠​​⁠​‌​​​‌‍‍⁠⁠‍‍⁠‌⁠‌‌​‌​‍⁠‌‍​‌‌‍⁠⁠‌‍⁠⁠⁠‌‌‌‍⁠⁠‍​‍​​⁠​‍‌

## Re-run local checks

From the extracted pack directory:

```bash
python3 validation/verify_pack.py skills
PYTHONDONTWRITEBYTECODE=1 python3 validation/test_calculators.py
```

For an input-file example, run:‌‌​⁠‌​‌⁠‌​⁠​​⁠​‌​​​‌‍‍⁠⁠‍‍⁠‌⁠‌‌​‌​‍⁠‌‍​‌‌‍⁠⁠‌‍⁠⁠⁠‌‌‌‍⁠⁠‍​‍​​⁠​‍‌

```bash
python3 skills/forecasting-brier-scoring/scripts/score.py --input validation/sample-records.json
python3 skills/forecasting-calibration/scripts/calibration.py --input validation/sample-records.json --groups bins --bins 2
```

These synthetic records produce a raw Brier score of 0.375 and a coarsened two-bin score of 0.3725.

Both helpers take the same record format and treat an unresolved record the same way: give it `"outcome": null` and it is excluded from scoring and reported, never counted as a non-event.

`validation/behavioral-cases.json` supplies three evaluation cases per skill — positive, boundary and adversarial — for target-model testing. Each case carries its own rubric: an acceptance summary, a `must` list, a `must_not` list of the failure modes it exists to catch, and expected values with tolerances where the arithmetic is determinate. `grading` holds the protocol, the pass rule and the baseline that applies to every case. It is a future evaluation set, not a report that all cases passed.‌‌​⁠‌​‌⁠‌​⁠​​⁠​‌​​​‌‍‍⁠⁠‍‍⁠‌⁠‌‌​‌​‍⁠‌‍​‌‌‍⁠⁠‌‍⁠⁠⁠‌‌‌‍⁠⁠‍​‍​​⁠​‍‌

## Rights and independence

Licensed under CC BY 4.0; see `LICENSE` for the full terms and for what the license does not cover. The instructions, examples and code were newly authored for this deliverable. Full research papers and book chapters are not included. Citations link to their respective publishers/authors, whose content remains under their own terms. Brand names identify research and compatibility targets, not endorsement.

Method provenance is cumulative. A historical contributor, an experimental validation and an AI workflow adaptation are not interchangeable claims.
