About this handbook
Supercast is a library for building forecasting agents that produce explicit probabilities, preserve their evidence and revisions, and learn from resolved outcomes. It supplies numerical methods and enforceable record contracts. The agent using it supplies domain knowledge, verified sources, model calls, and the judgment that connects observations to assumptions.
The handbook follows a forecast from question design to evaluation. You can read it in order or follow a narrower path:
| Your task | Start here |
|---|---|
| Produce a first binary forecast | A first forecast |
| Assess a proposed forecasting question | Question design |
| Update a probability from new evidence | Bayesian updates |
| Combine forecasters | Panels, then aggregation |
| Evaluate a track record | Scoring, calibration, validation |
| Choose which evidence to obtain | Indicators and decisions |
The companion Engineering Reference covers installation, Rust types, the JSON protocol, every operation, SQLite, errors, and release procedures. Both books are built into the local release and work without an external documentation service.
What a primitive can guarantee
A probability value can reject NaN or a number greater than one. A journal can reject a revision that cites the wrong predecessor. A bootstrap can keep observations from the same event cluster together. Those are properties of the software and its supplied inputs.
A primitive cannot prove that two articles have independent information, a reference class transfers to the present case, or a likelihood was honestly elicited. The handbook separates these modeling obligations from the computations so the agent can report both.
Supercast follows the thirteen method areas in the Superforecasting Skills Pack: question design, reference classes, decomposition, Bayesian updates, counterevidence, conditional indicators, elicitation, aggregation, update ledgers, scoring, calibration, validation, and the integrated workflow. Its software tests do not establish elite forecasting ability; that requires a prospective, appropriately controlled record.
Conventions used throughout
Probabilities are numbers between zero and one, not percentages. Timestamps are UTC Unix seconds. Binary Brier uses the one-component convention on a zero-to-one scale. Logarithms are natural logs; information gain is in nats. Examples involving product launches are synthetic and make no claim about real release outcomes.
Rust examples are compiled and run by the documentation gate. The JSON reference examples are executed against the actual command-line process. Both kinds of examples are maintained with the source code.
A first forecast
Suppose 12 of 40 comparable releases met their deadline. Before considering the current release, the empirical outside view is 0.30. An independently observed readiness signal is assessed as having probability 0.80 for an on-time release and 0.20 for a late one.
Bayes’ rule gives a posterior of approximately 0.632:
extern crate supercast;
use supercast::{Probability, bayes};
let prior = Probability::new(0.30)?;
let posterior = bayes::update(
prior,
Probability::new(0.80)?,
Probability::new(0.20)?,
)?;
assert!((posterior.get() - 12.0 / 19.0).abs() < 1e-12);
Ok::<(), supercast::Error>(())
The computation is conditional on the likelihood estimates and the relevance of the reference class. It does not turn those inputs into measured facts. A useful report might say:
My estimate is about 63%, starting from a 30% reference-class rate. The update assumes the readiness signal is four times as likely for an on-time release. A likelihood-ratio range of two to six gives an assumption envelope of roughly 46% to 72%.
The envelope is not a confidence interval. Its endpoints reflect alternative assumptions chosen by the forecaster.
Run the supplied workflow
From a source checkout:
CARGO_TARGET_DIR=target cargo build -p supercast-agent --locked
python3 examples/agent_demo.py
The demo creates a temporary SQLite database, registers a synthetic question, records a prior, applies a likelihood update, supplies an outcome, and scores the forecast. It restarts the process and retries the update to verify that the original receipt returns without creating another revision.
From a release bundle, use its included executable:
python3 examples/agent_demo.py --binary ./supercast-agent
To retain records, serve the JSON protocol with an explicit database path:
./supercast-agent --db forecasts.sqlite3
Write one JSON request per line. The process writes one response per nonblank input line. The command below is a calculation only; it does not append a forecast to the database.
{"version":1,"id":"first-calculation","command":{"op":"bayes","prior":0.3,"given_h":0.8,"given_not_h":0.2}}
Continue with question design before using a number like this in an actual forecasting record.
Design a resolvable question
A forecasting question needs a rule that a later adjudicator can apply without knowing which outcome the forecaster preferred. Define one proposition, its time window, the source that establishes the outcome, and what happens when the necessary observation is missing.
“Will the launch go well?” is not operational. “Will release X appear in the official general-availability archive strictly before timestamp T?” is closer: it identifies an event, a boundary, and an observation source. Whether that source is complete and accessible still needs investigation.
Freeze the contract
workflow::Question contains:
| Field | What to decide |
|---|---|
id, version | Stable target identity; increment the version when its meaning changes |
proposition | The exact claim being forecast |
opens_at, deadline | Allowed forecast window, including the start and excluding the deadline |
resolve_after | Earliest time an adjudication is permitted |
yes_rule, no_rule | Observable conditions establishing the two outcomes |
void_rule | Administrative exclusion conditions |
resolution_sources | Ordered source precedence, highest priority first |
The constructor checks nonempty fields, a positive version, and an ordered timeline. It does not parse the prose into a legal or causal predicate, check source accessibility, or prove that the yes/no rules form a complete partition. A source must be in the declared list for a resolution to be accepted; the adjudicator explains how precedence was applied.
Rehearse boundaries
Before forecasting, apply the proposed rules to hypothetical cases:
- The event happens exactly at the deadline. Does “before” exclude it?
- A source publishes after the deadline but documents an earlier event. Which timestamp controls the outcome?
- The product is renamed. Does the frozen identity still apply?
- The source stops publishing. Is the case unresolved, void, or resolvable from a fallback source?
- A report is revised. Does the first release or the revised value count?
Missing observations are not automatically evidence of “no.” Use unresolved status while the observation is insufficient. Void is an administrative state, not a third binary outcome to which an ordinary yes/no probability is assigned.
Forecasts and decisions
“Should we migrate the service?” is a decision. “Will the migration exceed its approved downtime window?” is a forecastable event. A small probability can justify mitigation if the loss is large; keep the event estimate separate from the utility calculation.
The storage layer freezes the question version. Editing a deadline or predicate means creating a new version and making that change visible. A correction to the arithmetic belongs in the forecast history rather than changing the question.
Start from an outside view
A reference class is a collection of past cases selected by a stated inclusion rule. It counters the tendency to focus entirely on the vivid details of the current case. Its value depends on the denominator: which attempts were included, which were omitted, and which have not resolved.
Record the class definition, inclusion rule, data source, observation window, successes, failures, unresolved cases, and transfer concerns. reference::ReferenceClass keeps these fields together and checks the declared counts for overflow.
Empirical and smoothed rates
If there are k successes and f failures, the empirical rate is k / (k + f). Unresolved cases are reported separately; they are not added as failures. No resolved cases means that an empirical rate is unavailable.
A Beta prior with parameters alpha and beta produces a posterior with parameters alpha + k and beta + f. Its mean is a next-case probability under the exchangeability assumption. The prior parameters are modeling choices, not observed extra cases.
extern crate supercast;
use supercast::bayes::BetaPrior;
let posterior = BetaPrior::new(1.0, 1.0)?.observe(12, 40)?;
assert!((posterior.mean().get() - 13.0 / 42.0).abs() < 1e-12);
assert!(posterior.variance() > 0.0);
Ok::<(), supercast::Error>(())
The reference_class JSON operation returns both the empirical rate and the smoothed posterior summary. With no resolved cases, empirical_rate is null while the supplied Beta prior still has a defined mean. That is a prior-driven estimate, not a historical success rate.
Three kinds of uncertainty
Distinguish sampling uncertainty from uncertainty about class selection and from structural change. More historical cases may reduce sampling noise without answering whether the class applies to the current setting. A product release by an experienced team may belong to a different class from a first release on a new platform.
When reasonable class definitions disagree, show their results separately. Overlapping classes are not independent evidence sources. Averaging them without accounting for overlap can make the record appear broader than it is.
Skewed quantities
For cost, duration, and capacity, retain the distribution’s tail. An empirical quantile is often more useful than a symmetric percentage band. Supercast uses linear interpolation at index (n - 1) * q, commonly called type 7:
extern crate supercast;
use supercast::{Probability, reference::quantile};
let mut duration_ratios = [1.0, 1.1, 1.2, 1.5, 4.0];
let p80 = quantile(&mut duration_ratios, Probability::new(0.8)?)?;
assert!((p80 - 2.0).abs() < 1e-12);
Ok::<(), supercast::Error>(())
The function sorts the supplied slice in place. P80 is a quantile of that empirical distribution, not an 80% confidence interval around its mean. Sampling design and transfer concerns still belong in the report.
Bayesian evidence updates
An update starts with a prior probability p and two likelihoods: a = P(E | H) and b = P(E | not H). The posterior is:
posterior = p*a / (p*a + (1-p)*b)
The ratio a/b describes how diagnostic the observation is. It is not the same as the source’s credibility, and neither likelihood is P(H | E). Ask the two likelihood questions separately before looking at the desired posterior.
Numerically stable computation
bayes::update computes nondegenerate cases in log-odds space. This avoids underflow from multiplying tiny probabilities. Exact impossible evidence is an error: if neither hypothesis can produce the observation under the supplied prior, the model has failed.
extern crate supercast;
use supercast::{Probability as P, bayes};
let p = P::new(0.5)?;
let tiny = P::new(1e-310)?;
assert_eq!(bayes::update(p, tiny, tiny)?, p);
assert!(bayes::update(p, P::ZERO, P::ZERO).is_err());
Ok::<(), supercast::Error>(())
A dogmatic prior of zero or one generally cannot learn from finite, nonzero likelihood ratios. Ordinary uncertain forecasts should avoid those endpoints unless the outcome is logically determined under the model. Supercast permits endpoints as valid probabilities but does not quietly replace them with an epsilon.
Condition on what was already known
For a second observation, use P(E2 | H, E1) and P(E2 | not H, E1). Multiplying two marginal likelihood ratios implicitly assumes conditional independence. Two articles copying the same press release are not two observations.
The in-memory UpdateSession accepts prior-origin IDs and rejects a reused evidence origin. The durable update_forecast operation collects origins from previous forecast evidence and stores every likelihood pair with the resulting revision. Repeated evidence in a batch and origins from earlier revisions are rejected.
An origin ID refers to the underlying observation. A continuing source can produce multiple distinct observations; give those distinct origin IDs and condition the likelihoods on earlier information. Renaming an old observation to apply it twice defeats the modeling contract even if the strings pass validation.
Sensitivity and qualitative updates
When likelihoods are judgmental, provide a range of defensible inputs. For a 0.30 prior, LR 2 yields about 0.462 and LR 6 yields 0.72. bayes::sensitivity returns these endpoints as an assumption envelope.
If the evidence does not support numeric likelihoods, record a qualitative assessment or a manual judgmental revision with that limitation. Do not back-solve a likelihood ratio merely to disguise a preferred final probability as a measured update.
Silence is evidence only when the observation process makes it diagnostic. An empty incomplete news feed does not establish that an incident did not happen.
Decompose without inventing independence
Decomposition makes assumptions inspectable. It improves a forecast only when the pieces are informative and their dependence is handled correctly.
Conjunctions
The chain rule is P(A and B) = P(A) * P(B | A). In a longer chain, each factor conditions on the relevant preceding events. conditional_chain multiplies supplied factors; its name makes the required interpretation explicit.
extern crate supercast;
use supercast::{Probability as P, decompose};
let launch = decompose::conditional_chain(&[
P::new(0.8)?, // dependency completes
P::new(0.6)?, // launch succeeds given dependency completion
])?;
assert!((launch.get() - 0.48).abs() < 1e-12);
Ok::<(), supercast::Error>(())
Do not supply two marginal estimates unless independence is actually justified. A common cause, such as a staffing shortage, can affect both components.
Unions and bounds
For two events, P(A or B) = P(A) + P(B) - P(A and B). If overlap is unknown, report bounds:
intersection: max(0, pA + pB - 1) .. min(pA, pB)
union: max(pA, pB) .. min(1, pA + pB)
decompose::union rejects a supplied intersection outside those bounds. The JSON union command returns both bounds and an exact union only if an intersection was supplied; otherwise its union field is null.
Scenario mixtures
A weighted mixture assumes disjoint, exhaustive scenarios. Weights must sum to one within the documented floating-point tolerance. If scenarios overlap, the weighted sum double-counts; if a pathway is missing, add an explicit “other” scenario.
For an indicator C, a simple mixture is P(U) = P(C)*P(U|C) + P(not C)*P(U|not C). The indicator primitive adds a coherence check and expected information gain to this model.
Counts and deadline events
A product of dimensional factors may estimate an expected count rather than the probability of any occurrence. Under a justified Poisson model with constant rate lambda, the probability of at least one event in duration t is 1 - exp(-lambda*t).
event_probability evaluates this expression with expm1 for precision at small rates. Its rate and duration must use compatible units. For a remaining-window forecast, the model conditions on survival to the window’s start. Clustering, bursts, or changing hazards can make a constant-hazard model unsuitable.
Stop decomposing when another factor would add unsupported precision without changing the conclusion or decision. More factors can multiply guesses as easily as they can combine evidence.
Indicators and information gain
An indicator is a resolvable event that changes a forecast for the ultimate outcome. Elicit three quantities: its chance c, the ultimate outcome’s probability if it occurs a, and that probability if it does not b.
The implied prior is u = c*a + (1-c)*b. Compare u to the stated prior before ranking the indicator. A dramatic conditional branch does not by itself make a coherent model.
extern crate supercast;
use supercast::{Probability as P, decompose::Indicator};
let indicator = Indicator {
chance: P::new(0.5)?,
if_yes: P::new(0.8)?,
if_no: P::new(0.3)?,
};
indicator.check(P::new(0.55)?, 1e-12)?;
assert!(indicator.information_gain() > 0.13);
assert_eq!(indicator.sensitivity(), [0.5, 0.5, 0.5]);
Ok::<(), supercast::Error>(())
Indicator::check rejects an incoherent stated prior using an explicit tolerance. It does not silently repair the beliefs. If the analyst prefers the prior, revisit the branch assumptions; if the branches are better supported, explain a prior revision.
Entropy reduction
Bernoulli entropy is h(p) = -p*ln(p) - (1-p)*ln(1-p), with endpoint entropy zero. Expected information gain is:
IG = h(u) - c*h(a) - (1-c)*h(b)
The result is in nats. An indicator with identical branches provides zero information. A perfect revelation of the target removes all prior entropy. A striking but extremely unlikely branch may have little expected value.
The returned sensitivity vector is ordered (chance, if_yes, if_no) and equals (a-b, c, 1-c). It identifies which elicited assumptions have the greatest local effect on the implied prior.
Information and decisions
Entropy reduction is not the same as decision value. An indicator can change a probability substantially while leaving the best action unchanged. Use value of information when deciding whether research is worth its cost.
These branches express conditional association. They do not establish the causal effect of intervening on the indicator. An event that predicts a product delay is not necessarily a lever that prevents it.
The current primitive is a two-branch model. More elaborate trees can be assembled from coherent mixtures, but their branch conditions and shared causes must remain explicit. The library does not infer a causal graph from a sequence of indicators.
Search for counterevidence
Once a forecast has a plausible explanation, ask what would make its favored conclusion false. The purpose is to test important assumptions, not to add a balancing sentence or mechanically move the estimate toward 50%.
Start with a contrary pathway that could materially change the outcome. For a launch forecast, the strongest opposing pathway might be that a readiness demo hides an unresolved integration dependency. Turn that concern into a falsifiable test: an independent integration run, an observable dependency milestone, or an operational readiness measure.
Record a testable challenge
workflow::Challenge records an alternative, a falsifiable test, the evidence needed, and an Indicator model. It is a planning record. It does not represent evidence that the alternative occurred, and creating it does not run a search or schedule a task.
| Element | Example |
|---|---|
| Alternative | The release is delayed by an integration defect |
| Falsifiable test | A specified end-to-end test fails before the review date |
| Evidence needed | An independently attributable test result with an observation date |
| Probability effect | Explicit conditional branches for pass and failure |
A useful challenge can reveal a missing scenario or change the reference class. It can also leave the probability unchanged after the concern is examined. Preserve the result and why it did or did not alter the estimate.
Avoid fabricated balance
A failure story is not an observation. A model-generated opposing narrative is not an independent expert. A second phrasing of the same source is not a new origin. Keep hypothetical pathways separate from dated, attributable evidence in the forecast record.
Prioritize cruxes with both material probability sensitivity and feasible tests. A remote speculation with no observable consequence may consume effort without improving the forecast. Indicators help quantify sensitivity; decision value helps determine whether obtaining the observation could change an action.
Stop deliberately
Stop when the highest-impact plausible contrary pathway has been examined to the available depth, when further sources repeat the same origin, or when an access limit prevents verification. Report the remaining uncertainty instead of inventing a countervailing fact.
The next review trigger should identify what new observation would justify reopening the issue. “Review again” is less useful than “review after independently verified completion of dependency Y.”
Elicit a forecast panel
A panel can combine information that no individual has alone. It can also amplify shared errors, social pressure, or a common source. A controlled elicitation process preserves initial judgments and reveals why respondents revise.
Collect private first estimates
Give every participant the same question version, evidence cutoff, and resolution rules. Collect a probability, strongest supporting reason, strongest opposing reason, main uncertainty, and observation that would cause an update. Avoid showing a prominent aggregate or senior participant’s estimate before this first round.
The Elicitation record preserves respondent identity, human/model metadata, an optional probability, and supporting/opposing reasons. None or JSON null means no response; it is never replaced automatically with 0.5 or the group mean.
The Respondent type distinguishes a human pseudonym from model family/version, prompt reference, and corpus reference. These metadata are supplied by the caller; they do not verify that a person participated or that a model was independently trained.
Give controlled feedback
summarize_panel returns invited, responded and missing counts, median, minimum and maximum. Duplicate respondent IDs and a panel with no probabilities are rejected. Its range is disagreement, not a confidence interval.
Present useful arguments and evidence origins alongside the distribution. Ask whether another participant supplied a new fact, corrected an interpretation, or merely expressed stronger conviction. Preserve dissent when it rests on an unresolved assumption.
A reasonable small exercise might use two predeclared rounds. More rounds should justify their cost rather than treating convergence as the goal. Compare initial and final aggregates using the same rule.
Independence and model panels
Repeated calls to one model can explore variability and alternative decompositions. They are not independent human experts, and their similar answers do not establish calibrated certainty. Shared retrieval corpora and prompt context can make errors strongly dependent.
Keep respondent-level records outside the summary. They are necessary for later analysis of influence, attrition, and performance. The library summarizes supplied responses; it neither contacts respondents nor fabricates interviews.
Use equal weights unless relevant held-out performance supports another policy. Familiarity, confidence, and prestige are not interchangeable with measured forecasting skill. Continue with aggregation once the elicitation record is clear.
Aggregate and recalibrate
Arithmetic pooling averages probabilities. Logit pooling averages log odds and transforms the result back. These are different algorithms and can give noticeably different answers.
For interior probabilities, logit(p) = ln(p) - ln(1-p). With nonnegative normalized weights w and declared adjustment alpha:
z = sum(w_i * logit(p_i))
aggregate = sigmoid(alpha * z)
Alpha one leaves the pooled log odds unchanged; values above one extremize, values between zero and one shrink toward 0.5, and zero returns 0.5. This is not extremization of the arithmetic mean.
extern crate supercast;
use supercast::{Probability as P, aggregate::{pool, WeightedForecast}};
let forecasts = [
WeightedForecast { probability: P::new(0.2)?, weight: 1.0 },
WeightedForecast { probability: P::new(0.8)?, weight: 1.0 },
];
let result = pool(&forecasts, 1.0)?;
assert!((result.logit.get() - 0.5).abs() < 1e-12);
assert_eq!(result.arithmetic, P::HALF);
Ok::<(), supercast::Error>(())
Numerical and participation policy
pool reports arithmetic and logit results, alpha, and positive-weight contributor count. It rescales weight magnitudes before summing to avoid overflow. Zero-weight rows do not contribute. All weights must be finite and nonnegative, with at least one positive weight.
Logit requires probabilities strictly inside zero and one. Positive-weight endpoints are rejected; there is no implicit clipping. If an application adopts an epsilon policy, it must state and validate that policy separately. Missing panel responses should be excluded with a coverage report, not invented as neutral forecasts.
When adjustment is justified
Extremization can help when participants bring sufficiently distinct information and the initial aggregate is underconfident. Correlated or overconfident inputs can make it harmful. A coefficient that helped another study or domain is not a universal constant.
aggregate::recalibrate applies sigmoid(intercept + slope*logit(p)) with finite intercept and nonnegative slope. It does not fit the parameters. evaluation::select_calibration selects from a declared candidate grid using training outcomes, then reports performance on a separate temporal/cluster holdout.
Include the identity candidate (0,1) when leaving forecasts unchanged is a legitimate option. Once the holdout is inspected, do not keep changing the grid against the same outcomes and call the final result held out. See validation for the evaluation contract.
Keep an auditable ledger
A forecast is a dated belief about a frozen question, supported by a record of what was known and how the estimate changed. Preserve each revision rather than overwriting the latest probability.
Supercast’s core Ledger owns one validated question version. Its public API exposes records for reading and accepts new forecasts or resolutions through validation. The SQLite adapter persists the same lifecycle as immutable events and replays the rules when loading them.
Revision fields
A forecast records its ID, as-of time, evidence cutoff, probability, predecessor ID, change type, method, rationale, assumptions, next review trigger, and evidence references. The initial forecast has no predecessor and uses initial. Later revisions name the current forecast ID.
Forecast times must increase strictly, cutoffs cannot regress, and no forecast can be submitted at or after the frozen deadline. A review_unchanged record must preserve the probability. Corrections append a new record and keep the old erroneous estimate visible.
The ledger does not require a probability to move after every review. New evidence can confirm a prior judgment or provide no diagnostic information. The rationale explains the assessment; small cosmetic changes do not make an agent more responsive.
Evidence timing
An evidence reference contains an ID, origin ID, locator, claim, publication time, and optional observation time. Both known times must respect the forecast’s evidence cutoff. A later article about an earlier event was not necessarily available at the earlier time.
The timestamp fields are caller assertions. The database also records local commit time, but historical imports remain possible. A credible prospective evaluation needs additional controls that freeze the cohort and enforce contemporaneous recording.
The source locator may refer to a supplied file rather than a URL. The library does not fetch it, prove its authenticity, or execute any instructions found inside it. The consuming application should treat source content as evidence data.
Resolution lifecycle
An adjudication must occur at or after resolve_after, name a declared source, and include a rationale. Outcomes are yes, no, unresolved, or void with a reason.
An unresolved adjudication may be followed by a later one. Yes, no, and void are terminal. Once adjudication begins, the ledger rejects additional forecasts. A delayed source release is handled as unresolved rather than quietly treating the absence as “no.”
Safe retries
Durable writes use the request ID as an idempotency key. After a lost response, resend the same decoded command and ID. The stored receipt returns without another event. Reusing the ID for a different command is a conflict.
Reading the current ledger before proposing a revision is not enough to prevent a race; another writer can append in between. The adapter rechecks the predecessor under the SQLite write transaction. If your revision is stale, read the new history and reconsider the update instead of merely replacing its predecessor string.
Score predictions properly
A proper scoring rule rewards an honest predictive distribution in expectation. A realized score evaluates a particular prediction/outcome pair; it does not reveal a hidden “true probability” from one outcome.
Use a scoring convention before seeing the outcomes, keep it consistent across comparisons, and report which unresolved or void questions were excluded.
Binary Brier and log loss
Binary Brier is (p-y)^2, where y is zero or one. Supercast uses the one-component zero-to-one convention. Lower is better.
extern crate supercast;
use supercast::{Probability as P, score};
let p = P::new(0.3)?;
assert!((score::brier(p, true) - 0.49).abs() < 1e-12);
assert!(score::log_loss(P::ZERO, true).is_infinite());
assert!(score::brier_skill(0.2, 0.0).is_err());
Ok::<(), supercast::Error>(())
Log loss is -ln(p) for yes and -ln(1-p) for no. A confidently wrong endpoint incurs positive infinity. JSON cannot encode infinity as a number, so binary_score returns a tagged value: positive_infinity, or finite with a numeric value.
Brier skill is 1 - model_score / reference_score. A zero reference score makes the ratio undefined; an unrepresentable ratio is also an error. Absolute scores and their difference can still be reported when relative skill is unavailable.
Categorical and ordered outcomes
A Categorical distribution has at least two entries and sums to one within tolerance. Categorical Brier sums squared errors across categories and ranges from zero to two. A two-category vector therefore produces twice the one-component binary score.
The ranked probability score instead compares cumulative probabilities over ordered categories. Supercast returns the unnormalized sum over K-1 thresholds, on a zero-to-K-1 scale. Use it for ordered categories such as low/medium/high, not unrelated labels.
Continuous outcomes
The empirical continuous ranked probability score, CRPS, evaluates a sample-based predictive distribution in the outcome’s units:
CRPS = mean(|sample - outcome|) - 0.5 * mean(|sample_i - sample_j|)
The implementation sorts the caller’s sample slice and integrates squared empirical-CDF error over adjacent gaps in O(n log n), an equivalent nonnegative formula that avoids a quadratic distance matrix and cancellation between large expectations. It preserves finite answers when some unscaled intermediate distances would overflow; truly unrepresentable results are errors.
extern crate supercast;
use supercast::score::empirical_crps;
let mut samples = [1.0, 2.0, 3.0];
let score = empirical_crps(&mut samples, 2.0)?;
assert!((score - 2.0 / 9.0).abs() < 1e-12);
Ok::<(), supercast::Error>(())
Repeated forecasts
Do not average every post as if each were an independent question. A frequent poster would otherwise receive different weighting. grid_brier carries the latest available forecast forward over a prespecified grid and averages within one question. Average those question scores according to the frozen cohort policy.
The grid must be strictly increasing, and a forecast must exist at its first time. The function does not fill a missing start with 0.5. Align grid times with the event window and evaluation protocol before calling it.
Diagnose calibration
Calibration asks whether events assigned probability near p occur near frequency p. Resolution asks whether forecasts distinguish groups with different outcome frequencies. Sharpness describes how concentrated the predictions are without reference to outcomes.
A constant base-rate predictor can be calibrated and still fail to distinguish individual cases. A sharp forecaster can be confidently wrong. Do not use either sharpness or ensemble agreement as a substitute for outcome-based evaluation.
Reliability tables
calibration::diagnose accepts one frozen binary forecast per question. Question IDs must be nonempty and unique. Unresolved outcomes are excluded and counted. It groups probabilities either exactly or into an explicitly chosen number of equal-width bins.
Each nonempty bin reports its count, mean forecast, and observed frequency. Empty bins are not evidence of perfect calibration. Choose the grouping before inspecting results, then report reasonable sensitivity to bin choices where appropriate.
extern crate supercast;
use supercast::{Probability as P, calibration::{Record, Grouping, diagnose}};
let rows = vec![
Record { question_id: "a".into(), probability: P::new(0.2)?, outcome: Some(false) },
Record { question_id: "b".into(), probability: P::new(0.8)?, outcome: Some(true) },
Record { question_id: "c".into(), probability: P::new(0.7)?, outcome: None },
];
let report = diagnose(&rows, Grouping::Exact)?;
assert_eq!(report.resolved, 2);
assert_eq!(report.excluded_unresolved, 1);
assert!(report.binning_residual.abs() < 1e-12);
Ok::<(), supercast::Error>(())
The decomposition and its trap
For bin k, let f be its mean forecast, o its outcome frequency, and w its fraction of all scored questions. With overall outcome frequency o-bar:
reliability = sum(w * (f-o)^2)
resolution = sum(w * (o-o_bar)^2)
uncertainty = o_bar * (1-o_bar)
For exact groups of identical probabilities, Brier equals reliability minus resolution plus uncertainty. For wider bins, that expression corresponds to the coarsened forecast system, where each original probability is replaced by its bin mean. It need not equal the raw Brier score.
Supercast therefore reports raw Brier, coarsened Brier, all three terms, and the binning residual raw - coarsened. Treating a nonzero residual as a bug in the forecast system would confuse a grouping transformation with the original predictions.
Sampling limits
A single 70% forecast that resolves yes is not proof of a stable 30-point bias. Small groups have noisy frequencies. Related questions can share outcomes or common causes, so an independent-Bernoulli interval may overstate precision.
The diagnostic does not fit a correction or manufacture intervals. Use cluster-level validation for a paired score comparison, and separate calibration training from evaluation. Coverage also matters: slow-resolving open questions may differ systematically from the resolved subset.
Validate forecasting skill
A winning score on a selected sample is not sufficient evidence of persistent skill. Many equally capable forecasters can produce an impressive winner by chance. Select a method on one set of outcomes and evaluate it on new, appropriately separated questions.
Freeze the comparison
Declare eligible questions, lead times or snapshots, resolution rules, baselines, weights, and exclusion rules before evaluation. Record attempted and missing questions as well as scored ones. Selective abstention can make a leaderboard misleading if coverage is hidden.
evaluation::Case pairs a model forecast and baseline on the same question, time, outcome, and event cluster. compare reports matched means and their difference. It rejects repeated question IDs and forecasts at or after resolution.
validate_holdout also requires all training labels by the cutoff, all holdout forecasts after it, and no event cluster shared across the split. The supplied IDs and timestamps must reflect the real information structure. The function cannot detect undisclosed pretraining or retrieval contamination.
Uncertainty with event clusters
cluster_bootstrap resamples entire declared event clusters with replacement, using the same draw for model and baseline. Questions inside each selected cluster remain together. The statistic remains the question-weighted paired mean difference.
It returns a percentile interval, bootstrap standard error, counts, confidence level, seed, replicate count, and a degeneracy flag. The seed makes an analysis reproducible. At least two clusters are computationally required; substantially more may be needed for useful inference.
extern crate supercast;
use supercast::{Probability as P, evaluation::{Case, cluster_bootstrap}};
let cases = [
Case { question_id: "a".into(), cluster_id: "a".into(), forecast_at: 1,
resolved_at: 5, model: P::ONE, baseline: P::HALF, outcome: true },
Case { question_id: "b".into(), cluster_id: "b".into(), forecast_at: 1,
resolved_at: 5, model: P::ZERO, baseline: P::HALF, outcome: true },
];
let report = cluster_bootstrap(&cases, 1000, 0.95, 7)?;
assert_eq!(report.clusters, 2);
assert_eq!(report.difference, 0.25); // positive means the model did worse
Ok::<(), supercast::Error>(())
This interval is about a performance statistic under sampling assumptions, not about the true probability of one event. A degenerate interval does not prove an absence of model uncertainty. Few independent clusters, biased selection, and temporal change can make percentile coverage unreliable.
Select a calibration candidate
select_calibration takes a prespecified grid of intercept/slope pairs. It picks the lowest training Brier, resolving exact ties by candidate order, then reports raw, calibrated, and baseline holdout scores. The chosen candidate remains chosen even if it performs worse on the holdout.
Changing the grid after inspecting that result consumes the holdout. Use a new evaluation period for a subsequent selection cycle. The library limits candidate and simulation workloads, but a computational limit is not a statistical adequacy rule.
Database evaluation
The evaluate JSON command reads a declared cohort at common forecast and resolution cutoffs. It reports which cases had no forecast, were unresolved, were void, or fell outside the allowed snapshot window. With no scored cases it returns a null summary rather than a fabricated zero score.
That command uses declared timestamps. Imported historical forecasts are not made prospective by assigning them earlier dates. A surrounding prospective process must freeze the plan and preserve real recording times. Claims of elite skill need that external evidence, not merely a passed software gate.
Choose research and actions
A probability is an input to a decision, not the decision itself. The preferred action depends on the consequences of each outcome and on the cost of learning more.
decision::Action specifies utility if the event occurs and if it does not. Expected utility is p*if_yes + (1-p)*if_no. best_action returns the zero-based index and utility of the highest-value action. Exact ties select the first action in the supplied order.
extern crate supercast;
use supercast::{Probability as P, decision::{Action, best_action}};
let actions = [
Action { if_yes: 10.0, if_no: -10.0 },
Action { if_yes: 0.0, if_no: 0.0 },
];
let (index, utility) = best_action(&actions, P::new(0.8)?)?;
assert_eq!(index, 0);
assert!((utility - 6.0).abs() < 1e-12);
Ok::<(), supercast::Error>(())
Utilities must share a scale. They can represent money, a scored objective, or another explicitly justified utility measure. The library does not determine an organization’s values or infer utilities from a probability.
Expected value of information
An indicator can be observed before choosing an action. The value of that observation is the expected best utility after observing it, minus the best utility available now. Subtract research cost to obtain net value.
extern crate supercast;
use supercast::{Probability as P, decision::{Action, value_of_information}, decompose::Indicator};
let actions = [
Action { if_yes: 10.0, if_no: -10.0 },
Action { if_yes: 0.0, if_no: 0.0 },
];
let perfect_signal = Indicator { chance: P::HALF, if_yes: P::ONE, if_no: P::ZERO };
assert_eq!(value_of_information(&actions, perfect_signal, 1.0)?, 4.0);
Ok::<(), supercast::Error>(())
The calculation uses the indicator’s implied prior. Its branches must therefore represent a coherent joint model. Research cost is nonnegative and uses the same utility units as the actions. A negative net value can be the correct answer.
Rank useful work
Use probability sensitivity to identify important assumptions, information gain to measure expected belief change, and decision value to assess whether resolving uncertainty could change an action. They answer different questions.
An indicator can carry information while having zero decision value because the same action is optimal in both branches. Conversely, a modest probability change near a decision threshold can be valuable. Do not rank research by how dramatic its narrative sounds.
This primitive assumes the observation arrives in time to act and has the supplied reliability. It does not model the causal effect of acquiring the information, a changing action set, or the opportunity cost of delay beyond the declared research cost.
A complete synthetic workflow
The source distribution includes a runnable Rust example that ties the primitives together. It creates a frozen release question, records a 30% prior, updates from an attributable readiness signal, reports a likelihood-ratio envelope, and supplies a synthetic yes outcome.
//! Synthetic workflow: these inputs are illustrative, not measured evidence.
use supercast::{Probability as P, Result, bayes, decompose::Indicator, workflow::*};
fn main() -> Result<()> {
let question = Question {
id: "synthetic-launch-001".into(),
version: 1,
proposition: "Will release X be generally available before the frozen deadline?".into(),
opens_at: 1_800_000_000,
deadline: 1_810_000_000,
resolve_after: 1_810_086_400,
yes_rule: "Official archive records general availability strictly before deadline".into(),
no_rule: "Complete official archive establishes no qualifying release by deadline".into(),
void_rule: "Release identity becomes undefined or archive permanently inaccessible".into(),
resolution_sources: vec!["official release archive".into()],
};
let mut ledger = Ledger::new(question)?;
let prior = P::new(12.0 / 40.0)?;
ledger.append(Forecast {
id: "f1".into(),
as_of: 1_800_000_000,
evidence_cutoff: 1_800_000_000,
probability: prior,
previous_id: None,
change: Change::Initial,
method: "empirical reference class".into(),
rationale: "12 of 40 comparable attempts succeeded".into(),
assumptions: vec!["Supplied cases are comparable".into()],
next_review_trigger: "independent readiness signal".into(),
evidence: vec![],
})?;
let evidence = Evidence {
id: "readiness-1".into(),
origin_id: "original-test-1".into(),
locator: "supplied:synthetic-readiness-test".into(),
claim: "Independent readiness test passed".into(),
published_at: 1_800_001_000,
observed_at: Some(1_800_000_900),
};
let mut updates = UpdateSession::new(prior, 1_800_001_000, &[])?;
let posterior = updates.apply(&evidence, P::new(0.8)?, P::new(0.2)?)?;
ledger.append(Forecast {
id: "f2".into(),
as_of: 1_800_001_000,
evidence_cutoff: 1_800_001_000,
probability: posterior,
previous_id: Some("f1".into()),
change: Change::Evidence,
method: "Bayesian likelihood update".into(),
rationale: "Readiness test is more likely for on-time releases".into(),
assumptions: vec!["Likelihoods 0.8 and 0.2 are illustrative judgments".into()],
next_review_trigger: "independently verified dependency completion".into(),
evidence: vec![evidence],
})?;
let (low, high) = bayes::sensitivity(prior, 2.0, 6.0)?;
let indicator = Indicator {
chance: P::new(0.5)?,
if_yes: P::new(0.8)?,
if_no: P::new(0.3)?,
};
println!(
"Synthetic forecast: {:.2}% -> {:.2}%",
100.0 * prior.get(),
100.0 * posterior.get()
);
println!(
"LR 2..6 assumption envelope: {:.2}%..{:.2}%",
100.0 * low.get(),
100.0 * high.get()
);
println!(
"Candidate indicator's separate model: prior {:.2}%, information {:.4} nats",
100.0 * indicator.implied_prior().get(),
indicator.information_gain()
);
ledger.resolve(Resolution {
at: 1_810_086_400,
source: "official release archive".into(),
rationale: "Synthetic archive confirms qualifying release".into(),
outcome: Outcome::Yes,
})?;
println!(
"Resolved yes; final Brier: {:.6}",
supercast::score::brier(posterior, true)
);
println!("Preserved {} forecast revisions", ledger.forecasts().len());
Ok(())
}
extern crate supercast;
The final posterior is approximately 63.16%, and the one-component Brier score after yes is approximately 0.135734. The 30% prior alone would score 0.49 for that outcome. This one synthetic case demonstrates arithmetic and workflow behavior; it says nothing about real forecasting advantage.
What is auditable
The initial forecast remains in the ledger. The revision points to it, carries a later as-of time and evidence cutoff, identifies the new observation’s origin, and records the modeling assumptions. The resolution names a source frozen in the question contract.
The candidate indicator shown near the end uses a separate hypothetical model. Its implied prior is 55%; it is not quietly substituted for the release forecast. A real use of that indicator would first reconcile its branches with the current forecast.
Run the durable version
The JSON fixture examples/agent-workflow.jsonl contains seven requests: question creation, initial forecast, Bayesian revision, sensitivity, resolution, evaluation, and ledger retrieval. The Python demonstration sends them through the actual process using a temporary database.
It then starts a new process, repeats the Bayesian request with the same ID, and reads the ledger. Both the original receipt and the full history must match. This covers an operational failure mode that a standalone arithmetic example cannot: retrying a committed write after the caller loses its response.
The Engineering Reference documents those operations individually, with setup requests and outputs produced by executing the examples. Use that reference when adapting this workflow to an agent tool loop.
Glossary
| Term | Meaning in Supercast |
|---|---|
| As-of time | The declared timestamp of a forecast revision |
| Base rate | An outside-view frequency or predictive probability from comparable cases |
| Brier score | Squared probability error under the stated binary or categorical convention |
| Calibration | Agreement between predicted probabilities and outcome frequencies |
| Cluster | A group of related questions kept together for splitting or resampling |
| Coarsened Brier | Score after replacing each forecast by its bin’s mean probability |
| Conditional likelihood | Probability of evidence under a hypothesis and earlier evidence |
| CRPS | Continuous ranked probability score; an outcome-unit measure of predictive distribution error |
| Cutoff | Latest allowed evidence or outcome availability time for an analysis |
| Degenerate bootstrap | A resampled distribution with no observed variation |
| Evidence origin | Identity of the underlying observation, independent of its republications |
| Holdout | Data excluded from selection and used to evaluate the selected method |
| Idempotency key | Request ID that identifies an exact retry of a durable mutation |
| Information gain | Expected entropy reduction from observing an indicator |
| Logit | Log odds, defined for probabilities strictly inside zero and one |
| Proper score | A rule whose expected optimum is honest distribution reporting |
| Reference class | Past cases selected using a stated comparability and inclusion rule |
| Resolution | An adjudicated outcome, distinct from the discriminating ability also called resolution in calibration analysis |
| Sensitivity envelope | Output range from alternative stated assumptions, not a statistical confidence interval |
| Sharpness | Concentration of predictive probabilities or distributions, independent of outcomes |
| Unresolved | The required outcome is not established; never an automatic no |
| Value of information | Improvement in expected decision utility from observing a signal, optionally less cost |
| Void | An administrative exclusion with an explicit reason |
When a term has more than one meaning, report the quantity and convention rather than relying on the label alone. “Resolution” is the most important example: journal adjudication and forecast discrimination are separate concepts.
Sources and attribution
The initial capability map and methodological safeguards follow Superforecasting Techniques — Skill Pack, reviewed at commit 17afb85d9ad56ba38a1402aa1b6093cceb2fd7ee. The pack is licensed under CC BY 4.0. Supercast implements new Rust routines, typed records, storage, transport, tests, and documentation rather than bundling the pack’s Python helper files.
The original pack’s method and provenance documents distinguish historical research, experimental validation, and workflow adaptations. Consult them for the underlying research attribution. Citing a method does not establish that this software has the same real-world performance as the participants in that research.
For implementation and statistical interpretation:
- SciPy bootstrap documentation describes percentile resampling intervals and contrasts them with other bootstrap methods.
- Stata panel bootstrap guidance explains keeping related observations together during resampling.
- Sebastiano Vigna’s SplitMix64 implementation supplies the public-domain generator adapted for deterministic sampling.
- SQLite transaction documentation describes the concurrency semantics used by the storage adapter.
- mdBook documentation describes the local searchable books and Rust example testing used here.
See the repository’s NOTICE.md and the release bundle’s dependency inventory for additional attribution. No affiliation with or endorsement by the source authors or forecasting organizations is implied.
- Scoring Rules theory documentation gives the equivalent CDF-integral and expectation forms of CRPS used to implement and independently test the routine.