Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Engineering reference

Supercast Supercast

This book describes the runnable implementation of Supercast: the dependency-free Rust core, the JSON adapter, SQLite journals, evaluation tools, and the local build and release process. It accompanies the Forecasting Handbook, which explains how to use the methods responsibly.

The API reference is generated from the executable’s request schema. Every operation page includes a request, any setup needed, an actual response with commit times redacted, and independently specified checks on the result. The documentation gate fails if an operation has no example, an example no longer works, or a generated page drifts.

Find the right level

NeedChapter
Build and run locallyInstallation
Call Rust functions directlyArchitecture and Rust API
Understand numerical corner casesNumerical contracts
Integrate a tool-calling processProtocol and operation index
Recover from an error or retryErrors
Understand persistence guaranteesStorage
Evaluate predictions correctlyEvaluation contracts
Maintain or release the projectTesting, operations, extension guide

The protocol is application-specific newline-delimited JSON, not MCP. It can be wrapped by an agent framework without coupling the core to a model provider. Source verification, model inference, scheduling and factual adjudication remain inputs from that application.

The implemented journal stores binary forecasts. Categorical and continuous scoring tools are also available as standalone calculations. Their existence does not imply a categorical/continuous journal type or an automatically fitted distribution model.

All documented commands are local. No build script publishes to a registry, sends forecasts to other people, or silently schedules future activity.

Install and build

Build from the workspace or use its local release bundle. Neither crate has been published to a package registry. The source uses Rust edition 2024; the verified toolchain is Rust 1.97.1 on x86_64 Linux. Other toolchains and targets require their own validation.

Source checkout

You need Rust and Cargo, a C compiler for bundled SQLite, and Python for the integration example and release tools. The release and documentation scripts require Python 3.11+. The books use mdBook 0.5.2.

CARGO_TARGET_DIR=target cargo build --workspace --locked
python3 examples/agent_demo.py

The core supercast crate has no runtime or development dependencies. supercast-agent adds Serde/JSON, Schemars, and rusqlite with SQLite bundled. The lockfile fixes the resolved dependency versions for reproducible builds.

With dependencies cached, use --offline. On a fresh connected machine, fetch them first:

cargo fetch --locked
cargo install mdbook --version 0.5.2 --locked
bash scripts/check.sh

The check script forces compilation into the checkout’s target directory unless SUPERCAST_TARGET_DIR is set. This prevents an inherited CARGO_TARGET_DIR pointing to an unavailable shared location from breaking the documented checks.

Rust dependency

For another local Rust project, add a path dependency:

[dependencies]
supercast = { path = "/path/to/supercast" }

Call the numerical and workflow modules directly. The core is synchronous and does not select an async runtime. For the adapter’s Store and protocol executor, depend on the separate crates/supercast-agent path.

Release bundle

Extract the host-specific archive and read RUN.md. The executable is optimized for the build host’s target and uses its system libraries; it is not a universal static binary. SQLite itself is bundled.

./supercast-agent --version
./supercast-agent --schema
python3 examples/agent_demo.py --binary ./supercast-agent
./supercast-agent --db forecasts.sqlite3

Use --memory when persistence is intentionally unnecessary. The database path is explicit; the process does not select an implicit home-directory database. Opening a new database initializes its schema, even before the first request.

Build the books

CARGO_TARGET_DIR=target cargo build --workspace --locked --offline
python3 scripts/docs.py

This executes all operation examples, generates the API chapters, builds both books, tests their Rust snippets, and checks rendered local links. Open the resulting local HTML files or serve them with mdBook during editing. python3 scripts/docs.py --check verifies generated source without rewriting it.

Architecture and Rust API

The dependency direction is intentionally simple:

validated probability types
           |
numerical methods and evaluation
           |
question / evidence / forecast lifecycle
           |
separate JSON + SQLite adapter
           |
calling agent application

The core cannot open a database, fetch a source, invoke a model, or start a background task. This keeps numerical behavior testable and lets a caller choose its own runtime and data infrastructure.

Core module map

ModuleMain types and operations
probabilityProbability, Categorical, logit/sigmoid, entropy
bayeslikelihood update, LR update, sensitivity, BetaPrior
referenceattributable case counts, empirical rates, quantiles
aggregateweighted arithmetic/logit pools, recalibration mapping
decomposeconditional chains, Frechet bounds, unions, mixtures, indicators, event-window probability
decisionexpected utility, best action, value of information
scoreBrier conventions, log loss, ranked score, CRPS, streaming Brier
calibrationfrozen question records, grouping, diagnostic report
evaluationpaired cases, holdout validation, time grids, cluster bootstrap, calibration selection
workflowquestions, evidence, forecasts, adjudications, ledger, updates, challenges and panels

Probability has a private representation. Construct it with new or TryFrom<f64> so invalid values cannot enter ordinary safe Rust code. Public workflow records use validated probability values but still require lifecycle validation when appended to a ledger.

extern crate supercast;
use supercast::Probability;

assert!(Probability::new(f64::NAN).is_err());
assert!(Probability::new(1.01).is_err());
let p = Probability::try_from(0.25)?;
assert_eq!(p.complement().get(), 0.75);
Ok::<(), supercast::Error>(())

Adapter modules

wire owns Serde/Schemars transport types. Its probability wrapper calls the core constructor during deserialization. protocol defines the request enum, response envelope and dispatch. analysis_tools translates the additional numerical/evaluation operations. store owns a SQLite connection, event replay and transactional writes. main implements bounded line input and flushed output.

The transport and core structs are separate by design. A future input format should convert through the same validated values rather than deserialize directly into private domain state. Ledger itself is not deserializable; persisted events must be replayed.

Error boundaries

The core returns supercast::Result<T> with explicit validation errors. The adapter maps those into structured error codes and messages, distinguishing malformed input, invalid models, storage failures, contention, conflicts and missing keys.

Core functions with mathematically bounded internal results sometimes construct a probability with an internal unwrap. External numbers still pass through fallible constructors. Unsafe Rust is forbidden in both project crates; dependencies such as SQLite have their own implementation contracts.

Generated Rust API pages from cargo doc --workspace --no-deps complement these books with source navigation. The JSON transport types and operation pages describe the cross-language boundary.

Numerical contracts

All core calculations use IEEE 754 f64. Results are finite unless a function explicitly represents an infinite loss or an endpoint log-odds transform. Computation is not arbitrary precision, and extreme values can saturate to endpoint probabilities.

Domain checks

QuantityAccepted domain
ProbabilityFinite closed interval [0,1]
Logit inputStrictly interior probability
Log-odds sigmoid inputAny value except NaN; infinities map to endpoints
Likelihood ratioFinite and strictly positive
Beta parametersPositive finite alpha/beta with a finite sum
Pool weightsFinite nonnegative values with positive total weight
Alpha / calibration slopeFinite and nonnegative
UtilityFinite; research cost also nonnegative
SamplesNonempty and finite
Scenario/category sumOne within 1e-12

Accepted category-sum roundoff is normalized at construction. Mixture weights are normalized by their accepted total. Probability inputs themselves are not silently clipped. A positive-weight zero/one input to logit pooling is an error rather than an invented epsilon.

Stable paths

Bayesian updates use log odds for nondegenerate cases and handle zero likelihoods explicitly. An observation impossible under both weighted hypotheses is rejected. A stable sigmoid avoids overflowing exp by choosing a branch according to the sign of its input.

Poisson event probability uses expm1 to retain precision for a very small rate-duration product. Empirical CRPS sorts the data and evaluates adjacent gaps; it scales before subtraction when opposite-sign extremes would overflow an intermediate distance. Truly unrepresentable scores still return errors.

extern crate supercast;
use supercast::score::{brier_skill, empirical_crps};

let mut extremes = [-f64::MAX, f64::MAX];
assert_eq!(empirical_crps(&mut extremes, 0.0)?, f64::MAX / 2.0);
assert!(brier_skill(1.0, f64::from_bits(1)).is_err());
Ok::<(), supercast::Error>(())

Score conventions

Binary Brier is one-component and ranges from zero to one. Categorical Brier sums over categories and ranges from zero to two. Ranked probability score is unnormalized and ranges from zero to K-1. CRPS has the outcome’s units. Log loss uses natural logs and permits positive infinity for a confidently wrong endpoint.

brier_skill rejects a zero baseline and an unrepresentable ratio. A database evaluation with a zero baseline returns null relative skill while preserving absolute means and their difference. Nonfinite JSON loss is tagged, not silently replaced by null.

Precision limits

Counts above 2^53 cannot retain unit precision when converted to f64. Very small probabilities may underflow, and extreme sigmoid results can round to zero or one. Sequence and source IDs are not floating-point values. The data model uses integer UTC seconds for timestamps.

Numerical tolerance is not a modeling tolerance. A sum differing by 1e-13 from one may be accepted as roundoff; two overlapping scenarios remain invalid even if their weights numerically sum to one. The caller is responsible for that semantic distinction.

Tests cover analytic cases, symmetry and proper-score properties, independent quadratic CRPS comparison, extreme finite values, endpoint behavior, binning decomposition, and seeded bootstrap behavior. Passing them establishes specified numerical behavior within these contracts, not real-world forecast validity.

Protocol and process lifecycle

Run supercast-agent --db PATH for durable storage or --memory for an ephemeral session. The process accepts newline-delimited JSON: one request object per line and one response per nonblank line. Output is flushed after every response so an agent can keep a subprocess open across tool calls.

The protocol is versioned independently from the crate release. Requests currently require version: 1, an ID, and a command object with an op discriminator. Unknown fields are rejected, including nested typed fields.

{"version":1,"id":"request-001","command":{"op":"binary_score","probability":0.3,"outcome":true}}

Success contains result; failure contains error. Every decoded request retains its ID in the response. Malformed or untyped input errors use a null ID because the envelope was not successfully decoded.

IDs and retries

IDs must contain nonwhitespace content and at most 200 UTF-8 bytes. For calculations and reads, an ID is response correlation only. For successful durable mutations, it is a database-wide idempotency key.

Use a new ID for a new intended mutation. After a timeout or lost response, resend the same decoded command and ID. The original receipt is returned, including its original commit timestamp. Whitespace and object-key order do not distinguish decoded commands. Reusing a committed ID with different content is a conflict.

The schema describes structural constraints. Runtime validation adds UTF-8 byte limits, numerical bounds, event lifecycle checks and cross-record relationships that JSON Schema alone does not capture.

Process limits and exits

Each input line is limited to 1 MiB including its newline. An oversized line receives request_too_large and ends the stream before its command executes. Blank lines are ignored. Other request errors return a response and processing continues.

Exit codeMeaning
0Every processed request succeeded
1One or more request errors, including oversized input
2Startup or I/O failure

During serving, stdout is reserved for protocol responses. Startup diagnostics go to stderr. --help, --version and --schema are standalone informational modes and do not open a database.

Streaming client obligations

Keep writes and response reads coordinated. Never assume a successful stdin write means the database committed. Confirm the response ID and inspect result versus error. On process failure, restart and retry unresolved mutations with their original IDs.

An application that allows parallel tool calls should serialize access to one process’s stdin/stdout or use one process/connection per worker. SQLite handles database write contention, but it does not multiplex responses for a poorly coordinated client.

The Python demo uses a finite batch, closes stdin, and checks response count and IDs. It also proves that a committed update can be retried after a process restart. It is a subprocess example rather than a native Python binding.

Errors and recovery

Treat an error as a result with a defined recovery path. Do not convert it to a neutral probability or silently drop a question from evaluation.

CodeTypical causeAppropriate response
invalid_jsonSyntax, unknown fields, wrong types, invalid probability during decodingCorrect the request; no command executed
validationImpossible evidence, incoherent probabilities, bad timeline, stale predecessorReconsider inputs or read current state; do not fabricate a substitute number
not_foundMissing question/version keyCheck target identity; create the intended question if appropriate
conflictReused request ID with different content, duplicate persisted key, SQL constraintReconcile intended mutation with existing state
busySQLite could not obtain a required lock within five secondsBack off and retry the identical mutation/ID
storageOther SQLite or journal failurePreserve the database, inspect the fault, and restore only through a verified procedure
request_too_largeLine exceeds 1 MiBReduce/split the operation where semantics permit; restart the process

Core validation errors are mapped to validation after a request is decoded. A numeric probability rejected by transport decoding instead appears as invalid_json. Error messages explain a specific failure but are not a stable machine enum; branch on the code and your operation context.

Failed writes are atomic

A mutation validates under the database transaction and commits its event and receipt together. Failure in either insert rolls back both. The test suite injects a receipt-write failure after the event insert to verify that no partial event survives.

A failed mutation does not consume the request ID. You can fix a previously uncommitted request and submit it under that ID, although recording a distinct attempt ID can be useful at the application layer. After a successful mutation, the ID is bound to its decoded content.

Stale revisions

If another worker appends after you read the ledger, your predecessor becomes stale. Reading the newest revision and blindly substituting its ID is unsafe: the evidence may already be incorporated or the assumptions may have changed. Recompute the intended update against the new history.

Impossible evidence

When a supplied observation has zero probability under the entire current model, Bayesian updating cannot produce a valid posterior. This is a model error. Inspect the event definition, prior, likelihoods, and observation claim. Returning 0.5 would hide the failure.

Startup and output failure

An invalid database, unsupported schema version, unavailable filesystem path, or failed output write can stop the process with exit code 2. These failures may not have a JSON response. If a response was lost, use idempotent replay to determine whether the intended mutation committed.

The library does not automatically repair an unknown schema or delete damaged history. Keep backups and the release version that created them. See storage and operations for the supported local workflow.

SQLite storage

Store owns one SQLite connection. New empty databases are initialized with an application ID and schema version. Existing nonempty foreign databases and unknown versions are rejected rather than silently repurposed.

The current application ID is 0x53555043, and database schema version is 1. Crate release versions and protocol versions are separate. A code update must not assume that changing a package version authorizes an on-disk migration.

Tables and ordering

events uses the composite primary key (question_id, version, sequence). Each row contains a local commit timestamp and JSON event payload. Sequence one contains the question contract; later events are forecasts or resolutions. receipts maps a successful request ID to its decoded command and result.

Both tables are WITHOUT ROWID. Their primary keys support journal lookup and request replay. Update/delete triggers reject accidental edits to existing events and receipts. There is no cryptographic claim against a database administrator who can drop those triggers or rewrite the file.

Write transaction

For a mutation, the adapter:

  1. Begins an immediate transaction and checks for an existing receipt.
  2. Returns an exact matching receipt or rejects mismatched reuse of its ID.
  3. Checks the connection-local data_version. Reuses a matching validated journal or loads and fully replays its history.
  4. Validates the new event, including its predecessor and evidence.
  5. Inserts the event and receipt.
  6. Commits both together, then retains the validated state only if exactly the expected two SQL row changes occurred.

The lock covers the read/validate/write sequence, so two writers cannot independently validate the same predecessor and both create competing branches. A five-second busy timeout handles transient contention; callers still need a bounded backoff policy.

Replay validation

Reloading checks question identity and sequence continuity, reconstructs the core ledger, and applies every lifecycle rule. For calculated Bayesian revisions it also verifies the recorded prior, recomputes the posterior, checks evidence correspondence, and rejects origins already incorporated.

Origin tracking uses one running set across replay and subsequent cached mutations. Each Store retains at most one question history, with O(history plus evidence) memory. It moves that state into a tentative mutation rather than cloning the entire ledger. Any validation, insertion or commit error after taking the state discards it; only a successful commit can restore it.

The cache token is read inside the immediate transaction. SQLite changes data_version when another connection commits, so even an external change to an earlier event invalidates the cache. Any external commit conservatively invalidates it, including an unrelated question’s write. Extra same-connection trigger changes suppress cache retention. Reads and evaluation always perform full replay; restart and question switches also take the cold path. There is no persistent checkpoint and no schema change.

See storage profiling and results for the measured gain, cache tests and remaining limits. Do not describe database writes using the core’s nanosecond arithmetic benchmark.

Durability

Connections use rollback journaling with synchronous=FULL. A successful response follows the SQLite commit. Durability still relies on SQLite, the filesystem, and storage hardware honoring their contracts. Local filesystems are the tested context; shared network filesystems are not part of the validated deployment model.

The implementation is based on SQLite’s transaction semantics and synchronous setting. Tests cover reopening, competing connections, numerical replay, protected rows and failed-write rollback. They do not emulate power loss.

For backups, stop writers before copying the database or use SQLite’s backup facilities. Do not copy only one file from an active database state without understanding its journal. Treat restored records as needing the same replay validation as ordinary reads.

Evaluation contracts

Supercast provides distinct evaluation tools for different tasks. Do not interpret one tool’s checks as evidence that another task’s assumptions have been established.

ToolWhat it establishes from supplied inputs
comparePaired score means on unique, pre-resolution question records
validate_holdoutQuestion uniqueness, event-cluster separation and temporal order
grid_scoreTime-grid Brier within one question using carry-forward forecasts
cluster_bootstrapA resampled paired performance interval under declared clusters
select_calibrationTraining-only candidate choice with subsequent holdout scoring
evaluateA database snapshot evaluation with explicit exclusions

Case-level evaluation

A Case holds question ID, cluster ID, forecast time, resolution time, model probability, baseline probability and binary outcome. The model/baseline pair shares one outcome and timestamp, preventing accidental comparison of unrelated rows within the record format.

validate_holdout requires nonempty training and holdout sets. Training labels are available by the cutoff; held-out forecasts occur after it. Every forecast precedes its own resolution. Related questions must use the same cluster ID so that cluster overlap can be rejected.

The baseline’s historical availability is not independently verified by the Case structure. The caller must freeze a legitimate benchmark. Do not use retrospective prevalence as if it were an operational ex-ante forecast without saying so.

Bootstrap implementation

Clusters are sorted by ID; rows inside them are sorted by question ID before their score differences are summed. Each replicate draws the same number of clusters with replacement and divides the total paired score difference by the total selected question count. Unequal cluster sizes therefore retain question weighting.

SplitMix64 provides seeded deterministic draws, with rejection sampling to remove modulo bias. At most 128 rejection attempts are allowed for an index; failure is explicit. Sampling permits 100–100,000 replicates and at most 10 million cluster draws. The computation allocates cluster summaries and one vector of replicate statistics, not a copied question matrix per replicate.

Percentile limits use type-7 interpolation. Standard error is the sample standard deviation of the replicate statistics. A constant distribution sets degenerate: true. These are percentile bootstrap intervals, not BCa intervals or posterior credible intervals.

Calibration selection

The candidate grid contains 1–1,000 finite intercept/nonnegative-slope pairs. At most 10 million training candidate/question evaluations are permitted. Each candidate transforms interior probabilities through a sigmoid of affine log odds and is scored on training outcomes.

Exact score ties use input candidate order. The chosen parameters are then applied to the holdout once; the operation returns all training scores and the selected method’s raw, calibrated and baseline holdout scores. It does not select a new candidate because the first performed poorly on held-out data.

The caller must prevent repeated adaptive use of the same holdout. A function cannot infer prior experiments that were not supplied.

Database snapshots

evaluate accepts 1–10,000 explicit keys, a common forecast cutoff, a resolution cutoff, and a fixed baseline with its declared as-of time. Duplicate question IDs are rejected even if the version differs. A missing key fails the request rather than shrinking the cohort.

It reads a consistent database transaction, selects the most recent forecast available by the cutoff, and selects the latest resolution available by the outcome cutoff. Cases outside the question’s forecast window, missing forecasts, unresolved outcomes and voids are reported as exclusions. No scored rows produces a null summary.

This operation labels its output declared_timestamp_snapshot. Supplied historical timestamps and baseline dates do not prove prospective participation. Commit-time enforcement, cohort preregistration, and independently observed future outcomes remain necessary for a prospective performance claim.

Performance and capacity

The core favors straightforward algorithms with explicit allocation behavior. Keep arithmetic latency, record-validation cost, database synchronization and model inference separate when measuring an agent.

OperationTimeExtra storage
Scalar Bayesian update or BrierO(1)O(1)
Weighted poolO(n)O(1)
Streaming BrierO(1) per updateCount and mean
Conditional chain / mixtureO(n)O(1)
Empirical quantile / CRPSO(n log n) sortIn-place sample slice
Calibration groupsO(n log g)IDs and nonempty groups
Time-grid scoreO(updates + grid)O(1)
Cluster bootstrapSorting plus O(replicates × clusters)Cluster summaries and replicates
Calibration grid selectionO(candidates × training + holdout)Candidate scores
Journal read / cold mutation validationO(history + evidence)Replayed history and IDs
Warm mutation validationAmortized O(new event + new evidence), excluding database workOne retained history and ID/origin sets

Some transport commands copy caller-owned arrays before invoking an in-place core routine. That cost belongs to the JSON adapter, not to the scalar numerical function. The adapter parses a whole request line; it does not stream individual JSON array entries.

Benchmark locally

CARGO_TARGET_DIR=target cargo bench --bench core --locked --offline

The harness warms each case for 1,000 calls and uses optimizer barriers around inputs and results. It measures a scalar Bayesian update, a pool of 100 equal-weight interior probabilities, and a running Brier update with mean access.

An early local run on an AMD Ryzen Threadripper PRO 5975WX with Rust 1.97.1 measured approximately 31 ns per Bayesian update, 1.5 microseconds per 100-forecast pool, and 4.6 ns per streaming-score operation. These are one-run observations, not stable latency guarantees or comparisons against another implementation.

Workload shape, compiler flags, thermal state, contention and input distribution affect measurements. Record those conditions for comparative claims. Do not infer forecasting accuracy from execution speed.

Capacity limits

The CLI limits a request line to 1 MiB. Listing is capped at 1,000 keys per page and database evaluation at 10,000 keys per request. Bootstrap and calibration selection have explicit work caps. The database journal can grow over time; the request-size limit does not impose a total history limit.

For long-running deployments, monitor journal length, mutation latency and cache hit rate. Repeated writes to the same journal on one Store use incremental validation. Restart, question switches and external commits require full replay. Memory retains one history per Store; it grows with that history. Concurrent or interleaved workloads can see many cache misses. See measured storage results.

Durable writes pay a filesystem synchronization cost. Use one connection per worker and expect write serialization under SQLite. Increasing process count does not make one local database accept unlimited simultaneous writers.

Measure the end-to-end agent separately, including retrieval, inference, evidence processing and persistence. Those application costs are likely to dominate the numerical primitives, but their actual contribution depends on the deployment.

Dataset replay and stress testing

The standard-library Python harness executes the compiled Supercast JSON adapter. It measures both numerical agreement with independent Python formulas and process/storage behavior. It does not call a language model or generate new factual forecasts.

Reproduce

Build the optimized binary, fetch pinned public inputs, then run the two workloads:

CARGO_TARGET_DIR=target cargo build -p supercast-agent --release --locked --offline
python3 scripts/fetch_gjp.py
python3 scripts/benchmark.py replay --output target/benchmarks/gjp-year1.json
python3 scripts/benchmark.py stress --events 100000 --revisions 500 --output target/benchmarks/stress-100k.json
python3 scripts/benchmark.py stress --events 1000000 --revisions 1000 --output target/benchmarks/stress-1m.json
python3 scripts/benchmark.py stress --events 10000000 --revisions 2000 --output target/benchmarks/stress-10m.json

Run timing workloads sequentially. Reports include the binary hash/version, seed, elapsed time and workload configuration. Fetching needs internet and curl; replay and stress are offline. Raw inputs and generated reports stay in ignored target/ and are not included in release source snapshots. The standard release gate includes three importer/oracle regression tests and a small synthetic integration run; downloading real data is not a prerequisite for that gate.

Source and provenance

Source: Good Judgment Project, GJP Data, Harvard Dataverse, doi:10.7910/DVN/BPCDH5. The dataset metadata identifies CC0 1.0. IARPA’s release announcement describes the four-year public archive.

The downloader pins three Dataverse file IDs: 2917330 (question metadata), 2917351 (year-one survey forecasts), and 2917350 (field documentation). It verifies SHA-256 fingerprints of the served representations and writes a manifest with download URLs and byte counts. A changed or corrupt existing file fails verification instead of being overwritten. Dataverse’s ingested TSV representation differs from the original CSV, so its original-file MD5 is not the checksum of the downloaded TSV.

Only question IDs, option counts, question type/status, dates and outcome labels are consumed from ifps.csv. That file is decoded with byte-preserving Latin-1 because it contains mixed non-UTF-8 prose; its prose is not rewritten or imported into a factual journal. Forecast TSV fields are UTF-8. No demographic or individual-difference data is downloaded.

Frozen replay protocol

  1. Retain ordinary (q_type=0), closed, two-option questions with resolved option a or b. Conditional, categorical, void and unresolved questions are excluded.
  2. Freeze each question seven days after its recorded opening. Exclude questions already suspended or resolved at that instant. This is a fixed age from opening, not a fixed distance from resolution.
  3. Keep only option-a rows between opening and snapshot, inclusive. For each question/forecaster use the latest (timestamp, forecast_id); forecast ID deterministically breaks timestamp ties. A latest withdrawal removes that forecaster from the panel. Complementary option-b rows are not additional observations.
  4. Interpret outcome a as true. This is the event “option a occurs,” which need not be the literal word “yes.” Apply declared endpoint regularization epsilon=0.001 to both pooling methods. Count adjustments. Every retained forecaster has equal weight regardless of update frequency or experimental condition.
  5. Training questions must resolve strictly before 2012-01-01. Holdout snapshots must occur on or after that cutoff and have no training cluster overlap. Other snapshots are excluded. Cluster IDs use the base IFP identifier before its suffix.
  6. Select among a prespecified grid of intercepts [0,-0.5,0.5] and slopes [1,0.5,1.5,2], minimizing training Brier only. Ties choose the first candidate. Evaluate the frozen selection once on holdout questions.
  7. Report equal-question Brier, natural-log loss, ten-bin calibration diagnostics and paired Brier differences against arithmetic pooling. Include a neutral 0.5 baseline. Cluster percentile intervals use 2,000 resamples, 95% confidence and seed 20260916.

Independent Python formulas verify arithmetic/logit pooling, per-question Brier/log loss, candidate selection, calibrated holdout score and calibration’s raw Brier. The bootstrap implementation is exercised, but its full sampling distribution is not independently reimplemented here. Reports retain the actual training and holdout cases for inspection.

Archive timestamps lack an explicit timezone in these files. The importer preserves their nominal ordering by interpreting all dates on one UTC-labelled clock; it does not claim to reconstruct historically accurate UTC instants. Forecasts recorded before a question’s metadata opening date are excluded conservatively. The exclusion counts make this choice visible.

Related geopolitical questions can remain dependent even with distinct base IDs. The confidence intervals assume those clusters are independent, and condition on the selected calibration model; they do not include training-selection uncertainty. Seventeen training questions are a small sample. This is a transparent regression/aggregation benchmark, not a reproduction of the tournament’s full daily scoring protocol, a randomized treatment comparison, or evidence that an LLM can predict these events prospectively.

Synthetic stress protocol

--events counts binary-scoring JSON requests; --revisions separately controls the length of one durable journal. Ten million numerical requests do not imply ten million SQLite writes. Requests are generated incrementally with a fixed seed. Normal draws are mixed with exact endpoints and near-endpoint values. Python verifies every Brier result and checks finite or infinite log-loss semantics.

Measurements include end-to-end single-client throughput and round-trip p50/p95/p99. Quantiles use a seeded uniform reservoir capped at 10,000 timings; they are estimates, not exact quantiles over all requests. Throughput includes Python serialization, validation, IPC and sampling overhead. Linux agent RSS high-water marks are sampled every 1,000 numerical requests; they do not measure the Python driver or subsequent storage workers.

The separate storage workload appends an increasing revision chain, repeats exact requests, rejects a stale update, kills/reopens the process, and verifies original receipts and journal length. It also kills a process after sending a request but before reading acknowledgement, retries that request and verifies exactly one additional revision. This permits either pre-commit or post-commit interruption; it does not guarantee a kill occurred inside SQLite’s commit or simulate power loss. Four workers create 16 additional questions in one database, followed by SQLite integrity checking.

Storage reports show write latency, first/last-quarter means and final database size. The v0.2.0 baseline replays history on each write. The v0.2.1 adapter reuses validated state for consecutive writes to one journal, while cold or invalidated caches still replay. See incremental-validation measurements. Millions of writes on one journal would be a separate, substantially more expensive capacity experiment. Concurrent creation is a contention smoke test, not a saturation benchmark. Timing results are single-host observations without performance pass/fail thresholds.

See the checked-in observed results for the completed local run.

Observed benchmark results

These local observations use optimized supercast-agent 0.2.0 on x86_64 Linux. They are reproducible workload results, not stable performance guarantees or prospective agent accuracy. Full machine-readable reports are under target/benchmarks/.

GJP year-one replay

Read 308,232 answer rows, yielding 88 frozen question snapshots: 17 training, 46 holdout, and 25 outside the temporal split. Independent numerical checks passed.

MethodHoldout Brier (lower is better)Log lossPaired Brier difference vs arithmetic, 95% interval
neutral0.2500000.6931470.087682 [0.059452, 0.116920]
arithmetic0.1623180.5058640.000000 [0.000000, 0.000000]
logit0.1447060.457483-0.017612 [-0.027226, -0.007341]
recalibrated0.1144800.364550-0.047837 [-0.080911, -0.011487]

Training selected {'intercept': -0.5, 'slope': 2.0} from the prespecified grid. Endpoint regularization adjusted 1944 retained forecaster probabilities. The neutral baseline is 0.5 for every question. Binary Brier uses the [0,1] convention, half the two-component binary Brier used in some GJP publications.

These results apply to the selected historical cohort. Base-question clustering does not establish independence between related geopolitical events. Calibration selection used only 17 training questions; the reported intervals condition on that fitted selection. See the protocol and exclusions.

JSON scoring stress

Every requested binary score matched the Python Brier oracle. The ten-million run additionally checked log-loss values and infinite endpoint penalties on every request. Quantiles are from a bounded 10,000-observation reservoir. Timing includes driver overhead.

RequestsSecondsRequests/sp95 round trip (ms)p99 (ms)Agent peak RSS (KiB)
100,0002.8834,6910.03190.03715964
1,000,00028.5135,0800.03180.03746088
10,000,000306.3432,6430.03340.04485952

Separate durable journal workloads

Revisionsp95 write (ms)First-quarter mean (ms)Last-quarter mean (ms)Final database bytes
5001.0510.3250.996835,584
1,0001.9360.4341.8111,564,672
2,0003.4220.6523.1703,108,864

All three storage workloads passed exact retry, stale-update rejection, acknowledged and unacknowledged-write restart checks, 16 concurrent question creations with four workers, and SQLite integrity checking. Database sizes include the additional interrupted revision and concurrent questions.

Historical v0.2.0 limitation: writes slowed with history length because every mutation replayed the journal. The subsequent v0.2.1 incremental-validation change addresses consecutive warm writes; see the before/after results. These original measurements remain unchanged. Ten million scoring requests do not establish ten-million-write capacity.

Provenance

Binary SHA-256: a6a9d5484ed1cda43b47754a5c3107cc808e13896fe2225d797126b228d96609

Seed: 20260916. Python: 3.14.4. Dataset: doi:10.7910/DVN/BPCDH5.

  • ifps.csv SHA-256: 1f64e5483741e656ddc667b310799e9aec8a8c5dbb7e12f9af5c5a60fd146286
  • survey_fcasts.yr1.tsv SHA-256: a070cc0e87eda8aba63058669308b41022b38f4030e421ee34173528cbb9e2b2

Incremental storage validation

The v0.2.1 adapter removes repeated full-history work from consecutive writes to the same journal on one Store. The numerical core remains v0.2.0; database schema and JSON protocol remain version 1. SQLite durability settings, immutable events and atomic retry receipts are unchanged.

Why this change

A stage profile of the original 2,000-revision workload identified history retrieval/JSON decoding as the dominant cost. The following are mean microseconds per mutation for the last 500 revisions. These timings exclude subprocess transport and use synthetic fixed-size manual forecasts.

StageBeforeAfter
Begin/write lock12.7812.81
Receipt lookup7.506.85
Read and decode history2130.590.00
Replay history343.460.00
Other validation/state handling215.462.91
Insert event and receipt72.0864.67
Commit38.6337.49
Total mutation2821.83125.99

All profiled post-change appends used validated state. Skipping the old scan addresses the observed growth with history length. SIMD would not eliminate that scan or its increasing amount of work. The remaining time is largely transaction/index/serialization and durability work.

Same-workload comparison

The saved unmodified v0.2.0 binary and new v0.2.1 binary each ran the same seeded JSON/SQLite stress harness three times at each journal length. Runs were sequential, with order alternated between repeats. Both binaries used rollback journaling and synchronous=FULL. The table reports the median of three per-run means for the last quarter of writes; p95 columns are medians of the three per-run p95 values.

RevisionsBefore last-quarter mean (ms)After (ms)SpeedupBefore p95 (ms)After p95 (ms)
5000.9850.1805.5×1.0480.227
1,0001.7300.1769.8×1.8650.229
2,0003.2210.17318.6×3.4850.230

An additional 10,000-revision candidate run passed all recovery and integrity checks. Its first-quarter mean was 0.180 ms and last-quarter mean 0.194 ms; p95 was 0.235 ms. The storage process RSS high-water mark was 12,996 KiB, sampled before recovery tests; database size was 15,273,984 bytes after recovery/concurrency checks. This is one local run, not a universal latency or memory guarantee.

The comparison uses only 100 numerical requests per run to fix the seeded setup and focus on journal behavior. It does not repeat the ten-million scoring run: no numerical algorithm changed. Every storage run still checks exact retries, rejected stale writes, acknowledged and unacknowledged-write recovery, concurrent creations, journal lengths and SQLite integrity.

Cache invariants

  1. Each Store owns one private SQLite connection and retains at most one fully validated question history, forecast-ID set and evidence-origin set. Retained memory grows with that history. There is no whole-ledger clone per append.
  2. Every mutation acquires BEGIN IMMEDIATE before consulting the cache. A cached question/version is reusable only when its connection-local PRAGMA data_version matches the token saved during validation. Another writer cannot commit while this transaction holds the write lock.
  3. Any external commit invalidates the cache, including writes to unrelated questions or historical edits. On a miss, the complete journal is read and replayed, preserving sequence, lifecycle, identity, origin and Bayesian audit checks. Tokens are never compared between connections.
  4. Cached state is moved into a tentative mutation. Validation/insertion/commit failures discard that tentative state. It is installed again only after commit succeeds. Request-ID receipts remain authoritative for exact retries, and retries do not advance cached state.
  5. The normal mutation changes exactly two rows: one event and one receipt. Extra same-connection changes, including trigger side effects, suppress cache retention. SQLite total_changes supplies this additional check.
  6. Journal reads and evaluation always fully replay stored history. Reopening, switching questions or discarding the cache also requires replay. No persistent checkpoint or schema migration is introduced.

These rules follow SQLite’s data-version semantics, immediate transaction semantics, and total-change counting. The cache does not authenticate raw database-file edits outside SQLite or a malicious database administrator.

Correctness evidence

New tests verify warm hits versus cold/question-switch misses, stale-writer invalidation, corruption of an earlier row while the cache is warm, receipt-insert rollback, commit-time deferred-constraint failure, manual and Bayesian evidence-origin tracking, same-connection trigger changes and agreement between warm writers and writers reopened before every append. Existing fork-prevention, replay, schema and durability tests continue to run. No correctness test depends on a timing threshold.

Remaining limits

Warm lifecycle validation is amortized O(new event plus new evidence), using retained ID/origin sets. The complete database operation still includes indexed lookups/inserts and a durable commit. A cold mutation is O(history plus evidence). Repeated external commits or interleaving different journals on one Store can force repeated cold replay; this optimization does not claim constant latency for those workloads. One active history remains allocated per Store. Persistent checkpoints, a bounded multi-question cache and transaction batching are separate future changes.

Reproduce

Preserve a v0.2.0 binary before rebuilding, or extract it from the existing v0.2.0 release archive. The comparison must use two different binaries; the report records their SHA-256 hashes. Then run:

CARGO_TARGET_DIR=target cargo build -p supercast-agent --release --locked --offline
python3 scripts/benchmark_storage.py --baseline target/benchmarks/baseline-agent --binary target/release/supercast-agent --output target/benchmarks/incremental-validation
CARGO_TARGET_DIR=target cargo run -p supercast-agent --example profile_store --release --locked --offline -- 2000 > target/benchmarks/store-profile-after.json
python3 scripts/benchmark.py stress --events 100 --revisions 10000 --output target/benchmarks/incremental-validation/candidate-10000.json

Store::last_mutation_profile() exposes microsecond stage timings for the last successful new mutation. Mutation attempts clear it; exact retries leave no profile. It does not change JSON output. The example emits one JSON report containing all mutation profiles and verifies the completed journal with a full audit. Baseline phase timings came from an instrumented pre-cache build; the end-to-end comparison uses the unmodified saved baseline.

Full run reports are in target/benchmarks/incremental-validation/; the phase profiles are in target/benchmarks/store-profile-{before,after}.json.

Baseline binary SHA-256: a6a9d5484ed1cda43b47754a5c3107cc808e13896fe2225d797126b228d96609

Candidate binary SHA-256: 79b85c50d5cbe57c6dde874cd30715f26dfd21163dbfb20b0a86f2548cd101d0

Testing and documentation

Run the release gate from the workspace:

bash scripts/check.sh

It checks formatting, workspace tests, Clippy with warnings denied, Rust API documentation with warnings denied, the subprocess restart/retry example, and the mdBook documentation gate. Dependencies are resolved from the lockfile and local cache.

What the tests cover

Numerical tests include analytic results, probability-domain rejection, Bayesian symmetry, expected Brier optimality, CRPS against an independent quadratic formula, finite extreme values, reference-class denominators, and calibration raw/coarsened distinctions.

Workflow tests cover revision identity, evidence cutoffs, unchanged reviews, origin reuse, terminal adjudications and time-grid weighting. Adapter tests exercise strict decoding, persisted replay, request idempotency across restarts, competing writers, SQL immutability, failure rollback and explicit evaluation exclusions. Incremental-validation tests cover external commits, historical corruption, question switches, warm/cold equivalence, evidence-origin reuse, trigger side effects and commit-time constraint failure.

Evaluation tests check whole-cluster resampling against a simple known distribution, unequal cluster sizes, degeneracy, workload bounds and selection that does not use held-out labels. This verifies algorithmic behavior rather than empirical interval coverage for arbitrary real datasets.

Executable books

CARGO_TARGET_DIR=target cargo build --workspace --locked --offline
python3 scripts/docs.py
python3 scripts/docs.py --check

The first command writes updated generated reference pages. The check mode fails on drift. Both execute every JSON operation example, check explicit expected fields, build the two mdBooks, run their Rust code examples and verify local links and fragments in the rendered HTML.

The example fixture is scripts/doc_cases.py. Each case names one operation, setup commands if needed, an actual request, explanatory contracts and independently specified expected output fields. The set of cases must equal the runtime schema’s operation set exactly; a new operation without documentation fails the gate.

Commit timestamps are redacted in generated response examples so reference files are stable. Other outputs come from the binary. Expected float values use a 1e-12 comparison tolerance. Full result snapshots are displayed for inspectability.

Rust book examples use a copy of the Cargo-built core rlib in an isolated temporary search directory. This avoids accidentally selecting stale Clippy/check metadata from the shared build cache. Examples remain compiled against the current core implementation.

Boundaries of verification

The suite does not prove that a supplied source is authentic, an event cluster assignment is valid, a historical forecast was actually made before its outcome, or a utility function is appropriate. Those are application/data obligations. Storage tests do not simulate power loss or adversarial administrators.

A passing gate is necessary for a release, but inspect the changed behavior and artifacts too. A test that merely repeats an implementation can preserve the same error in two places; independent examples and invariants are included to reduce that risk.

Packaging and operations

python3 scripts/release.py checks the workspace, builds an optimized host executable, prepares a local bundle, smoke-tests the staged executable, and writes an archive plus SHA-256 checksum under target/releases.

The bundle includes the binary, Markdown documentation, both rendered books, schema, examples, source snapshot, dependency notices/inventory and a per-file checksum manifest. RUN.md explains the executable and documentation entrypoints. Nothing is published to a registry or remote hosting service.

Release procedure

  1. Make source and documentation changes together.
  2. Regenerate the request schema if the command types changed.
  3. Run python3 scripts/docs.py to update generated operation pages.
  4. Run the release script. It reruns the full gate before packaging.
  5. Inspect its output and verify the archive checksum before distributing it in an authorized context.

Existing release archives are not silently overwritten. Use a new release version when the distributable changes. A project’s source license and publication policy are separate from successful local packaging; this project has not selected a license for its new code.

Running the process

Use an explicit database path on a supported local filesystem. Keep the executable version and frozen schema information with backups. The process runs in the foreground; a service manager, model runtime or scheduler belongs to the consuming application.

Monitor stderr, exit status, database size and operation latency. Ordinary request errors return JSON and allow later requests to proceed, so a long-lived client must inspect individual responses rather than waiting for process exit.

Backup and restore

Stop writers before copying a database, or use SQLite’s backup facilities. An in-flight journal and database file must not be treated as unrelated files. Restore to a separate path, open it with the matching release, and read representative or all registered journals through get_ledger to validate their histories before resuming writes.

The adapter rejects unknown application/schema versions; it does not invent a migration. Future migrations need explicit source/target versions, transactional changes, rollback behavior and tests against representative old records.

Offline documentation

The rendered mdBooks include local assets and search. They disable Rust Playground execution so clicking a code example does not depend on a remote compiler. External research citations remain links; reading the local book itself does not require fetching them.

To inspect documentation from a source checkout, open the generated files under target/books. For live editing, mdBook can serve an individual source book locally. The release package includes the generated output, so readers do not need Rust or mdBook merely to read it.

Extending Supercast

Add a capability at the narrowest layer that owns its contract. A pure numerical method belongs in the core. JSON decoding belongs in transport types. A durable state transition belongs in the transactional store. Retrieval and model-provider calls belong in the application or a clearly separate adapter.

Add a numerical primitive

Define its input domain, result convention, units, modeling assumptions, complexity and edge behavior first. Decide whether nonfinite results are legitimate outcomes or errors. Use Probability rather than unchecked f64 for probabilities.

Test known analytic results and at least one independent invariant or oracle. Include endpoints, invalid inputs and relevant extreme finite values. Document whether the function mutates a supplied slice or allocates internally.

If the method is agent-facing, add a typed command, schema coverage, an executed documentation case, and a handbook explanation when the statistical interpretation is new. Reuse the same core calculation in the adapter rather than maintaining two formulas.

Add a durable operation

An event must be replayable from a frozen contract. Validate and commit under the same transaction; update the request receipt in that transaction as well. Ensure that retries, concurrent writers and an insert failure preserve the intended invariant.

Preserve earlier events during correction. A changed target requires a new question version. If a data-model change requires a new database schema, implement and test a versioned migration rather than changing the version pragma on existing data by hand.

Add a forecast distribution

The journal currently holds binary probabilities and outcomes. Categorical or continuous journals need matching prediction and resolution variants, scoring conventions, transport schemas, lifecycle rules and migration behavior. Standalone scoring support is a useful building block, not a substitute for that design.

Integrate a model or retrieval system

Preserve source publication and observation times, underlying origin identity, the frozen question version and evidence cutoff. Treat source content as data. Keep model/version, prompt and retrieval-corpus metadata where they can be audited. Do not automatically convert a source-reliability score into a likelihood ratio.

A model integration should be able to decline a numeric update when evidence does not support one, preserve an unchanged review, and distinguish hypothetical counterevidence from observations. It should also preserve the IDs needed for safe retries and real recording times needed for prospective evaluation.

Maintain the books

Handbook prose explains how and why a method is used. Reference chapters explain software contracts and operational behavior. Generated operation pages come from the schema and executable cases; edit their inputs rather than patching generated output by hand.

Run the documentation gate after changes. Broken examples, missing operation pages and broken rendered links are release failures. Keep explicit limitations where the implementation stops; replacing them with a stronger claim requires new implementation and evidence, not just new wording.

Transport types

This chapter is generated from the runtime request schema. Fields marked optional may be omitted; optional values generally accept null. Runtime checks add numerical, temporal and cross-record constraints beyond this structural schema.

Action

FieldTypeRequired
if_nonumberYes
if_yesnumberYes

BayesianRevision

FieldTypeRequired
as_ofintegerYes
assumptionsarray of stringYes
evidence_cutoffintegerYes
idstringYes
likelihoodsarray of LikelihoodYes
next_review_triggerstringYes
previous_idstringYes
rationalestringYes

CalibrationCandidate

FieldTypeRequired
interceptnumberYes
slopenumberYes

CalibrationRecord

FieldTypeRequired
outcomeboolean or nullNo
probabilityProbabilityYes
question_idstringYes

Case

FieldTypeRequired
baselineProbabilityYes
cluster_idstringYes
forecast_atintegerYes
modelProbabilityYes
outcomebooleanYes
question_idstringYes
resolved_atintegerYes

Change

{
  "enum": [
    "initial",
    "evidence",
    "assumption",
    "correction",
    "review_unchanged"
  ],
  "type": "string"
}

Elicitation

FieldTypeRequired
opposing_reasonstringYes
probabilityProbability or nullNo
respondentRespondentYes
respondent_idstringYes
supporting_reasonstringYes

Evidence

FieldTypeRequired
claimstringYes
idstringYes
locatorstringYes
observed_atinteger or nullNo
origin_idstringYes
published_atintegerYes

Forecast

FieldTypeRequired
as_ofintegerYes
assumptionsarray of stringYes
changeChangeYes
evidencearray of EvidenceYes
evidence_cutoffintegerYes
idstringYes
methodstringYes
next_review_triggerstringYes
previous_idstring or nullNo
probabilityProbabilityYes
rationalestringYes

Indicator

FieldTypeRequired
chanceProbabilityYes
if_noProbabilityYes
if_yesProbabilityYes

Key

FieldTypeRequired
question_idstringYes
versionintegerYes

Likelihood

FieldTypeRequired
evidenceEvidenceYes
given_hProbabilityYes
given_not_hProbabilityYes

Outcome

{
  "oneOf": [
    {
      "additionalProperties": false,
      "properties": {
        "status": {
          "const": "yes",
          "type": "string"
        }
      },
      "required": [
        "status"
      ],
      "type": "object"
    },
    {
      "additionalProperties": false,
      "properties": {
        "status": {
          "const": "no",
          "type": "string"
        }
      },
      "required": [
        "status"
      ],
      "type": "object"
    },
    {
      "additionalProperties": false,
      "properties": {
        "status": {
          "const": "unresolved",
          "type": "string"
        }
      },
      "required": [
        "status"
      ],
      "type": "object"
    },
    {
      "additionalProperties": false,
      "properties": {
        "reason": {
          "type": "string"
        },
        "status": {
          "const": "void",
          "type": "string"
        }
      },
      "required": [
        "status",
        "reason"
      ],
      "type": "object"
    }
  ]
}

Probability

{
  "maximum": 1.0,
  "minimum": 0.0,
  "type": "number"
}

Question

FieldTypeRequired
deadlineintegerYes
idstringYes
no_rulestringYes
opens_atintegerYes
propositionstringYes
resolution_sourcesarray of stringYes
resolve_afterintegerYes
versionintegerYes
void_rulestringYes
yes_rulestringYes

ReferenceClass

FieldTypeRequired
definitionstringYes
failuresintegerYes
inclusion_rulestringYes
observed_fromintegerYes
observed_throughintegerYes
sourcestringYes
successesintegerYes
transfer_concernsarray of stringYes
unresolvedintegerYes

Resolution

FieldTypeRequired
atintegerYes
outcomeOutcomeYes
rationalestringYes
sourcestringYes

Respondent

{
  "oneOf": [
    {
      "additionalProperties": false,
      "properties": {
        "kind": {
          "const": "human",
          "type": "string"
        },
        "pseudonym": {
          "type": "string"
        }
      },
      "required": [
        "kind",
        "pseudonym"
      ],
      "type": "object"
    },
    {
      "additionalProperties": false,
      "properties": {
        "corpus_ref": {
          "type": "string"
        },
        "family": {
          "type": "string"
        },
        "kind": {
          "const": "model",
          "type": "string"
        },
        "prompt_ref": {
          "type": "string"
        },
        "version": {
          "type": "string"
        }
      },
      "required": [
        "kind",
        "family",
        "version",
        "prompt_ref",
        "corpus_ref"
      ],
      "type": "object"
    }
  ]
}

Snapshot

FieldTypeRequired
atintegerYes
probabilityProbabilityYes

WeightedForecast

FieldTypeRequired
probabilityProbabilityYes
weightnumberYes

Operation index

Every command below has an executed example, setup when required, generated field types, and a checked response. All requests use the version 1 envelope.

OperationPurpose
append_forecastAppend a supplied forecast after validating its lifecycle and evidence.
bayesUpdate a binary probability from two conditional evidence likelihoods.
binary_scoreScore one binary forecast with Brier and natural-log loss.
calibrationDiagnose reliability and resolution while preserving unresolved coverage.
categorical_scoreScore a categorical distribution using categorical Brier and the ordered ranked rule.
cluster_bootstrapEstimate a paired score-difference interval by resampling whole event clusters.
compareCompare model and baseline Brier on paired, uniquely identified questions.
conditional_chainMultiply a nonempty chain of conditional event probabilities.
create_questionRegister a frozen question version in the durable journal.
crpsEvaluate an empirical continuous predictive distribution in the outcome’s units.
decisionSelect the action with highest expected utility for a binary forecast.
evaluateEvaluate a declared cohort at common forecast and resolution cutoffs.
event_probabilityEstimate an event probability in a remaining constant-hazard window.
get_ledgerRetrieve the complete, validated journal for one question version.
grid_scoreCarry forecasts forward over a frozen time grid and average within one question.
indicatorValidate a two-branch indicator model and measure expected information gain.
list_questionsList question/version keys using stable lexicographic pagination.
mixtureCombine weighted conditional probabilities over exhaustive scenarios.
panelSummarize supplied human or model elicitation records without inventing missing answers.
poolReturn arithmetic and logit aggregates under declared weights and alpha.
quantileCompute a type-7 empirical quantile with linear interpolation.
recalibrateApply a declared intercept/slope mapping in log-odds space.
reference_classReport empirical counts and a declared Beta-Binomial outside view.
resolve_questionAppend an adjudication from a source in the frozen question contract.
select_calibrationSelect a calibration grid candidate on training outcomes, then score the untouched holdout.
sensitivityCompute a posterior assumption envelope from ordered positive likelihood-ratio bounds.
unionReturn coherent union/intersection bounds and an exact union when overlap is known.
update_forecastCompute and append an evidence-based Bayesian revision atomically.
validate_holdoutValidate declared question, event-cluster and time separation.
value_of_informationMeasure net expected decision value of observing an indicator before acting.

append_forecast

Append a supplied forecast after validating its lifecycle and evidence.

Contract

The initial forecast has no predecessor. Subsequent forecasts must identify the current predecessor, increase as_of, and preserve a nonregressing cutoff. A manual probability is accepted as a declared estimate, not as proof of its method.

Fields

FieldTypeRequired
forecastForecastYes
keyKeyYes
op"append_forecast"Yes

Setup

Start a fresh --memory process and send these requests, one per line, before the example. The same sequence also works with a new SQLite database.

{"version":1,"id":"setup-1","command":{"op":"create_question","question":{"id":"doc-launch","version":1,"proposition":"Release before timestamp 100","opens_at":0,"deadline":100,"resolve_after":110,"yes_rule":"Archive confirms release strictly before 100","no_rule":"Complete archive confirms no qualifying release","void_rule":"Archive permanently unavailable","resolution_sources":["official archive"]}}}

Request

Send this object on one line. It is expanded below for readability.

{
  "version": 1,
  "id": "example",
  "command": {
    "op": "append_forecast",
    "key": {
      "question_id": "doc-launch",
      "version": 1
    },
    "forecast": {
      "id": "f1",
      "as_of": 1,
      "evidence_cutoff": 1,
      "probability": 0.3,
      "previous_id": null,
      "change": "initial",
      "method": "empirical outside view",
      "rationale": "12 of 40 comparable attempts",
      "assumptions": [
        "cases are comparable"
      ],
      "next_review_trigger": "readiness test",
      "evidence": []
    }
  }
}

Response

This response is generated by executing the example against the current binary. Commit timestamps, when present, are shown as <runtime UTC seconds>; the actual protocol returns integer UTC seconds.

{
  "version": 1,
  "id": "example",
  "result": {
    "event": {
      "calculation": null,
      "forecast": {
        "as_of": 1,
        "assumptions": [
          "cases are comparable"
        ],
        "change": "initial",
        "evidence": [],
        "evidence_cutoff": 1,
        "id": "f1",
        "method": "empirical outside view",
        "next_review_trigger": "readiness test",
        "previous_id": null,
        "probability": 0.3,
        "rationale": "12 of 40 comparable attempts"
      },
      "kind": "forecast"
    },
    "key": {
      "question_id": "doc-launch",
      "version": 1
    },
    "recorded_at": "<runtime UTC seconds>",
    "sequence": 2
  }
}

Verification

The documentation gate executes the setup and request, then checks the following result fields against independently specified expectations:

JSON pointer within resultExpected
/sequence2
/event/forecast/probability0.3

Float comparisons use a 1e-12 tolerance. Request schema validation and cross-record validation still apply. See errors and recovery for failure handling.

bayes

Update a binary probability from two conditional evidence likelihoods.

Contract

Likelihoods describe the supplied observation under each hypothesis, conditional on earlier evidence. Impossible evidence is an error. This calculation does not append a journal revision.

Fields

FieldTypeRequired
given_hProbabilityYes
given_not_hProbabilityYes
op"bayes"Yes
priorProbabilityYes

This operation needs no existing question or forecast. Use --memory for a calculation-only session.

Request

Send this object on one line. It is expanded below for readability.

{
  "version": 1,
  "id": "example",
  "command": {
    "op": "bayes",
    "prior": 0.3,
    "given_h": 0.8,
    "given_not_h": 0.2
  }
}

Response

This response is generated by executing the example against the current binary. Commit timestamps, when present, are shown as <runtime UTC seconds>; the actual protocol returns integer UTC seconds.

{
  "version": 1,
  "id": "example",
  "result": {
    "posterior": 0.631578947368421
  }
}

Verification

The documentation gate executes the setup and request, then checks the following result fields against independently specified expectations:

JSON pointer within resultExpected
/posterior0.631578947368421

Float comparisons use a 1e-12 tolerance. Request schema validation and cross-record validation still apply. See errors and recovery for failure handling.

binary_score

Score one binary forecast with Brier and natural-log loss.

Contract

Binary Brier uses the one-component [0,1] convention. Confidently wrong endpoints produce a tagged positive_infinity log loss, never JSON null.

Fields

FieldTypeRequired
op"binary_score"Yes
outcomebooleanYes
probabilityProbabilityYes

This operation needs no existing question or forecast. Use --memory for a calculation-only session.

Request

Send this object on one line. It is expanded below for readability.

{
  "version": 1,
  "id": "example",
  "command": {
    "op": "binary_score",
    "probability": 0.3,
    "outcome": true
  }
}

Response

This response is generated by executing the example against the current binary. Commit timestamps, when present, are shown as <runtime UTC seconds>; the actual protocol returns integer UTC seconds.

{
  "version": 1,
  "id": "example",
  "result": {
    "binary_brier": 0.48999999999999994,
    "log_loss": {
      "kind": "finite",
      "value": 1.2039728043259361
    }
  }
}

Verification

The documentation gate executes the setup and request, then checks the following result fields against independently specified expectations:

JSON pointer within resultExpected
/binary_brier0.49
/log_loss/kind"finite"

Float comparisons use a 1e-12 tolerance. Request schema validation and cross-record validation still apply. See errors and recovery for failure handling.

calibration

Diagnose reliability and resolution while preserving unresolved coverage.

Contract

Use one frozen forecast per unique question ID. Null bins chooses exact groups; a positive bin count chooses equal-width groups. Raw and coarsened Brier are reported separately.

Fields

FieldTypeRequired
binsinteger or nullNo
op"calibration"Yes
recordsarray of CalibrationRecordYes

This operation needs no existing question or forecast. Use --memory for a calculation-only session.

Request

Send this object on one line. It is expanded below for readability.

{
  "version": 1,
  "id": "example",
  "command": {
    "op": "calibration",
    "records": [
      {
        "question_id": "a",
        "probability": 0.2,
        "outcome": false
      },
      {
        "question_id": "b",
        "probability": 0.8,
        "outcome": true
      },
      {
        "question_id": "c",
        "probability": 0.7,
        "outcome": null
      }
    ],
    "bins": null
  }
}

Response

This response is generated by executing the example against the current binary. Commit timestamps, when present, are shown as <runtime UTC seconds>; the actual protocol returns integer UTC seconds.

{
  "version": 1,
  "id": "example",
  "result": {
    "binning_residual": 1.3877787807814457e-17,
    "bins": [
      {
        "count": 1,
        "mean_forecast": 0.2,
        "observed_frequency": 0.0
      },
      {
        "count": 1,
        "mean_forecast": 0.8,
        "observed_frequency": 1.0
      }
    ],
    "coarsened_brier": 0.03999999999999998,
    "excluded_unresolved": 1,
    "raw_brier": 0.039999999999999994,
    "reliability": 0.039999999999999994,
    "resolution": 0.25,
    "resolved": 2,
    "uncertainty": 0.25
  }
}

Verification

The documentation gate executes the setup and request, then checks the following result fields against independently specified expectations:

JSON pointer within resultExpected
/resolved2
/excluded_unresolved1
/raw_brier0.04

Float comparisons use a 1e-12 tolerance. Request schema validation and cross-record validation still apply. See errors and recovery for failure handling.

categorical_score

Score a categorical distribution using categorical Brier and the ordered ranked rule.

Contract

Probabilities must sum to one; outcome is a zero-based index. Brier is on [0,2]. Use the unnormalized ranked score only when category order has meaning.

Fields

FieldTypeRequired
op"categorical_score"Yes
outcomeintegerYes
probabilitiesarray of ProbabilityYes

This operation needs no existing question or forecast. Use --memory for a calculation-only session.

Request

Send this object on one line. It is expanded below for readability.

{
  "version": 1,
  "id": "example",
  "command": {
    "op": "categorical_score",
    "probabilities": [
      0.3,
      0.7
    ],
    "outcome": 0
  }
}

Response

This response is generated by executing the example against the current binary. Commit timestamps, when present, are shown as <runtime UTC seconds>; the actual protocol returns integer UTC seconds.

{
  "version": 1,
  "id": "example",
  "result": {
    "categorical_brier": 0.9799999999999999,
    "ranked_probability_unnormalized": 0.48999999999999994
  }
}

Verification

The documentation gate executes the setup and request, then checks the following result fields against independently specified expectations:

JSON pointer within resultExpected
/categorical_brier0.98
/ranked_probability_unnormalized0.49

Float comparisons use a 1e-12 tolerance. Request schema validation and cross-record validation still apply. See errors and recovery for failure handling.

cluster_bootstrap

Estimate a paired score-difference interval by resampling whole event clusters.

Contract

At least two clusters and 100..100000 replicates are required, capped at 10 million draws. Clusters must be independent under the model. This example intentionally shows a degenerate interval; it is not proof of no uncertainty.

Fields

FieldTypeRequired
casesarray of CaseYes
confidencenumberYes
op"cluster_bootstrap"Yes
resamplesintegerYes
seedintegerYes

This operation needs no existing question or forecast. Use --memory for a calculation-only session.

Request

Send this object on one line. It is expanded below for readability.

{
  "version": 1,
  "id": "example",
  "command": {
    "op": "cluster_bootstrap",
    "cases": [
      {
        "question_id": "a",
        "cluster_id": "a",
        "forecast_at": 1,
        "resolved_at": 5,
        "model": 0.5,
        "baseline": 0.5,
        "outcome": true
      },
      {
        "question_id": "b",
        "cluster_id": "b",
        "forecast_at": 1,
        "resolved_at": 5,
        "model": 0.5,
        "baseline": 0.5,
        "outcome": false
      }
    ],
    "resamples": 100,
    "confidence": 0.95,
    "seed": 7
  }
}

Response

This response is generated by executing the example against the current binary. Commit timestamps, when present, are shown as <runtime UTC seconds>; the actual protocol returns integer UTC seconds.

{
  "version": 1,
  "id": "example",
  "result": {
    "clusters": 2,
    "confidence": 0.95,
    "degenerate": true,
    "difference": 0.0,
    "high": 0.0,
    "low": 0.0,
    "method": "paired_cluster_percentile",
    "questions": 2,
    "resamples": 100,
    "seed": 7,
    "standard_error": 0.0
  }
}

Verification

The documentation gate executes the setup and request, then checks the following result fields against independently specified expectations:

JSON pointer within resultExpected
/clusters2
/low0.0
/high0.0
/degeneratetrue

Float comparisons use a 1e-12 tolerance. Request schema validation and cross-record validation still apply. See errors and recovery for failure handling.

compare

Compare model and baseline Brier on paired, uniquely identified questions.

Contract

Forecasts must precede resolution. The difference is model minus baseline, so negative favors the model. This is a descriptive comparison without an interval.

Fields

FieldTypeRequired
casesarray of CaseYes
op"compare"Yes

This operation needs no existing question or forecast. Use --memory for a calculation-only session.

Request

Send this object on one line. It is expanded below for readability.

{
  "version": 1,
  "id": "example",
  "command": {
    "op": "compare",
    "cases": [
      {
        "question_id": "a",
        "cluster_id": "a",
        "forecast_at": 1,
        "resolved_at": 5,
        "model": 0.5,
        "baseline": 0.5,
        "outcome": true
      },
      {
        "question_id": "b",
        "cluster_id": "b",
        "forecast_at": 1,
        "resolved_at": 5,
        "model": 0.5,
        "baseline": 0.5,
        "outcome": false
      }
    ]
  }
}

Response

This response is generated by executing the example against the current binary. Commit timestamps, when present, are shown as <runtime UTC seconds>; the actual protocol returns integer UTC seconds.

{
  "version": 1,
  "id": "example",
  "result": {
    "baseline_brier": 0.25,
    "difference": 0.0,
    "model_brier": 0.25,
    "questions": 2
  }
}

Verification

The documentation gate executes the setup and request, then checks the following result fields against independently specified expectations:

JSON pointer within resultExpected
/questions2
/model_brier0.25
/difference0.0

Float comparisons use a 1e-12 tolerance. Request schema validation and cross-record validation still apply. See errors and recovery for failure handling.

conditional_chain

Multiply a nonempty chain of conditional event probabilities.

Contract

Every later factor must condition on relevant preceding events. Supplying marginals invents independence unless that assumption has been justified separately.

Fields

FieldTypeRequired
factorsarray of ProbabilityYes
op"conditional_chain"Yes

This operation needs no existing question or forecast. Use --memory for a calculation-only session.

Request

Send this object on one line. It is expanded below for readability.

{
  "version": 1,
  "id": "example",
  "command": {
    "op": "conditional_chain",
    "factors": [
      0.8,
      0.6
    ]
  }
}

Response

This response is generated by executing the example against the current binary. Commit timestamps, when present, are shown as <runtime UTC seconds>; the actual protocol returns integer UTC seconds.

{
  "version": 1,
  "id": "example",
  "result": {
    "probability": 0.48
  }
}

Verification

The documentation gate executes the setup and request, then checks the following result fields against independently specified expectations:

JSON pointer within resultExpected
/probability0.48

Float comparisons use a 1e-12 tolerance. Request schema validation and cross-record validation still apply. See errors and recovery for failure handling.

create_question

Register a frozen question version in the durable journal.

Contract

Requires a resolvable contract with ordered timestamps. Structural validation does not prove the prose predicate is complete. Existing question/version keys cannot be overwritten.

Fields

FieldTypeRequired
op"create_question"Yes
questionQuestionYes

This operation needs no existing question or forecast. Use --memory for a calculation-only session.

Request

Send this object on one line. It is expanded below for readability.

{
  "version": 1,
  "id": "example",
  "command": {
    "op": "create_question",
    "question": {
      "id": "doc-launch",
      "version": 1,
      "proposition": "Release before timestamp 100",
      "opens_at": 0,
      "deadline": 100,
      "resolve_after": 110,
      "yes_rule": "Archive confirms release strictly before 100",
      "no_rule": "Complete archive confirms no qualifying release",
      "void_rule": "Archive permanently unavailable",
      "resolution_sources": [
        "official archive"
      ]
    }
  }
}

Response

This response is generated by executing the example against the current binary. Commit timestamps, when present, are shown as <runtime UTC seconds>; the actual protocol returns integer UTC seconds.

{
  "version": 1,
  "id": "example",
  "result": {
    "event": {
      "kind": "question",
      "question": {
        "deadline": 100,
        "id": "doc-launch",
        "no_rule": "Complete archive confirms no qualifying release",
        "opens_at": 0,
        "proposition": "Release before timestamp 100",
        "resolution_sources": [
          "official archive"
        ],
        "resolve_after": 110,
        "version": 1,
        "void_rule": "Archive permanently unavailable",
        "yes_rule": "Archive confirms release strictly before 100"
      }
    },
    "key": {
      "question_id": "doc-launch",
      "version": 1
    },
    "recorded_at": "<runtime UTC seconds>",
    "sequence": 1
  }
}

Verification

The documentation gate executes the setup and request, then checks the following result fields against independently specified expectations:

JSON pointer within resultExpected
/sequence1
/event/kind"question"

Float comparisons use a 1e-12 tolerance. Request schema validation and cross-record validation still apply. See errors and recovery for failure handling.

crps

Evaluate an empirical continuous predictive distribution in the outcome’s units.

Contract

Samples must be nonempty and finite. Sorting gives O(n log n) computation instead of a quadratic pairwise matrix. Repeated sample values are allowed.

Fields

FieldTypeRequired
op"crps"Yes
outcomenumberYes
samplesarray of numberYes

This operation needs no existing question or forecast. Use --memory for a calculation-only session.

Request

Send this object on one line. It is expanded below for readability.

{
  "version": 1,
  "id": "example",
  "command": {
    "op": "crps",
    "samples": [
      1.0,
      2.0,
      3.0
    ],
    "outcome": 2.0
  }
}

Response

This response is generated by executing the example against the current binary. Commit timestamps, when present, are shown as <runtime UTC seconds>; the actual protocol returns integer UTC seconds.

{
  "version": 1,
  "id": "example",
  "result": {
    "empirical_crps": 0.2222222222222222
  }
}

Verification

The documentation gate executes the setup and request, then checks the following result fields against independently specified expectations:

JSON pointer within resultExpected
/empirical_crps0.2222222222222222

Float comparisons use a 1e-12 tolerance. Request schema validation and cross-record validation still apply. See errors and recovery for failure handling.

decision

Select the action with highest expected utility for a binary forecast.

Contract

Utilities are finite and share a scale. The result uses a zero-based action index. Exact ties retain the first action; the library does not choose an organization’s utility function.

Fields

FieldTypeRequired
actionsarray of ActionYes
op"decision"Yes
probabilityProbabilityYes

This operation needs no existing question or forecast. Use --memory for a calculation-only session.

Request

Send this object on one line. It is expanded below for readability.

{
  "version": 1,
  "id": "example",
  "command": {
    "op": "decision",
    "actions": [
      {
        "if_yes": 10.0,
        "if_no": -10.0
      },
      {
        "if_yes": 0.0,
        "if_no": 0.0
      }
    ],
    "probability": 0.8
  }
}

Response

This response is generated by executing the example against the current binary. Commit timestamps, when present, are shown as <runtime UTC seconds>; the actual protocol returns integer UTC seconds.

{
  "version": 1,
  "id": "example",
  "result": {
    "action_index": 0,
    "expected_utility": 6.0
  }
}

Verification

The documentation gate executes the setup and request, then checks the following result fields against independently specified expectations:

JSON pointer within resultExpected
/action_index0
/expected_utility6.0

Float comparisons use a 1e-12 tolerance. Request schema validation and cross-record validation still apply. See errors and recovery for failure handling.

evaluate

Evaluate a declared cohort at common forecast and resolution cutoffs.

Contract

The latest available forecast is selected per question; all exclusions are reported. One version per question is allowed. This uses caller-declared historical timestamps and is not proof of prospective performance.

Fields

FieldTypeRequired
baselineProbabilityYes
baseline_as_ofintegerYes
forecast_cutoffintegerYes
keysarray of KeyYes
op"evaluate"Yes
resolution_cutoffintegerYes

Setup

Start a fresh --memory process and send these requests, one per line, before the example. The same sequence also works with a new SQLite database.

{"version":1,"id":"setup-1","command":{"op":"create_question","question":{"id":"doc-launch","version":1,"proposition":"Release before timestamp 100","opens_at":0,"deadline":100,"resolve_after":110,"yes_rule":"Archive confirms release strictly before 100","no_rule":"Complete archive confirms no qualifying release","void_rule":"Archive permanently unavailable","resolution_sources":["official archive"]}}}
{"version":1,"id":"setup-2","command":{"op":"append_forecast","key":{"question_id":"doc-launch","version":1},"forecast":{"id":"f1","as_of":1,"evidence_cutoff":1,"probability":0.3,"previous_id":null,"change":"initial","method":"empirical outside view","rationale":"12 of 40 comparable attempts","assumptions":["cases are comparable"],"next_review_trigger":"readiness test","evidence":[]}}}
{"version":1,"id":"setup-3","command":{"op":"update_forecast","key":{"question_id":"doc-launch","version":1},"revision":{"id":"f2","previous_id":"f1","as_of":20,"evidence_cutoff":20,"rationale":"diagnostic readiness evidence","assumptions":["illustrative likelihoods"],"next_review_trigger":"dependency completion","likelihoods":[{"evidence":{"id":"readiness","origin_id":"test-1","locator":"supplied:test-1","claim":"Readiness test passed","published_at":10,"observed_at":9},"given_h":0.8,"given_not_h":0.2}]}}}
{"version":1,"id":"setup-4","command":{"op":"resolve_question","key":{"question_id":"doc-launch","version":1},"resolution":{"at":120,"source":"official archive","rationale":"synthetic archive confirms release","outcome":{"status":"yes"}}}}

Request

Send this object on one line. It is expanded below for readability.

{
  "version": 1,
  "id": "example",
  "command": {
    "op": "evaluate",
    "keys": [
      {
        "question_id": "doc-launch",
        "version": 1
      }
    ],
    "forecast_cutoff": 30,
    "resolution_cutoff": 120,
    "baseline": 0.3,
    "baseline_as_of": 0
  }
}

Response

This response is generated by executing the example against the current binary. Commit timestamps, when present, are shown as <runtime UTC seconds>; the actual protocol returns integer UTC seconds.

{
  "version": 1,
  "id": "example",
  "result": {
    "baseline": 0.3,
    "baseline_as_of": 0,
    "eligible": 1,
    "evaluation_kind": "declared_timestamp_snapshot",
    "excluded": 0,
    "exclusions": [],
    "forecast_cutoff": 30,
    "resolution_cutoff": 120,
    "rows": [
      {
        "as_of": 20,
        "baseline_brier": 0.48999999999999994,
        "binary_brier": 0.1357340720221607,
        "forecast_id": "f2",
        "key": {
          "question_id": "doc-launch",
          "version": 1
        },
        "outcome": true,
        "probability": 0.631578947368421
      }
    ],
    "scored": 1,
    "summary": {
      "baseline_brier": 0.48999999999999994,
      "brier_skill": 0.7229916897506924,
      "calibration": {
        "binning_residual": 0.0,
        "bins": [
          {
            "count": 1,
            "mean_forecast": 0.631578947368421,
            "observed_frequency": 1.0
          }
        ],
        "coarsened_brier": 0.1357340720221607,
        "excluded_unresolved": 0,
        "raw_brier": 0.1357340720221607,
        "reliability": 0.1357340720221607,
        "resolution": 0.0,
        "resolved": 1,
        "uncertainty": 0.0
      },
      "difference": -0.35426592797783923,
      "model_brier": 0.1357340720221607
    }
  }
}

Verification

The documentation gate executes the setup and request, then checks the following result fields against independently specified expectations:

JSON pointer within resultExpected
/scored1
/excluded0
/rows/0/forecast_id"f2"
/summary/model_brier0.13573407202216065

Float comparisons use a 1e-12 tolerance. Request schema validation and cross-record validation still apply. See errors and recovery for failure handling.

event_probability

Estimate an event probability in a remaining constant-hazard window.

Contract

Rate and duration use compatible units and are finite and nonnegative. The Poisson/constant-hazard assumption is explicit; bursty or changing hazards need another model.

Fields

FieldTypeRequired
durationnumberYes
op"event_probability"Yes
ratenumberYes

This operation needs no existing question or forecast. Use --memory for a calculation-only session.

Request

Send this object on one line. It is expanded below for readability.

{
  "version": 1,
  "id": "example",
  "command": {
    "op": "event_probability",
    "rate": 0.1,
    "duration": 10.0
  }
}

Response

This response is generated by executing the example against the current binary. Commit timestamps, when present, are shown as <runtime UTC seconds>; the actual protocol returns integer UTC seconds.

{
  "version": 1,
  "id": "example",
  "result": {
    "model": "constant_hazard_poisson",
    "probability": 0.6321205588285577
  }
}

Verification

The documentation gate executes the setup and request, then checks the following result fields against independently specified expectations:

JSON pointer within resultExpected
/probability0.6321205588285577

Float comparisons use a 1e-12 tolerance. Request schema validation and cross-record validation still apply. See errors and recovery for failure handling.

get_ledger

Retrieve the complete, validated journal for one question version.

Contract

Reloading checks sequence continuity, question identity and lifecycle rules, and recomputes stored Bayesian audits. A missing key is an error; direct database tampering is not cryptographically authenticated.

Fields

FieldTypeRequired
keyKeyYes
op"get_ledger"Yes

Setup

Start a fresh --memory process and send these requests, one per line, before the example. The same sequence also works with a new SQLite database.

{"version":1,"id":"setup-1","command":{"op":"create_question","question":{"id":"doc-launch","version":1,"proposition":"Release before timestamp 100","opens_at":0,"deadline":100,"resolve_after":110,"yes_rule":"Archive confirms release strictly before 100","no_rule":"Complete archive confirms no qualifying release","void_rule":"Archive permanently unavailable","resolution_sources":["official archive"]}}}
{"version":1,"id":"setup-2","command":{"op":"append_forecast","key":{"question_id":"doc-launch","version":1},"forecast":{"id":"f1","as_of":1,"evidence_cutoff":1,"probability":0.3,"previous_id":null,"change":"initial","method":"empirical outside view","rationale":"12 of 40 comparable attempts","assumptions":["cases are comparable"],"next_review_trigger":"readiness test","evidence":[]}}}
{"version":1,"id":"setup-3","command":{"op":"update_forecast","key":{"question_id":"doc-launch","version":1},"revision":{"id":"f2","previous_id":"f1","as_of":20,"evidence_cutoff":20,"rationale":"diagnostic readiness evidence","assumptions":["illustrative likelihoods"],"next_review_trigger":"dependency completion","likelihoods":[{"evidence":{"id":"readiness","origin_id":"test-1","locator":"supplied:test-1","claim":"Readiness test passed","published_at":10,"observed_at":9},"given_h":0.8,"given_not_h":0.2}]}}}

Request

Send this object on one line. It is expanded below for readability.

{
  "version": 1,
  "id": "example",
  "command": {
    "op": "get_ledger",
    "key": {
      "question_id": "doc-launch",
      "version": 1
    }
  }
}

Response

This response is generated by executing the example against the current binary. Commit timestamps, when present, are shown as <runtime UTC seconds>; the actual protocol returns integer UTC seconds.

{
  "version": 1,
  "id": "example",
  "result": [
    {
      "event": {
        "kind": "question",
        "question": {
          "deadline": 100,
          "id": "doc-launch",
          "no_rule": "Complete archive confirms no qualifying release",
          "opens_at": 0,
          "proposition": "Release before timestamp 100",
          "resolution_sources": [
            "official archive"
          ],
          "resolve_after": 110,
          "version": 1,
          "void_rule": "Archive permanently unavailable",
          "yes_rule": "Archive confirms release strictly before 100"
        }
      },
      "recorded_at": "<runtime UTC seconds>",
      "sequence": 1
    },
    {
      "event": {
        "calculation": null,
        "forecast": {
          "as_of": 1,
          "assumptions": [
            "cases are comparable"
          ],
          "change": "initial",
          "evidence": [],
          "evidence_cutoff": 1,
          "id": "f1",
          "method": "empirical outside view",
          "next_review_trigger": "readiness test",
          "previous_id": null,
          "probability": 0.3,
          "rationale": "12 of 40 comparable attempts"
        },
        "kind": "forecast"
      },
      "recorded_at": "<runtime UTC seconds>",
      "sequence": 2
    },
    {
      "event": {
        "calculation": {
          "likelihoods": [
            {
              "evidence": {
                "claim": "Readiness test passed",
                "id": "readiness",
                "locator": "supplied:test-1",
                "observed_at": 9,
                "origin_id": "test-1",
                "published_at": 10
              },
              "given_h": 0.8,
              "given_not_h": 0.2
            }
          ],
          "prior": 0.3
        },
        "forecast": {
          "as_of": 20,
          "assumptions": [
            "illustrative likelihoods"
          ],
          "change": "evidence",
          "evidence": [
            {
              "claim": "Readiness test passed",
              "id": "readiness",
              "locator": "supplied:test-1",
              "observed_at": 9,
              "origin_id": "test-1",
              "published_at": 10
            }
          ],
          "evidence_cutoff": 20,
          "id": "f2",
          "method": "conditional Bayesian likelihood update",
          "next_review_trigger": "dependency completion",
          "previous_id": "f1",
          "probability": 0.631578947368421,
          "rationale": "diagnostic readiness evidence"
        },
        "kind": "forecast"
      },
      "recorded_at": "<runtime UTC seconds>",
      "sequence": 3
    }
  ]
}

Verification

The documentation gate executes the setup and request, then checks the following result fields against independently specified expectations:

JSON pointer within resultExpected
/2/sequence3
/2/event/forecast/probability0.631578947368421

Float comparisons use a 1e-12 tolerance. Request schema validation and cross-record validation still apply. See errors and recovery for failure handling.

grid_score

Carry forecasts forward over a frozen time grid and average within one question.

Contract

Updates and grid times must increase strictly. A forecast must exist at the first grid point. Extra posts between grid points do not gain additional weight; the caller chooses a valid event-window grid.

Fields

FieldTypeRequired
gridarray of integerYes
op"grid_score"Yes
outcomebooleanYes
updatesarray of SnapshotYes

This operation needs no existing question or forecast. Use --memory for a calculation-only session.

Request

Send this object on one line. It is expanded below for readability.

{
  "version": 1,
  "id": "example",
  "command": {
    "op": "grid_score",
    "updates": [
      {
        "at": 0,
        "probability": 0.2
      },
      {
        "at": 5,
        "probability": 0.7
      }
    ],
    "grid": [
      0,
      10,
      20
    ],
    "outcome": true
  }
}

Response

This response is generated by executing the example against the current binary. Commit timestamps, when present, are shown as <runtime UTC seconds>; the actual protocol returns integer UTC seconds.

{
  "version": 1,
  "id": "example",
  "result": {
    "binary_brier": 0.27333333333333343,
    "grid_points": 3
  }
}

Verification

The documentation gate executes the setup and request, then checks the following result fields against independently specified expectations:

JSON pointer within resultExpected
/binary_brier0.2733333333333333
/grid_points3

Float comparisons use a 1e-12 tolerance. Request schema validation and cross-record validation still apply. See errors and recovery for failure handling.

indicator

Validate a two-branch indicator model and measure expected information gain.

Contract

The implied prior must match the stated prior within the declared tolerance. Information is in nats; sensitivities are ordered chance, if_yes, if_no. Association is not an intervention effect.

Fields

FieldTypeRequired
indicatorIndicatorYes
op"indicator"Yes
stated_priorProbabilityYes
tolerancenumberYes

This operation needs no existing question or forecast. Use --memory for a calculation-only session.

Request

Send this object on one line. It is expanded below for readability.

{
  "version": 1,
  "id": "example",
  "command": {
    "op": "indicator",
    "indicator": {
      "chance": 0.5,
      "if_yes": 0.8,
      "if_no": 0.3
    },
    "stated_prior": 0.55,
    "tolerance": 1e-12
  }
}

Response

This response is generated by executing the example against the current binary. Commit timestamps, when present, are shown as <runtime UTC seconds>; the actual protocol returns integer UTC seconds.

{
  "version": 1,
  "id": "example",
  "result": {
    "implied_prior": 0.55,
    "information_gain_nats": 0.1325054509170478,
    "sensitivity": [
      0.5,
      0.5,
      0.5
    ]
  }
}

Verification

The documentation gate executes the setup and request, then checks the following result fields against independently specified expectations:

JSON pointer within resultExpected
/implied_prior0.55
/sensitivity/00.5

Float comparisons use a 1e-12 tolerance. Request schema validation and cross-record validation still apply. See errors and recovery for failure handling.

list_questions

List question/version keys using stable lexicographic pagination.

Contract

Limit must be 1..1000. For the next page, supply the last returned key as after. Listing identifies registered keys; get_ledger validates their complete histories.

Fields

FieldTypeRequired
afterKey or nullNo
limitintegerYes
op"list_questions"Yes

Setup

Start a fresh --memory process and send these requests, one per line, before the example. The same sequence also works with a new SQLite database.

{"version":1,"id":"setup-1","command":{"op":"create_question","question":{"id":"doc-launch","version":1,"proposition":"Release before timestamp 100","opens_at":0,"deadline":100,"resolve_after":110,"yes_rule":"Archive confirms release strictly before 100","no_rule":"Complete archive confirms no qualifying release","void_rule":"Archive permanently unavailable","resolution_sources":["official archive"]}}}

Request

Send this object on one line. It is expanded below for readability.

{
  "version": 1,
  "id": "example",
  "command": {
    "op": "list_questions",
    "after": null,
    "limit": 10
  }
}

Response

This response is generated by executing the example against the current binary. Commit timestamps, when present, are shown as <runtime UTC seconds>; the actual protocol returns integer UTC seconds.

{
  "version": 1,
  "id": "example",
  "result": {
    "questions": [
      {
        "question_id": "doc-launch",
        "version": 1
      }
    ]
  }
}

Verification

The documentation gate executes the setup and request, then checks the following result fields against independently specified expectations:

JSON pointer within resultExpected
/questions/0/question_id"doc-launch"
/questions/0/version1

Float comparisons use a 1e-12 tolerance. Request schema validation and cross-record validation still apply. See errors and recovery for failure handling.

mixture

Combine weighted conditional probabilities over exhaustive scenarios.

Contract

Each pair is [scenario_weight, conditional_probability]. Weights sum to one within tolerance. The caller must establish that scenarios are disjoint and exhaustive.

Fields

FieldTypeRequired
branchesarray of array of tuple [Probability, Probability]Yes
op"mixture"Yes

This operation needs no existing question or forecast. Use --memory for a calculation-only session.

Request

Send this object on one line. It is expanded below for readability.

{
  "version": 1,
  "id": "example",
  "command": {
    "op": "mixture",
    "branches": [
      [
        0.5,
        0.2
      ],
      [
        0.5,
        0.8
      ]
    ]
  }
}

Response

This response is generated by executing the example against the current binary. Commit timestamps, when present, are shown as <runtime UTC seconds>; the actual protocol returns integer UTC seconds.

{
  "version": 1,
  "id": "example",
  "result": {
    "probability": 0.5
  }
}

Verification

The documentation gate executes the setup and request, then checks the following result fields against independently specified expectations:

JSON pointer within resultExpected
/probability0.5

Float comparisons use a 1e-12 tolerance. Request schema validation and cross-record validation still apply. See errors and recovery for failure handling.

panel

Summarize supplied human or model elicitation records without inventing missing answers.

Contract

Respondent IDs must be unique. At least one probability is needed. The returned range is disagreement, not calibrated uncertainty; this command does not contact respondents.

Fields

FieldTypeRequired
op"panel"Yes
responsesarray of ElicitationYes

This operation needs no existing question or forecast. Use --memory for a calculation-only session.

Request

Send this object on one line. It is expanded below for readability.

{
  "version": 1,
  "id": "example",
  "command": {
    "op": "panel",
    "responses": [
      {
        "respondent_id": "a",
        "respondent": {
          "kind": "human",
          "pseudonym": "A"
        },
        "probability": 0.3,
        "supporting_reason": "relevant reference cases",
        "opposing_reason": "integration risk"
      },
      {
        "respondent_id": "b",
        "respondent": {
          "kind": "human",
          "pseudonym": "B"
        },
        "probability": null,
        "supporting_reason": "",
        "opposing_reason": ""
      }
    ]
  }
}

Response

This response is generated by executing the example against the current binary. Commit timestamps, when present, are shown as <runtime UTC seconds>; the actual protocol returns integer UTC seconds.

{
  "version": 1,
  "id": "example",
  "result": {
    "invited": 2,
    "maximum": 0.3,
    "median": 0.3,
    "minimum": 0.3,
    "missing": 1,
    "responded": 1
  }
}

Verification

The documentation gate executes the setup and request, then checks the following result fields against independently specified expectations:

JSON pointer within resultExpected
/invited2
/responded1
/missing1
/median0.3

Float comparisons use a 1e-12 tolerance. Request schema validation and cross-record validation still apply. See errors and recovery for failure handling.

pool

Return arithmetic and logit aggregates under declared weights and alpha.

Contract

Weights are finite and nonnegative with positive total weight. Positive-weight endpoints are rejected. Alpha is explicit and is not fitted by this operation; missing panel responses require a separate coverage report.

Fields

FieldTypeRequired
alphanumberYes
forecastsarray of WeightedForecastYes
op"pool"Yes

This operation needs no existing question or forecast. Use --memory for a calculation-only session.

Request

Send this object on one line. It is expanded below for readability.

{
  "version": 1,
  "id": "example",
  "command": {
    "op": "pool",
    "forecasts": [
      {
        "probability": 0.2,
        "weight": 1.0
      },
      {
        "probability": 0.8,
        "weight": 1.0
      }
    ],
    "alpha": 1.0
  }
}

Response

This response is generated by executing the example against the current binary. Commit timestamps, when present, are shown as <runtime UTC seconds>; the actual protocol returns integer UTC seconds.

{
  "version": 1,
  "id": "example",
  "result": {
    "alpha": 1.0,
    "arithmetic": 0.5,
    "contributors": 2,
    "logit": 0.5
  }
}

Verification

The documentation gate executes the setup and request, then checks the following result fields against independently specified expectations:

JSON pointer within resultExpected
/arithmetic0.5
/logit0.5
/contributors2

Float comparisons use a 1e-12 tolerance. Request schema validation and cross-record validation still apply. See errors and recovery for failure handling.

quantile

Compute a type-7 empirical quantile with linear interpolation.

Contract

The output describes the empirical distribution, not a confidence interval for a parameter. Samples must be finite and nonempty; q includes endpoints zero and one.

Fields

FieldTypeRequired
op"quantile"Yes
qProbabilityYes
samplesarray of numberYes

This operation needs no existing question or forecast. Use --memory for a calculation-only session.

Request

Send this object on one line. It is expanded below for readability.

{
  "version": 1,
  "id": "example",
  "command": {
    "op": "quantile",
    "samples": [
      1.0,
      1.1,
      1.2,
      1.5,
      4.0
    ],
    "q": 0.8
  }
}

Response

This response is generated by executing the example against the current binary. Commit timestamps, when present, are shown as <runtime UTC seconds>; the actual protocol returns integer UTC seconds.

{
  "version": 1,
  "id": "example",
  "result": {
    "method": "linear_type_7",
    "quantile": 2.0000000000000004
  }
}

Verification

The documentation gate executes the setup and request, then checks the following result fields against independently specified expectations:

JSON pointer within resultExpected
/quantile2.0
/method"linear_type_7"

Float comparisons use a 1e-12 tolerance. Request schema validation and cross-record validation still apply. See errors and recovery for failure handling.

recalibrate

Apply a declared intercept/slope mapping in log-odds space.

Contract

This operation does not fit coefficients. Probability must be interior; intercept and nonnegative slope must be finite. Validate fitted parameters on separate outcomes.

Fields

FieldTypeRequired
interceptnumberYes
op"recalibrate"Yes
probabilityProbabilityYes
slopenumberYes

This operation needs no existing question or forecast. Use --memory for a calculation-only session.

Request

Send this object on one line. It is expanded below for readability.

{
  "version": 1,
  "id": "example",
  "command": {
    "op": "recalibrate",
    "probability": 0.8,
    "intercept": 0.0,
    "slope": 0.0
  }
}

Response

This response is generated by executing the example against the current binary. Commit timestamps, when present, are shown as <runtime UTC seconds>; the actual protocol returns integer UTC seconds.

{
  "version": 1,
  "id": "example",
  "result": {
    "probability": 0.5
  }
}

Verification

The documentation gate executes the setup and request, then checks the following result fields against independently specified expectations:

JSON pointer within resultExpected
/probability0.5

Float comparisons use a 1e-12 tolerance. Request schema validation and cross-record validation still apply. See errors and recovery for failure handling.

reference_class

Report empirical counts and a declared Beta-Binomial outside view.

Contract

Unresolved cases are not failures. No resolved cases produces null empirical_rate while the prior-driven posterior remains available. Comparability and exchangeability are caller assumptions.

Fields

FieldTypeRequired
alphanumberYes
betanumberYes
classReferenceClassYes
op"reference_class"Yes

This operation needs no existing question or forecast. Use --memory for a calculation-only session.

Request

Send this object on one line. It is expanded below for readability.

{
  "version": 1,
  "id": "example",
  "command": {
    "op": "reference_class",
    "class": {
      "definition": "comparable releases",
      "inclusion_rule": "all registered attempts",
      "source": "supplied:reference",
      "observed_from": 0,
      "observed_through": 10,
      "successes": 12,
      "failures": 28,
      "unresolved": 2,
      "transfer_concerns": [
        "team composition changed"
      ]
    },
    "alpha": 1.0,
    "beta": 1.0
  }
}

Response

This response is generated by executing the example against the current binary. Commit timestamps, when present, are shown as <runtime UTC seconds>; the actual protocol returns integer UTC seconds.

{
  "version": 1,
  "id": "example",
  "result": {
    "empirical_rate": 0.3,
    "posterior_alpha": 13.0,
    "posterior_beta": 29.0,
    "posterior_mean": 0.30952380952380953,
    "posterior_variance": 0.004970205136318093,
    "resolved": 40,
    "transfer_concerns": [
      "team composition changed"
    ],
    "unresolved": 2
  }
}

Verification

The documentation gate executes the setup and request, then checks the following result fields against independently specified expectations:

JSON pointer within resultExpected
/empirical_rate0.3
/posterior_mean0.30952380952380953
/unresolved2

Float comparisons use a 1e-12 tolerance. Request schema validation and cross-record validation still apply. See errors and recovery for failure handling.

resolve_question

Append an adjudication from a source in the frozen question contract.

Contract

Resolution must occur at or after resolve_after. Unresolved may be followed by a later adjudication. Yes, no and void are terminal; void requires a reason. The caller establishes the factual outcome.

Fields

FieldTypeRequired
keyKeyYes
op"resolve_question"Yes
resolutionResolutionYes

Setup

Start a fresh --memory process and send these requests, one per line, before the example. The same sequence also works with a new SQLite database.

{"version":1,"id":"setup-1","command":{"op":"create_question","question":{"id":"doc-launch","version":1,"proposition":"Release before timestamp 100","opens_at":0,"deadline":100,"resolve_after":110,"yes_rule":"Archive confirms release strictly before 100","no_rule":"Complete archive confirms no qualifying release","void_rule":"Archive permanently unavailable","resolution_sources":["official archive"]}}}
{"version":1,"id":"setup-2","command":{"op":"append_forecast","key":{"question_id":"doc-launch","version":1},"forecast":{"id":"f1","as_of":1,"evidence_cutoff":1,"probability":0.3,"previous_id":null,"change":"initial","method":"empirical outside view","rationale":"12 of 40 comparable attempts","assumptions":["cases are comparable"],"next_review_trigger":"readiness test","evidence":[]}}}

Request

Send this object on one line. It is expanded below for readability.

{
  "version": 1,
  "id": "example",
  "command": {
    "op": "resolve_question",
    "key": {
      "question_id": "doc-launch",
      "version": 1
    },
    "resolution": {
      "at": 120,
      "source": "official archive",
      "rationale": "synthetic archive confirms release",
      "outcome": {
        "status": "yes"
      }
    }
  }
}

Response

This response is generated by executing the example against the current binary. Commit timestamps, when present, are shown as <runtime UTC seconds>; the actual protocol returns integer UTC seconds.

{
  "version": 1,
  "id": "example",
  "result": {
    "event": {
      "kind": "resolution",
      "resolution": {
        "at": 120,
        "outcome": {
          "status": "yes"
        },
        "rationale": "synthetic archive confirms release",
        "source": "official archive"
      }
    },
    "key": {
      "question_id": "doc-launch",
      "version": 1
    },
    "recorded_at": "<runtime UTC seconds>",
    "sequence": 3
  }
}

Verification

The documentation gate executes the setup and request, then checks the following result fields against independently specified expectations:

JSON pointer within resultExpected
/sequence3
/event/resolution/outcome/status"yes"

Float comparisons use a 1e-12 tolerance. Request schema validation and cross-record validation still apply. See errors and recovery for failure handling.

select_calibration

Select a calibration grid candidate on training outcomes, then score the untouched holdout.

Contract

Temporal and cluster separation are checked before fitting. Exact ties choose the first candidate. The selected candidate is not changed after seeing its holdout performance; repeated tuning on that holdout would consume it.

Fields

FieldTypeRequired
candidatesarray of CalibrationCandidateYes
cutoffintegerYes
holdoutarray of CaseYes
op"select_calibration"Yes
trainingarray of CaseYes

This operation needs no existing question or forecast. Use --memory for a calculation-only session.

Request

Send this object on one line. It is expanded below for readability.

{
  "version": 1,
  "id": "example",
  "command": {
    "op": "select_calibration",
    "training": [
      {
        "question_id": "a",
        "cluster_id": "a",
        "forecast_at": 1,
        "resolved_at": 5,
        "model": 0.5,
        "baseline": 0.5,
        "outcome": true
      },
      {
        "question_id": "b",
        "cluster_id": "b",
        "forecast_at": 1,
        "resolved_at": 5,
        "model": 0.5,
        "baseline": 0.5,
        "outcome": false
      }
    ],
    "holdout": [
      {
        "question_id": "held-a",
        "cluster_id": "held-a",
        "forecast_at": 11,
        "resolved_at": 20,
        "model": 0.5,
        "baseline": 0.5,
        "outcome": true
      },
      {
        "question_id": "held-b",
        "cluster_id": "held-b",
        "forecast_at": 11,
        "resolved_at": 20,
        "model": 0.5,
        "baseline": 0.5,
        "outcome": false
      }
    ],
    "cutoff": 10,
    "candidates": [
      {
        "intercept": 0.0,
        "slope": 1.0
      },
      {
        "intercept": 0.0,
        "slope": 0.0
      }
    ]
  }
}

Response

This response is generated by executing the example against the current binary. Commit timestamps, when present, are shown as <runtime UTC seconds>; the actual protocol returns integer UTC seconds.

{
  "version": 1,
  "id": "example",
  "result": {
    "baseline_holdout_brier": 0.25,
    "calibrated_holdout_brier": 0.25,
    "holdout_questions": 2,
    "raw_holdout_brier": 0.25,
    "selected": {
      "intercept": 0.0,
      "slope": 1.0
    },
    "training_brier": [
      0.25,
      0.25
    ],
    "training_questions": 2
  }
}

Verification

The documentation gate executes the setup and request, then checks the following result fields against independently specified expectations:

JSON pointer within resultExpected
/selected/slope1.0
/training_questions2
/calibrated_holdout_brier0.25

Float comparisons use a 1e-12 tolerance. Request schema validation and cross-record validation still apply. See errors and recovery for failure handling.

sensitivity

Compute a posterior assumption envelope from ordered positive likelihood-ratio bounds.

Contract

The interval is sensitivity to supplied assumptions, not a confidence or credible interval. Finite positive ratios and ordered bounds are required.

Fields

FieldTypeRequired
high_lrnumberYes
low_lrnumberYes
op"sensitivity"Yes
priorProbabilityYes

This operation needs no existing question or forecast. Use --memory for a calculation-only session.

Request

Send this object on one line. It is expanded below for readability.

{
  "version": 1,
  "id": "example",
  "command": {
    "op": "sensitivity",
    "prior": 0.3,
    "low_lr": 2.0,
    "high_lr": 6.0
  }
}

Response

This response is generated by executing the example against the current binary. Commit timestamps, when present, are shown as <runtime UTC seconds>; the actual protocol returns integer UTC seconds.

{
  "version": 1,
  "id": "example",
  "result": {
    "kind": "assumption_envelope",
    "posterior_high": 0.72,
    "posterior_low": 0.4615384615384615
  }
}

Verification

The documentation gate executes the setup and request, then checks the following result fields against independently specified expectations:

JSON pointer within resultExpected
/kind"assumption_envelope"
/posterior_low0.46153846153846156
/posterior_high0.72

Float comparisons use a 1e-12 tolerance. Request schema validation and cross-record validation still apply. See errors and recovery for failure handling.

union

Return coherent union/intersection bounds and an exact union when overlap is known.

Contract

Unknown intersection leaves union null; bounds are still available. A supplied intersection outside Frechet bounds is rejected rather than silently corrected.

Fields

FieldTypeRequired
aProbabilityYes
bProbabilityYes
intersectionProbability or nullNo
op"union"Yes

This operation needs no existing question or forecast. Use --memory for a calculation-only session.

Request

Send this object on one line. It is expanded below for readability.

{
  "version": 1,
  "id": "example",
  "command": {
    "op": "union",
    "a": 0.4,
    "b": 0.5,
    "intersection": null
  }
}

Response

This response is generated by executing the example against the current binary. Commit timestamps, when present, are shown as <runtime UTC seconds>; the actual protocol returns integer UTC seconds.

{
  "version": 1,
  "id": "example",
  "result": {
    "intersection_bounds": [
      0.0,
      0.4
    ],
    "union": null,
    "union_bounds": [
      0.5,
      0.9
    ]
  }
}

Verification

The documentation gate executes the setup and request, then checks the following result fields against independently specified expectations:

JSON pointer within resultExpected
/unionnull
/union_bounds/00.5
/union_bounds/10.9

Float comparisons use a 1e-12 tolerance. Request schema validation and cross-record validation still apply. See errors and recovery for failure handling.

update_forecast

Compute and append an evidence-based Bayesian revision atomically.

Contract

The probability is computed, not supplied. Each likelihood conditions on earlier evidence. Reused origins, impossible evidence, stale predecessors and future evidence are rejected; all inputs are stored in the calculation audit.

Fields

FieldTypeRequired
keyKeyYes
op"update_forecast"Yes
revisionBayesianRevisionYes

Setup

Start a fresh --memory process and send these requests, one per line, before the example. The same sequence also works with a new SQLite database.

{"version":1,"id":"setup-1","command":{"op":"create_question","question":{"id":"doc-launch","version":1,"proposition":"Release before timestamp 100","opens_at":0,"deadline":100,"resolve_after":110,"yes_rule":"Archive confirms release strictly before 100","no_rule":"Complete archive confirms no qualifying release","void_rule":"Archive permanently unavailable","resolution_sources":["official archive"]}}}
{"version":1,"id":"setup-2","command":{"op":"append_forecast","key":{"question_id":"doc-launch","version":1},"forecast":{"id":"f1","as_of":1,"evidence_cutoff":1,"probability":0.3,"previous_id":null,"change":"initial","method":"empirical outside view","rationale":"12 of 40 comparable attempts","assumptions":["cases are comparable"],"next_review_trigger":"readiness test","evidence":[]}}}

Request

Send this object on one line. It is expanded below for readability.

{
  "version": 1,
  "id": "example",
  "command": {
    "op": "update_forecast",
    "key": {
      "question_id": "doc-launch",
      "version": 1
    },
    "revision": {
      "id": "f2",
      "previous_id": "f1",
      "as_of": 20,
      "evidence_cutoff": 20,
      "rationale": "diagnostic readiness evidence",
      "assumptions": [
        "illustrative likelihoods"
      ],
      "next_review_trigger": "dependency completion",
      "likelihoods": [
        {
          "evidence": {
            "id": "readiness",
            "origin_id": "test-1",
            "locator": "supplied:test-1",
            "claim": "Readiness test passed",
            "published_at": 10,
            "observed_at": 9
          },
          "given_h": 0.8,
          "given_not_h": 0.2
        }
      ]
    }
  }
}

Response

This response is generated by executing the example against the current binary. Commit timestamps, when present, are shown as <runtime UTC seconds>; the actual protocol returns integer UTC seconds.

{
  "version": 1,
  "id": "example",
  "result": {
    "event": {
      "calculation": {
        "likelihoods": [
          {
            "evidence": {
              "claim": "Readiness test passed",
              "id": "readiness",
              "locator": "supplied:test-1",
              "observed_at": 9,
              "origin_id": "test-1",
              "published_at": 10
            },
            "given_h": 0.8,
            "given_not_h": 0.2
          }
        ],
        "prior": 0.3
      },
      "forecast": {
        "as_of": 20,
        "assumptions": [
          "illustrative likelihoods"
        ],
        "change": "evidence",
        "evidence": [
          {
            "claim": "Readiness test passed",
            "id": "readiness",
            "locator": "supplied:test-1",
            "observed_at": 9,
            "origin_id": "test-1",
            "published_at": 10
          }
        ],
        "evidence_cutoff": 20,
        "id": "f2",
        "method": "conditional Bayesian likelihood update",
        "next_review_trigger": "dependency completion",
        "previous_id": "f1",
        "probability": 0.631578947368421,
        "rationale": "diagnostic readiness evidence"
      },
      "kind": "forecast"
    },
    "key": {
      "question_id": "doc-launch",
      "version": 1
    },
    "recorded_at": "<runtime UTC seconds>",
    "sequence": 3
  }
}

Verification

The documentation gate executes the setup and request, then checks the following result fields against independently specified expectations:

JSON pointer within resultExpected
/sequence3
/event/forecast/probability0.631578947368421
/event/calculation/prior0.3

Float comparisons use a 1e-12 tolerance. Request schema validation and cross-record validation still apply. See errors and recovery for failure handling.

validate_holdout

Validate declared question, event-cluster and time separation.

Contract

Training outcomes must be available by cutoff; holdout forecasts must follow it. Questions are unique across both sets and event clusters cannot cross the boundary. Hidden model-training contamination cannot be detected.

Fields

FieldTypeRequired
cutoffintegerYes
holdoutarray of CaseYes
op"validate_holdout"Yes
trainingarray of CaseYes

This operation needs no existing question or forecast. Use --memory for a calculation-only session.

Request

Send this object on one line. It is expanded below for readability.

{
  "version": 1,
  "id": "example",
  "command": {
    "op": "validate_holdout",
    "training": [
      {
        "question_id": "a",
        "cluster_id": "a",
        "forecast_at": 1,
        "resolved_at": 5,
        "model": 0.5,
        "baseline": 0.5,
        "outcome": true
      },
      {
        "question_id": "b",
        "cluster_id": "b",
        "forecast_at": 1,
        "resolved_at": 5,
        "model": 0.5,
        "baseline": 0.5,
        "outcome": false
      }
    ],
    "holdout": [
      {
        "question_id": "held-a",
        "cluster_id": "held-a",
        "forecast_at": 11,
        "resolved_at": 20,
        "model": 0.5,
        "baseline": 0.5,
        "outcome": true
      },
      {
        "question_id": "held-b",
        "cluster_id": "held-b",
        "forecast_at": 11,
        "resolved_at": 20,
        "model": 0.5,
        "baseline": 0.5,
        "outcome": false
      }
    ],
    "cutoff": 10
  }
}

Response

This response is generated by executing the example against the current binary. Commit timestamps, when present, are shown as <runtime UTC seconds>; the actual protocol returns integer UTC seconds.

{
  "version": 1,
  "id": "example",
  "result": {
    "holdout_questions": 2,
    "training_questions": 2,
    "valid": true
  }
}

Verification

The documentation gate executes the setup and request, then checks the following result fields against independently specified expectations:

JSON pointer within resultExpected
/validtrue
/training_questions2
/holdout_questions2

Float comparisons use a 1e-12 tolerance. Request schema validation and cross-record validation still apply. See errors and recovery for failure handling.

value_of_information

Measure net expected decision value of observing an indicator before acting.

Contract

The indicator’s implied prior defines the baseline. Cost is nonnegative in the same utility units as actions. Negative net value is legitimate; information gain alone need not change an action.

Fields

FieldTypeRequired
actionsarray of ActionYes
costnumberYes
indicatorIndicatorYes
op"value_of_information"Yes

This operation needs no existing question or forecast. Use --memory for a calculation-only session.

Request

Send this object on one line. It is expanded below for readability.

{
  "version": 1,
  "id": "example",
  "command": {
    "op": "value_of_information",
    "actions": [
      {
        "if_yes": 10.0,
        "if_no": -10.0
      },
      {
        "if_yes": 0.0,
        "if_no": 0.0
      }
    ],
    "indicator": {
      "chance": 0.5,
      "if_yes": 1.0,
      "if_no": 0.0
    },
    "cost": 1.0
  }
}

Response

This response is generated by executing the example against the current binary. Commit timestamps, when present, are shown as <runtime UTC seconds>; the actual protocol returns integer UTC seconds.

{
  "version": 1,
  "id": "example",
  "result": {
    "net_value": 4.0
  }
}

Verification

The documentation gate executes the setup and request, then checks the following result fields against independently specified expectations:

JSON pointer within resultExpected
/net_value4.0

Float comparisons use a 1e-12 tolerance. Request schema validation and cross-record validation still apply. See errors and recovery for failure handling.