Dataset replay and stress testing
The standard-library Python harness executes the compiled Supercast JSON adapter. It measures both numerical agreement with independent Python formulas and process/storage behavior. It does not call a language model or generate new factual forecasts.
Reproduce
Build the optimized binary, fetch pinned public inputs, then run the two workloads:
CARGO_TARGET_DIR=target cargo build -p supercast-agent --release --locked --offline
python3 scripts/fetch_gjp.py
python3 scripts/benchmark.py replay --output target/benchmarks/gjp-year1.json
python3 scripts/benchmark.py stress --events 100000 --revisions 500 --output target/benchmarks/stress-100k.json
python3 scripts/benchmark.py stress --events 1000000 --revisions 1000 --output target/benchmarks/stress-1m.json
python3 scripts/benchmark.py stress --events 10000000 --revisions 2000 --output target/benchmarks/stress-10m.json
Run timing workloads sequentially. Reports include the binary hash/version, seed, elapsed time and workload configuration. Fetching needs internet and curl; replay and stress are offline. Raw inputs and generated reports stay in ignored target/ and are not included in release source snapshots. The standard release gate includes three importer/oracle regression tests and a small synthetic integration run; downloading real data is not a prerequisite for that gate.
Source and provenance
Source: Good Judgment Project, GJP Data, Harvard Dataverse, doi:10.7910/DVN/BPCDH5. The dataset metadata identifies CC0 1.0. IARPA’s release announcement describes the four-year public archive.
The downloader pins three Dataverse file IDs: 2917330 (question metadata), 2917351 (year-one survey forecasts), and 2917350 (field documentation). It verifies SHA-256 fingerprints of the served representations and writes a manifest with download URLs and byte counts. A changed or corrupt existing file fails verification instead of being overwritten. Dataverse’s ingested TSV representation differs from the original CSV, so its original-file MD5 is not the checksum of the downloaded TSV.
Only question IDs, option counts, question type/status, dates and outcome labels are consumed from ifps.csv. That file is decoded with byte-preserving Latin-1 because it contains mixed non-UTF-8 prose; its prose is not rewritten or imported into a factual journal. Forecast TSV fields are UTF-8. No demographic or individual-difference data is downloaded.
Frozen replay protocol
- Retain ordinary (
q_type=0), closed, two-option questions with resolved optionaorb. Conditional, categorical, void and unresolved questions are excluded. - Freeze each question seven days after its recorded opening. Exclude questions already suspended or resolved at that instant. This is a fixed age from opening, not a fixed distance from resolution.
- Keep only option-
arows between opening and snapshot, inclusive. For each question/forecaster use the latest(timestamp, forecast_id); forecast ID deterministically breaks timestamp ties. A latest withdrawal removes that forecaster from the panel. Complementary option-brows are not additional observations. - Interpret outcome
aas true. This is the event “option a occurs,” which need not be the literal word “yes.” Apply declared endpoint regularizationepsilon=0.001to both pooling methods. Count adjustments. Every retained forecaster has equal weight regardless of update frequency or experimental condition. - Training questions must resolve strictly before 2012-01-01. Holdout snapshots must occur on or after that cutoff and have no training cluster overlap. Other snapshots are excluded. Cluster IDs use the base IFP identifier before its suffix.
- Select among a prespecified grid of intercepts
[0,-0.5,0.5]and slopes[1,0.5,1.5,2], minimizing training Brier only. Ties choose the first candidate. Evaluate the frozen selection once on holdout questions. - Report equal-question Brier, natural-log loss, ten-bin calibration diagnostics and paired Brier differences against arithmetic pooling. Include a neutral 0.5 baseline. Cluster percentile intervals use 2,000 resamples, 95% confidence and seed
20260916.
Independent Python formulas verify arithmetic/logit pooling, per-question Brier/log loss, candidate selection, calibrated holdout score and calibration’s raw Brier. The bootstrap implementation is exercised, but its full sampling distribution is not independently reimplemented here. Reports retain the actual training and holdout cases for inspection.
Archive timestamps lack an explicit timezone in these files. The importer preserves their nominal ordering by interpreting all dates on one UTC-labelled clock; it does not claim to reconstruct historically accurate UTC instants. Forecasts recorded before a question’s metadata opening date are excluded conservatively. The exclusion counts make this choice visible.
Related geopolitical questions can remain dependent even with distinct base IDs. The confidence intervals assume those clusters are independent, and condition on the selected calibration model; they do not include training-selection uncertainty. Seventeen training questions are a small sample. This is a transparent regression/aggregation benchmark, not a reproduction of the tournament’s full daily scoring protocol, a randomized treatment comparison, or evidence that an LLM can predict these events prospectively.
Synthetic stress protocol
--events counts binary-scoring JSON requests; --revisions separately controls the length of one durable journal. Ten million numerical requests do not imply ten million SQLite writes. Requests are generated incrementally with a fixed seed. Normal draws are mixed with exact endpoints and near-endpoint values. Python verifies every Brier result and checks finite or infinite log-loss semantics.
Measurements include end-to-end single-client throughput and round-trip p50/p95/p99. Quantiles use a seeded uniform reservoir capped at 10,000 timings; they are estimates, not exact quantiles over all requests. Throughput includes Python serialization, validation, IPC and sampling overhead. Linux agent RSS high-water marks are sampled every 1,000 numerical requests; they do not measure the Python driver or subsequent storage workers.
The separate storage workload appends an increasing revision chain, repeats exact requests, rejects a stale update, kills/reopens the process, and verifies original receipts and journal length. It also kills a process after sending a request but before reading acknowledgement, retries that request and verifies exactly one additional revision. This permits either pre-commit or post-commit interruption; it does not guarantee a kill occurred inside SQLite’s commit or simulate power loss. Four workers create 16 additional questions in one database, followed by SQLite integrity checking.
Storage reports show write latency, first/last-quarter means and final database size. The v0.2.0 baseline replays history on each write. The v0.2.1 adapter reuses validated state for consecutive writes to one journal, while cold or invalidated caches still replay. See incremental-validation measurements. Millions of writes on one journal would be a separate, substantially more expensive capacity experiment. Concurrent creation is a contention smoke test, not a saturation benchmark. Timing results are single-host observations without performance pass/fail thresholds.
See the checked-in observed results for the completed local run.