Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Observed benchmark results

These local observations use optimized supercast-agent 0.2.0 on x86_64 Linux. They are reproducible workload results, not stable performance guarantees or prospective agent accuracy. Full machine-readable reports are under target/benchmarks/.

GJP year-one replay

Read 308,232 answer rows, yielding 88 frozen question snapshots: 17 training, 46 holdout, and 25 outside the temporal split. Independent numerical checks passed.

MethodHoldout Brier (lower is better)Log lossPaired Brier difference vs arithmetic, 95% interval
neutral0.2500000.6931470.087682 [0.059452, 0.116920]
arithmetic0.1623180.5058640.000000 [0.000000, 0.000000]
logit0.1447060.457483-0.017612 [-0.027226, -0.007341]
recalibrated0.1144800.364550-0.047837 [-0.080911, -0.011487]

Training selected {'intercept': -0.5, 'slope': 2.0} from the prespecified grid. Endpoint regularization adjusted 1944 retained forecaster probabilities. The neutral baseline is 0.5 for every question. Binary Brier uses the [0,1] convention, half the two-component binary Brier used in some GJP publications.

These results apply to the selected historical cohort. Base-question clustering does not establish independence between related geopolitical events. Calibration selection used only 17 training questions; the reported intervals condition on that fitted selection. See the protocol and exclusions.

JSON scoring stress

Every requested binary score matched the Python Brier oracle. The ten-million run additionally checked log-loss values and infinite endpoint penalties on every request. Quantiles are from a bounded 10,000-observation reservoir. Timing includes driver overhead.

RequestsSecondsRequests/sp95 round trip (ms)p99 (ms)Agent peak RSS (KiB)
100,0002.8834,6910.03190.03715964
1,000,00028.5135,0800.03180.03746088
10,000,000306.3432,6430.03340.04485952

Separate durable journal workloads

Revisionsp95 write (ms)First-quarter mean (ms)Last-quarter mean (ms)Final database bytes
5001.0510.3250.996835,584
1,0001.9360.4341.8111,564,672
2,0003.4220.6523.1703,108,864

All three storage workloads passed exact retry, stale-update rejection, acknowledged and unacknowledged-write restart checks, 16 concurrent question creations with four workers, and SQLite integrity checking. Database sizes include the additional interrupted revision and concurrent questions.

Historical v0.2.0 limitation: writes slowed with history length because every mutation replayed the journal. The subsequent v0.2.1 incremental-validation change addresses consecutive warm writes; see the before/after results. These original measurements remain unchanged. Ten million scoring requests do not establish ten-million-write capacity.

Provenance

Binary SHA-256: a6a9d5484ed1cda43b47754a5c3107cc808e13896fe2225d797126b228d96609

Seed: 20260916. Python: 3.14.4. Dataset: doi:10.7910/DVN/BPCDH5.

  • ifps.csv SHA-256: 1f64e5483741e656ddc667b310799e9aec8a8c5dbb7e12f9af5c5a60fd146286
  • survey_fcasts.yr1.tsv SHA-256: a070cc0e87eda8aba63058669308b41022b38f4030e421ee34173528cbb9e2b2