Observed benchmark results
These local observations use optimized supercast-agent 0.2.0 on x86_64 Linux. They are reproducible workload results, not stable performance guarantees or prospective agent accuracy. Full machine-readable reports are under target/benchmarks/.
GJP year-one replay
Read 308,232 answer rows, yielding 88 frozen question snapshots: 17 training, 46 holdout, and 25 outside the temporal split. Independent numerical checks passed.
| Method | Holdout Brier (lower is better) | Log loss | Paired Brier difference vs arithmetic, 95% interval |
|---|---|---|---|
| neutral | 0.250000 | 0.693147 | 0.087682 [0.059452, 0.116920] |
| arithmetic | 0.162318 | 0.505864 | 0.000000 [0.000000, 0.000000] |
| logit | 0.144706 | 0.457483 | -0.017612 [-0.027226, -0.007341] |
| recalibrated | 0.114480 | 0.364550 | -0.047837 [-0.080911, -0.011487] |
Training selected {'intercept': -0.5, 'slope': 2.0} from the prespecified grid. Endpoint regularization adjusted 1944 retained forecaster probabilities. The neutral baseline is 0.5 for every question. Binary Brier uses the [0,1] convention, half the two-component binary Brier used in some GJP publications.
These results apply to the selected historical cohort. Base-question clustering does not establish independence between related geopolitical events. Calibration selection used only 17 training questions; the reported intervals condition on that fitted selection. See the protocol and exclusions.
JSON scoring stress
Every requested binary score matched the Python Brier oracle. The ten-million run additionally checked log-loss values and infinite endpoint penalties on every request. Quantiles are from a bounded 10,000-observation reservoir. Timing includes driver overhead.
| Requests | Seconds | Requests/s | p95 round trip (ms) | p99 (ms) | Agent peak RSS (KiB) |
|---|---|---|---|---|---|
| 100,000 | 2.88 | 34,691 | 0.0319 | 0.0371 | 5964 |
| 1,000,000 | 28.51 | 35,080 | 0.0318 | 0.0374 | 6088 |
| 10,000,000 | 306.34 | 32,643 | 0.0334 | 0.0448 | 5952 |
Separate durable journal workloads
| Revisions | p95 write (ms) | First-quarter mean (ms) | Last-quarter mean (ms) | Final database bytes |
|---|---|---|---|---|
| 500 | 1.051 | 0.325 | 0.996 | 835,584 |
| 1,000 | 1.936 | 0.434 | 1.811 | 1,564,672 |
| 2,000 | 3.422 | 0.652 | 3.170 | 3,108,864 |
All three storage workloads passed exact retry, stale-update rejection, acknowledged and unacknowledged-write restart checks, 16 concurrent question creations with four workers, and SQLite integrity checking. Database sizes include the additional interrupted revision and concurrent questions.
Historical v0.2.0 limitation: writes slowed with history length because every mutation replayed the journal. The subsequent v0.2.1 incremental-validation change addresses consecutive warm writes; see the before/after results. These original measurements remain unchanged. Ten million scoring requests do not establish ten-million-write capacity.
Provenance
Binary SHA-256: a6a9d5484ed1cda43b47754a5c3107cc808e13896fe2225d797126b228d96609
Seed: 20260916. Python: 3.14.4. Dataset: doi:10.7910/DVN/BPCDH5.
ifps.csvSHA-256:1f64e5483741e656ddc667b310799e9aec8a8c5dbb7e12f9af5c5a60fd146286survey_fcasts.yr1.tsvSHA-256:a070cc0e87eda8aba63058669308b41022b38f4030e421ee34173528cbb9e2b2