Performance and capacity
The core favors straightforward algorithms with explicit allocation behavior. Keep arithmetic latency, record-validation cost, database synchronization and model inference separate when measuring an agent.
| Operation | Time | Extra storage |
|---|---|---|
| Scalar Bayesian update or Brier | O(1) | O(1) |
| Weighted pool | O(n) | O(1) |
| Streaming Brier | O(1) per update | Count and mean |
| Conditional chain / mixture | O(n) | O(1) |
| Empirical quantile / CRPS | O(n log n) sort | In-place sample slice |
| Calibration groups | O(n log g) | IDs and nonempty groups |
| Time-grid score | O(updates + grid) | O(1) |
| Cluster bootstrap | Sorting plus O(replicates × clusters) | Cluster summaries and replicates |
| Calibration grid selection | O(candidates × training + holdout) | Candidate scores |
| Journal read / cold mutation validation | O(history + evidence) | Replayed history and IDs |
| Warm mutation validation | Amortized O(new event + new evidence), excluding database work | One retained history and ID/origin sets |
Some transport commands copy caller-owned arrays before invoking an in-place core routine. That cost belongs to the JSON adapter, not to the scalar numerical function. The adapter parses a whole request line; it does not stream individual JSON array entries.
Benchmark locally
CARGO_TARGET_DIR=target cargo bench --bench core --locked --offline
The harness warms each case for 1,000 calls and uses optimizer barriers around inputs and results. It measures a scalar Bayesian update, a pool of 100 equal-weight interior probabilities, and a running Brier update with mean access.
An early local run on an AMD Ryzen Threadripper PRO 5975WX with Rust 1.97.1 measured approximately 31 ns per Bayesian update, 1.5 microseconds per 100-forecast pool, and 4.6 ns per streaming-score operation. These are one-run observations, not stable latency guarantees or comparisons against another implementation.
Workload shape, compiler flags, thermal state, contention and input distribution affect measurements. Record those conditions for comparative claims. Do not infer forecasting accuracy from execution speed.
Capacity limits
The CLI limits a request line to 1 MiB. Listing is capped at 1,000 keys per page and database evaluation at 10,000 keys per request. Bootstrap and calibration selection have explicit work caps. The database journal can grow over time; the request-size limit does not impose a total history limit.
For long-running deployments, monitor journal length, mutation latency and cache hit rate. Repeated writes to the same journal on one Store use incremental validation. Restart, question switches and external commits require full replay. Memory retains one history per Store; it grows with that history. Concurrent or interleaved workloads can see many cache misses. See measured storage results.
Durable writes pay a filesystem synchronization cost. Use one connection per worker and expect write serialization under SQLite. Increasing process count does not make one local database accept unlimited simultaneous writers.
Measure the end-to-end agent separately, including retrieval, inference, evidence processing and persistence. Those application costs are likely to dominate the numerical primitives, but their actual contribution depends on the deployment.