evaluate
Evaluate a declared cohort at common forecast and resolution cutoffs.
Contract
The latest available forecast is selected per question; all exclusions are reported. One version per question is allowed. This uses caller-declared historical timestamps and is not proof of prospective performance.
Fields
| Field | Type | Required |
|---|---|---|
baseline | Probability | Yes |
baseline_as_of | integer | Yes |
forecast_cutoff | integer | Yes |
keys | array of Key | Yes |
op | "evaluate" | Yes |
resolution_cutoff | integer | Yes |
Setup
Start a fresh --memory process and send these requests, one per line, before the example. The same sequence also works with a new SQLite database.
{"version":1,"id":"setup-1","command":{"op":"create_question","question":{"id":"doc-launch","version":1,"proposition":"Release before timestamp 100","opens_at":0,"deadline":100,"resolve_after":110,"yes_rule":"Archive confirms release strictly before 100","no_rule":"Complete archive confirms no qualifying release","void_rule":"Archive permanently unavailable","resolution_sources":["official archive"]}}}
{"version":1,"id":"setup-2","command":{"op":"append_forecast","key":{"question_id":"doc-launch","version":1},"forecast":{"id":"f1","as_of":1,"evidence_cutoff":1,"probability":0.3,"previous_id":null,"change":"initial","method":"empirical outside view","rationale":"12 of 40 comparable attempts","assumptions":["cases are comparable"],"next_review_trigger":"readiness test","evidence":[]}}}
{"version":1,"id":"setup-3","command":{"op":"update_forecast","key":{"question_id":"doc-launch","version":1},"revision":{"id":"f2","previous_id":"f1","as_of":20,"evidence_cutoff":20,"rationale":"diagnostic readiness evidence","assumptions":["illustrative likelihoods"],"next_review_trigger":"dependency completion","likelihoods":[{"evidence":{"id":"readiness","origin_id":"test-1","locator":"supplied:test-1","claim":"Readiness test passed","published_at":10,"observed_at":9},"given_h":0.8,"given_not_h":0.2}]}}}
{"version":1,"id":"setup-4","command":{"op":"resolve_question","key":{"question_id":"doc-launch","version":1},"resolution":{"at":120,"source":"official archive","rationale":"synthetic archive confirms release","outcome":{"status":"yes"}}}}
Request
Send this object on one line. It is expanded below for readability.
{
"version": 1,
"id": "example",
"command": {
"op": "evaluate",
"keys": [
{
"question_id": "doc-launch",
"version": 1
}
],
"forecast_cutoff": 30,
"resolution_cutoff": 120,
"baseline": 0.3,
"baseline_as_of": 0
}
}
Response
This response is generated by executing the example against the current binary. Commit timestamps, when present, are shown as <runtime UTC seconds>; the actual protocol returns integer UTC seconds.
{
"version": 1,
"id": "example",
"result": {
"baseline": 0.3,
"baseline_as_of": 0,
"eligible": 1,
"evaluation_kind": "declared_timestamp_snapshot",
"excluded": 0,
"exclusions": [],
"forecast_cutoff": 30,
"resolution_cutoff": 120,
"rows": [
{
"as_of": 20,
"baseline_brier": 0.48999999999999994,
"binary_brier": 0.1357340720221607,
"forecast_id": "f2",
"key": {
"question_id": "doc-launch",
"version": 1
},
"outcome": true,
"probability": 0.631578947368421
}
],
"scored": 1,
"summary": {
"baseline_brier": 0.48999999999999994,
"brier_skill": 0.7229916897506924,
"calibration": {
"binning_residual": 0.0,
"bins": [
{
"count": 1,
"mean_forecast": 0.631578947368421,
"observed_frequency": 1.0
}
],
"coarsened_brier": 0.1357340720221607,
"excluded_unresolved": 0,
"raw_brier": 0.1357340720221607,
"reliability": 0.1357340720221607,
"resolution": 0.0,
"resolved": 1,
"uncertainty": 0.0
},
"difference": -0.35426592797783923,
"model_brier": 0.1357340720221607
}
}
}
Verification
The documentation gate executes the setup and request, then checks the following result fields against independently specified expectations:
JSON pointer within result | Expected |
|---|---|
/scored | 1 |
/excluded | 0 |
/rows/0/forecast_id | "f2" |
/summary/model_brier | 0.13573407202216065 |
Float comparisons use a 1e-12 tolerance. Request schema validation and cross-record validation still apply. See errors and recovery for failure handling.