compare
Compare model and baseline Brier on paired, uniquely identified questions.
Contract
Forecasts must precede resolution. The difference is model minus baseline, so negative favors the model. This is a descriptive comparison without an interval.
Fields
| Field | Type | Required |
|---|---|---|
cases | array of Case | Yes |
op | "compare" | Yes |
This operation needs no existing question or forecast. Use --memory for a calculation-only session.
Request
Send this object on one line. It is expanded below for readability.
{
"version": 1,
"id": "example",
"command": {
"op": "compare",
"cases": [
{
"question_id": "a",
"cluster_id": "a",
"forecast_at": 1,
"resolved_at": 5,
"model": 0.5,
"baseline": 0.5,
"outcome": true
},
{
"question_id": "b",
"cluster_id": "b",
"forecast_at": 1,
"resolved_at": 5,
"model": 0.5,
"baseline": 0.5,
"outcome": false
}
]
}
}
Response
This response is generated by executing the example against the current binary. Commit timestamps, when present, are shown as <runtime UTC seconds>; the actual protocol returns integer UTC seconds.
{
"version": 1,
"id": "example",
"result": {
"baseline_brier": 0.25,
"difference": 0.0,
"model_brier": 0.25,
"questions": 2
}
}
Verification
The documentation gate executes the setup and request, then checks the following result fields against independently specified expectations:
JSON pointer within result | Expected |
|---|---|
/questions | 2 |
/model_brier | 0.25 |
/difference | 0.0 |
Float comparisons use a 1e-12 tolerance. Request schema validation and cross-record validation still apply. See errors and recovery for failure handling.