calibration
Diagnose reliability and resolution while preserving unresolved coverage.
Contract
Use one frozen forecast per unique question ID. Null bins chooses exact groups; a positive bin count chooses equal-width groups. Raw and coarsened Brier are reported separately.
Fields
| Field | Type | Required |
|---|---|---|
bins | integer or null | No |
op | "calibration" | Yes |
records | array of CalibrationRecord | Yes |
This operation needs no existing question or forecast. Use --memory for a calculation-only session.
Request
Send this object on one line. It is expanded below for readability.
{
"version": 1,
"id": "example",
"command": {
"op": "calibration",
"records": [
{
"question_id": "a",
"probability": 0.2,
"outcome": false
},
{
"question_id": "b",
"probability": 0.8,
"outcome": true
},
{
"question_id": "c",
"probability": 0.7,
"outcome": null
}
],
"bins": null
}
}
Response
This response is generated by executing the example against the current binary. Commit timestamps, when present, are shown as <runtime UTC seconds>; the actual protocol returns integer UTC seconds.
{
"version": 1,
"id": "example",
"result": {
"binning_residual": 1.3877787807814457e-17,
"bins": [
{
"count": 1,
"mean_forecast": 0.2,
"observed_frequency": 0.0
},
{
"count": 1,
"mean_forecast": 0.8,
"observed_frequency": 1.0
}
],
"coarsened_brier": 0.03999999999999998,
"excluded_unresolved": 1,
"raw_brier": 0.039999999999999994,
"reliability": 0.039999999999999994,
"resolution": 0.25,
"resolved": 2,
"uncertainty": 0.25
}
}
Verification
The documentation gate executes the setup and request, then checks the following result fields against independently specified expectations:
JSON pointer within result | Expected |
|---|---|
/resolved | 2 |
/excluded_unresolved | 1 |
/raw_brier | 0.04 |
Float comparisons use a 1e-12 tolerance. Request schema validation and cross-record validation still apply. See errors and recovery for failure handling.