select_calibration
Select a calibration grid candidate on training outcomes, then score the untouched holdout.
Contract
Temporal and cluster separation are checked before fitting. Exact ties choose the first candidate. The selected candidate is not changed after seeing its holdout performance; repeated tuning on that holdout would consume it.
Fields
| Field | Type | Required |
|---|---|---|
candidates | array of CalibrationCandidate | Yes |
cutoff | integer | Yes |
holdout | array of Case | Yes |
op | "select_calibration" | Yes |
training | array of Case | Yes |
This operation needs no existing question or forecast. Use --memory for a calculation-only session.
Request
Send this object on one line. It is expanded below for readability.
{
"version": 1,
"id": "example",
"command": {
"op": "select_calibration",
"training": [
{
"question_id": "a",
"cluster_id": "a",
"forecast_at": 1,
"resolved_at": 5,
"model": 0.5,
"baseline": 0.5,
"outcome": true
},
{
"question_id": "b",
"cluster_id": "b",
"forecast_at": 1,
"resolved_at": 5,
"model": 0.5,
"baseline": 0.5,
"outcome": false
}
],
"holdout": [
{
"question_id": "held-a",
"cluster_id": "held-a",
"forecast_at": 11,
"resolved_at": 20,
"model": 0.5,
"baseline": 0.5,
"outcome": true
},
{
"question_id": "held-b",
"cluster_id": "held-b",
"forecast_at": 11,
"resolved_at": 20,
"model": 0.5,
"baseline": 0.5,
"outcome": false
}
],
"cutoff": 10,
"candidates": [
{
"intercept": 0.0,
"slope": 1.0
},
{
"intercept": 0.0,
"slope": 0.0
}
]
}
}
Response
This response is generated by executing the example against the current binary. Commit timestamps, when present, are shown as <runtime UTC seconds>; the actual protocol returns integer UTC seconds.
{
"version": 1,
"id": "example",
"result": {
"baseline_holdout_brier": 0.25,
"calibrated_holdout_brier": 0.25,
"holdout_questions": 2,
"raw_holdout_brier": 0.25,
"selected": {
"intercept": 0.0,
"slope": 1.0
},
"training_brier": [
0.25,
0.25
],
"training_questions": 2
}
}
Verification
The documentation gate executes the setup and request, then checks the following result fields against independently specified expectations:
JSON pointer within result | Expected |
|---|---|
/selected/slope | 1.0 |
/training_questions | 2 |
/calibrated_holdout_brier | 0.25 |
Float comparisons use a 1e-12 tolerance. Request schema validation and cross-record validation still apply. See errors and recovery for failure handling.