Aggregate and recalibrate
Arithmetic pooling averages probabilities. Logit pooling averages log odds and transforms the result back. These are different algorithms and can give noticeably different answers.
For interior probabilities, logit(p) = ln(p) - ln(1-p). With nonnegative normalized weights w and declared adjustment alpha:
z = sum(w_i * logit(p_i))
aggregate = sigmoid(alpha * z)
Alpha one leaves the pooled log odds unchanged; values above one extremize, values between zero and one shrink toward 0.5, and zero returns 0.5. This is not extremization of the arithmetic mean.
extern crate supercast;
use supercast::{Probability as P, aggregate::{pool, WeightedForecast}};
let forecasts = [
WeightedForecast { probability: P::new(0.2)?, weight: 1.0 },
WeightedForecast { probability: P::new(0.8)?, weight: 1.0 },
];
let result = pool(&forecasts, 1.0)?;
assert!((result.logit.get() - 0.5).abs() < 1e-12);
assert_eq!(result.arithmetic, P::HALF);
Ok::<(), supercast::Error>(())
Numerical and participation policy
pool reports arithmetic and logit results, alpha, and positive-weight contributor count. It rescales weight magnitudes before summing to avoid overflow. Zero-weight rows do not contribute. All weights must be finite and nonnegative, with at least one positive weight.
Logit requires probabilities strictly inside zero and one. Positive-weight endpoints are rejected; there is no implicit clipping. If an application adopts an epsilon policy, it must state and validate that policy separately. Missing panel responses should be excluded with a coverage report, not invented as neutral forecasts.
When adjustment is justified
Extremization can help when participants bring sufficiently distinct information and the initial aggregate is underconfident. Correlated or overconfident inputs can make it harmful. A coefficient that helped another study or domain is not a universal constant.
aggregate::recalibrate applies sigmoid(intercept + slope*logit(p)) with finite intercept and nonnegative slope. It does not fit the parameters. evaluation::select_calibration selects from a declared candidate grid using training outcomes, then reports performance on a separate temporal/cluster holdout.
Include the identity candidate (0,1) when leaving forecasts unchanged is a legitimate option. Once the holdout is inspected, do not keep changing the grid against the same outcomes and call the final result held out. See validation for the evaluation contract.