Exact binary transformation
For probabilities strictly inside (0,1): logit(p)=ln(p/(1-p)); z=sum_i w_i logit(p_i), with nonnegative weights summing to one; pooled probability=sigmoid(alphaz)=1/(1+exp(-alphaz)).
Equal weights w_i=1/n recover the simple aggregator in Satopää et al. (2014). Alpha=1 produces a geometric mean of odds, not an arithmetic mean of probabilities. Alpha>1 moves the pooled probability away from 0.5; 0<alpha<1 shrinks it toward 0.5. Alpha=0 returns 0.5. The implemented helper accepts nonnegative alpha.
Extremizing an arithmetic mean is a different algorithm. Do not conflate sigmoid(alphamean(logit(p_i))) with sigmoid(alphalogit(mean(p_i))).
Why adjustment might help
Pooling forecasters with different information can yield an underconfident aggregate. That is a hypothesis to test, not a license to make any consensus more extreme. Correlated information and overconfident inputs can make extremization harmful.
The method’s original idealized model treats log-odds errors as interchangeable noise around attenuated latent log odds. It is not a universal model of expert or LLM behavior. Do not interpret “alpha=2 worked in a study” as a reusable constant.
Validation procedure
Declare the evaluation unit and cutoff before fitting. Split by question and time, not individual updates randomly. Keep related questions in the same fold when shared outcomes would leak information.
Compare a small prespecified alpha grid using a proper score on training data, then evaluate the selected choice on held-out data. Choosing the grid or winner repeatedly based on the test set consumes the holdout. Report uncertainty in the score difference and domain/sample limitations.
Prefer equal weights without enough relevant performance history. An arbitrary weight based on prestige is not calibrated expertise. Custom weight fitting increases degrees of freedom and requires extra validation. The helper supports declared weights but does not fit them.
Endpoints and missingness
Logit(0) and logit(1) are infinite. The helper rejects endpoints by default. An explicit epsilon option clips to [epsilon,1-epsilon] and reports clipping; this is a numerical/model choice, not newly learned evidence. It can strongly affect conflicting extreme forecasts.
Do not use 0.5 to fill missing responses silently. Exclude them with a coverage report or use a prespecified method that is separately evaluated.
Helper and output
python3 scripts/aggregate.py --probabilities 0.2 0.6 0.8 --alpha 1
Returns arithmetic baseline, logit pool, alpha, weights and clipping metadata. It never fits alpha automatically. Keep model-only estimates labeled as such. Low ensemble spread does not establish calibrated uncertainty.