Back to the log

Estimating Base Rates with the Outside View: Method‌‌​⁠‌​‌⁠‌​⁠​​⁠​‌​​​‌‌‌⁠‌⁠​‌​⁠‌​​⁠⁠‌‍⁠‌⁠⁠​⁠​⁠‌‌‌‍‍⁠​‍⁠‌⁠‍⁠‌​‍‌‍​‌

Class selection protocol

Define similarity by causal relevance, not surface resemblance. For a delivery deadline, scope, team continuity, dependency count and project maturity may matter more than company branding. Select these criteria before seeing which class yields the desired forecast.

Inventory missing and failed cases. A dataset of announced successes cannot estimate the probability of success among all attempts. A denominator of completed projects cannot estimate on-time delivery when unfinished projects have disappeared from the sample.‌‌​⁠‌​‌⁠‌​⁠​​⁠​‌​​​‌‌‌⁠‌⁠​‌​⁠‌​​⁠⁠‌‍⁠‌⁠⁠​⁠​⁠‌‌‌‍‍⁠​‍⁠‌⁠‍⁠‌​‍‌‍​‌

Align exposure. A one-year failure probability is not a quarterly probability. For a constant hazard assumption, convert p over horizon T into 1-(1-p)^(t/T), but label the constant-hazard assumption. Do not divide a large annual probability by four without checking the approximation.

Estimation

For comparable completed binary trials, use k/n as the empirical rate. For small samples, distinguish that descriptive rate from a smoothed predictive estimate.‌‌​⁠‌​‌⁠‌​⁠​​⁠​‌​​​‌‌‌⁠‌⁠​‌​⁠‌​​⁠⁠‌‍⁠‌⁠⁠​⁠​⁠‌‌‌‍‍⁠​‍⁠‌⁠‍⁠‌​‍‌‍​‌

An optional Beta-Binomial model gives posterior Beta(alpha+k, beta+n-k) and next-trial mean (alpha+k)/(alpha+beta+n). Alpha and beta are modeling choices, not extra observed cases. A uniform Beta(1,1) or Jeffreys Beta(0.5,0.5) prior is not universally correct. The model assumes exchangeability; shared projects, teams or regimes can invalidate simple independent-trial uncertainty calculations.

For cost or time estimates, retain skewness. Use empirical quantiles of realized-to-estimated ratios when appropriate; do not replace a skewed tail with a symmetric percentage band. A P80 budget is a decision quantile, not an 80% confidence interval around the mean.

Transferring the class‌‌​⁠‌​‌⁠‌​⁠​​⁠​‌​​​‌‌‌⁠‌⁠​‌​⁠‌​​⁠⁠‌‍⁠‌⁠⁠​⁠​⁠‌‌‌‍‍⁠​‍⁠‌⁠‍⁠‌​‍‌‍​‌

Separate three uncertainties: sampling error in the dataset; uncertainty over which class applies; structural changes between past and future. Increasing sample size helps the first, not automatically the others.

Use a coarse sensitivity table when class selection is disputed. If plausible classes imply 20%, 35% and 60%, expose that range and its assumptions. Do not average overlapping classes as if they were independent evidence sources.

Adjustment methods include a documented likelihood ratio, a validated predictive model, or clearly labeled judgment. Do not fabricate a measured adjustment coefficient. If credible data are absent, provide an elicited prior and an evidence acquisition plan; never present an invented denominator.‌‌​⁠‌​‌⁠‌​⁠​​⁠​‌​​​‌‌‌⁠‌⁠​‌​⁠‌​​⁠⁠‌‍⁠‌⁠⁠​⁠​⁠‌‌‌‍‍⁠​‍⁠‌⁠‍⁠‌​‍‌‍​‌

Deliverable

Include the class definition, count table, observation dates, empirical rate, any smoothing choice, transfer concerns and case-specific adjustment. Label a sensitivity envelope as such; it is not automatically a confidence or credible interval.