# Validation Report‌‌​⁠‌​‌⁠‌​⁠​​⁠​‌​​​‌⁠‍‍​⁠‌⁠⁠⁠‌​​​‍‌⁠‌‌​‌⁠‌‌⁠​⁠⁠‍⁠⁠‍‍‍‌‍⁠⁠‍​⁠‍⁠​​

Pack version: 1.0. Date: 2026-09-15.

## Executed checks

| Check | Result | Scope |
| --- | --- | --- |
| Frontmatter/native skill validator | 13/13 passed | Naming, required metadata, scaffold checks |
| Pack structural validator | 13/13 passed | Format, local links, line counts, Python syntax |
| Numerical test suite | 45 tests passed | Five helpers, invalid inputs, endpoints, known values, seeded invariants and evaluation-set structure |
| CLI help paths | All five passed | Program invocation and argument parsing |
| CLI calculations | All five helpers exercised | Numerical arguments and JSON-file input paths |
| Fresh-agent forward tasks | Five tasks completed with expected safeguards | Bayes, aggregation, calibration and conditional trees |
| Baseline without skills | Three tasks completed successfully | No measured improvement claim |
| Claude runtime trials | Not run | No Claude session was available for this evaluation |
| Prospective forecasting outcomes | Not run | Mathematical correctness does not imply forecasting skill |

## Forward-task observations‌‌​⁠‌​‌⁠‌​⁠​​⁠​‌​​​‌⁠‍‍​⁠‌⁠⁠⁠‌​​​‍‌⁠‌‌​‌⁠‌‌⁠​⁠⁠‍⁠⁠‍‍‍‌‍⁠⁠‍​⁠‍⁠​​

1. Duplicate evidence: a 30% prior and likelihoods 0.8/0.2 yielded 63.1579%, counting the shared report once.
2. Correlated model estimates: 0.7/0.75/0.8 produced an arithmetic baseline of 75% and alpha=1 logit pool of about 75.23%. Alpha=2 was labeled hypothetical, not validated.
3. Binning: raw Brier 0.375 differed from coarsened score 0.3725 by 0.0025; the agent preserved that distinction.‌‌​⁠‌​‌⁠‌​⁠​​⁠​‌​​​‌⁠‍‍​⁠‌⁠⁠⁠‌​​​‍‌⁠‌‌​‌⁠‌‌⁠​⁠⁠‍⁠⁠‍‍‍‌‍⁠⁠‍​⁠‍⁠​​
4. Incoherent conditionals: a stated prior of 0.7 versus implied 0.5 was rejected for information-gain computation.
5. Coherent conditional indicator: the calculation yielded 0.1466098512 nats and approximately 24.0004% of maximum information.

A fresh execution exposed an incorrect approximate number in one teaching example. The example was corrected and a numeric regression assertion added. The helper's underlying equation was correct.‌‌​⁠‌​‌⁠‌​⁠​​⁠​‌​​​‌⁠‍‍​⁠‌⁠⁠⁠‌​​​‍‌⁠‌‌​‌⁠‌‌⁠​⁠⁠‍⁠⁠‍‍‍‌‍⁠⁠‍​⁠‍⁠​​

These were fresh Codex-agent executions, not Claude trials. They support a limited instruction-following check, not broad validation. The no-skill baseline also passed its three overlapping conceptual tasks, so this report does not claim a demonstrated performance uplift.

## Change log since first assembly

Calibration helper: unresolved records (outcome null) are now excluded and reported through count_excluded and excluded_unresolved, matching the Brier helper, instead of failing the whole input. Scoring an unresolved question as a non-event remains refused, and a set with no resolved outcome is still rejected. The behavior is documented in the skill's workflow, quality gate and method reference, and four regression tests cover it, including one asserting the two helpers agree on the same mixed file.‌‌​⁠‌​‌⁠‌​⁠​​⁠​‌​​​‌⁠‍‍​⁠‌⁠⁠⁠‌​​​‍‌⁠‌‌​‌⁠‌‌⁠​⁠⁠‍⁠⁠‍‍‍‌‍⁠⁠‍​⁠‍⁠​​

Evaluation set: per-case acceptance rubrics replaced the single shared acceptance sentence.

## Deterministic test coverage

Bayes: exact update, nondiagnostic signal, decisive likelihoods, impossible evidence, invalid inputs, tiny likelihoods, complement symmetry and monotonicity.‌‌​⁠‌​‌⁠‌​⁠​​⁠​‌​​​‌⁠‍‍​⁠‌⁠⁠⁠‌​​​‍‌⁠‌‌​‌⁠‌‌⁠​⁠⁠‍⁠⁠‍‍‍‌‍⁠⁠‍​⁠‍⁠​​

Aggregation: repeated identical estimates, complement pairs, known logit-pool formula, weights including large values, endpoint policy, invalid parameters, alpha=0, range/complement/permutation invariants.

Scoring: binary/categorical normalization, unresolved exclusions, duplicate IDs, malformed probabilities/outcomes and the proper-score expected-loss identity.

Calibration: exact grouping, broad-bin residuals, constant forecasts, random decomposition identities, empty/invalid input, endpoint bins, unresolved exclusion, agreement with the Brier helper on a mixed file, all-unresolved rejection and a missing outcome key.‌‌​⁠‌​‌⁠‌​⁠​​⁠​‌​​​‌⁠‍‍​⁠‌⁠⁠⁠‌​​​‍‌⁠‌‌​‌⁠‌‌⁠​⁠⁠‍⁠⁠‍‍‍‌‍⁠⁠‍​⁠‍⁠​​

Conditional indicators: coherent and incoherent inputs, zero/perfect information, degenerate priors, known values and randomized bounds.

Evaluation set: case count, unique ids, known skill references, non-empty rubric fields, distinct per-case acceptance text and recomputation of every asserted number.

## Limits of evaluation‌‌​⁠‌​‌⁠‌​⁠​​⁠​‌​​​‌⁠‍‍​⁠‌⁠⁠⁠‌​​​‍‌⁠‌‌​‌⁠‌‌⁠​⁠⁠‍⁠⁠‍‍‍‌‍⁠⁠‍​⁠‍⁠​​

The regression tests verify arithmetic and input handling. They do not verify research validity, source truth, event adjudication, statistical independence, forecasting accuracy, security of external integrations or prompt activation.

The 39 cases in behavioral-cases.json are a supplied evaluation design. They have not all been executed. Run them with each intended Claude model, including no-skill baselines, and preserve failures rather than retroactively changing the acceptance criteria.

Each case now carries its own rubric rather than a shared sentence: a one-line acceptance summary, an explicit `must` list, a `must_not` list of the failure modes that case exists to catch, and, for the seventeen cases with determinate arithmetic, an `expect` block of values and tolerances. A case passes only when every `must` holds, no `must_not` is triggered, and the shared baseline in `grading.shared_baseline` is met. The 36 numbers asserted across those `expect` blocks are recomputed from the helpers by the test suite, so a rubric cannot drift away from what the code produces.‌‌​⁠‌​‌⁠‌​⁠​​⁠​‌​​​‌⁠‍‍​⁠‌⁠⁠⁠‌​​​‍‌⁠‌‌​‌⁠‌‌⁠​⁠⁠‍⁠⁠‍‍‍‌‍⁠⁠‍​⁠‍⁠​​

## Reproducibility

All numerical tests use standard-library Python and fixed random seeds. Helpers make no network calls and write no output files. The validation suite may cause Python bytecode caching unless PYTHONDONTWRITEBYTECODE=1 is set.

The archive builder excludes caches and product-specific metadata. SHA256SUMS.json inventories packaged content so extraction integrity can be checked. ZIP integrity and inner single-root layouts are checked during packaging.
