Block uncertainty
Resample 20-session blocks 2,000 times so overlapping outcomes are not treated as independent rows.
Run 5 is the post-hoc least-bad development candidate. Nine predeclared gates decide whether that ranking is strong enough to consider deployment.
These readings update independently. They are not model inputs, deployment evidence, or a trading recommendation.
TECHNICALS THROUGH · TIMESTAMPED FALLBACK
The answer depends on evidence, not rank. Run 5 can be least bad among the five inspected candidates and still fail against the prior, uncertainty, calibration, stability, and null benchmarks.
The family-wise null circularly shifts the frozen OOS label path against frozen forecasts. It excludes shifts within 63 sessions of zero and protects against selecting the strongest of five runs.
Resample 20-session blocks 2,000 times so overlapping outcomes are not treated as independent rows.
Circularly shift the frozen OOS label path and compare the strongest of five development runs.
Audit slope, intercept, five-bin calibration error, and confidence coverage.
Change block length, inspect regimes, remove extreme-event windows, and measure feature drift.
The primary 20-session bootstrap keeps zero inside the Brier improvement interval and 0.50 inside the AUC interval.
Run 5 · shallow boosting
Worse than the fold-local prior
Lower bound remains below 0.50
Frozen-forecast circular-shift null
Above the 5% evidence gate
Run 5 remains the post-hoc least-bad candidate, but that label only ranks the five inspected runs. Eight of nine deployment evidence gates fail, so no model should move forward.
These are conservative research guardrails, not nine independent hypothesis tests. The observed and required values stay visible.
| Gate | Rule | Observed | Required | Result |
|---|---|---|---|---|
| G1 | Positive Brier improvement with 20-session lower bound above zero | −0.00170; lower −0.00498 | Point > 0; lower > 0 | FAIL |
| G2 | ROC AUC and its 20-session lower bound clear no-skill | 0.4577; lower 0.3800 | AUC ≥ 0.52; lower > 0.50 | FAIL |
| G3 | Family-wise frozen-forecast circular-shift tail rate | 96.59% | < 5% | FAIL |
| G4 | Run 5 improves Brier score on Run 1 with paired lower bound | +0.00197; lower −0.00222 | Point > 0; lower > 0 | FAIL |
| G5 | Calibration slope | −1.319 | 0.70 to 1.30 | FAIL |
| G6 | Absolute calibration intercept | 0.098 | ≤ 0.10 | PASS |
| G7 | Five-bin expected calibration error | 6.44% | < 5% | FAIL |
| G8 | Outer folds with positive Brier improvement | 6 / 12 · 50% | ≥ 60% | FAIL |
| G9 | Block-length, event-window, and supported-regime sensitivity | Lower −0.00503; event −0.00242 | All sensitivity conditions pass | FAIL |
The repair program tests flatness, equal-frequency qcut calibration error, and fold consistency separately. Each row keeps its own nine-gate denominator.
| Candidate | Probability spread | qcut ECE | Positive folds | Individual result |
|---|---|---|---|---|
| Prior-rank Ridge family output | 0.4642 to 0.4976 Std. dev. 0.0108 Median within-fold std. dev. 0.000569 | 0.060283 | 9 / 12 | 2 / 9 / G4 / G8 |
| Empirical-bin family member | 0.4257 to 0.5433 Std. dev. 0.030792 Median within-fold std. dev. 0.027692 / maximum 0.039467 | 0.055738 | 8 / 12 | 2 / 9 / G6 / G8 |
| Non-flat state model | 0.3225 to 0.5410 Std. dev. 0.0370 Median within-fold std. dev. 0.021151 | 0.043606 | 6 / 12 | 1 / 9 / G7 |
| Neutral-anchor Core Ridge | Prior fallback in 9 / 12 folds Std. dev. 0.0120 Median within-fold std. dev. 0.000000 across all folds / 0.006596 active-only | 0.060468 | 2 / 12 | 1 / 9 / G4 |
| Legacy direction-aware Core Ridge | Compressed legacy output | 0.066126 | 4 / 12 | 1 / 9 / G4 |
The selected calibration family passes G4 and G8. The non-flat state model passes G7. The unselected empirical-bin member passes G6 and G8. Neutral-anchor and the corrected legacy direction-aware audit pass G4 only. These passes belong to different probability paths and cannot be combined.
The selected prior-rank Ridge family reaches 9 of 12 positive folds, yet its median within-fold forecast standard deviation is only 0.000569 and qcut ECE is 6.028%. The non-flat model spans 32.25% to 54.10%, has a 0.021151 median within-fold standard deviation, and passes qcut ECE at 4.361%, but wins only 6 of 12 folds and loses to the prior overall.
The empirical-bin family member was not selected. It passes G6 and G8, posts +0.000871 Brier improvement and 0.502218 AUC, spans 42.57% to 54.33%, and reaches 8 of 12 positive folds. Its median within-fold standard deviation is 0.027692, but qcut ECE is 5.5738%, missing G7 by 0.5738 percentage points. It remains development-only and does not raise the best gate count above 2 of 9.
This corrected audit uses the original frozen reference, qcut ECE, seed 20260730, and 5/20/63-session blocks. The superseded legacy gate file must not be used. Its only pass is G4: +0.003139 / lower +0.000281. qcut ECE is 0.066126, and the family-wise tail rate is 0.970779.
Calibration family gates / Non-flat state gates / Neutral-core gates / Corrected legacy direction-aware gates. The Phase A table above remains unchanged.
This is a direction-probability exercise, not a trading backtest. It does not model positions, overlapping trades, execution, transaction costs, or portfolio risk.
One frozen audit of uncertainty, family-wise null evidence, calibration, folds, regimes, drift, and event sensitivity.
Turn the evidence into a durable recommendation.