Nest the choices
Hyperparameters and calibration are learned from inner training chronology, never the outer test block.
Compare five development candidates inside the same nested, purged chronology—and let the training prior compete too.
These readings update independently. They are not model inputs, deployment evidence, or a trading recommendation.
TECHNICALS THROUGH · TIMESTAMPED FALLBACK
“Best model” means the lowest out-of-sample probability error among the five candidates—not permission to deploy. The smoothed fold-local prior remains the benchmark every run must beat.
Nested tuning is slower than one global fit. That cost buys a cleaner answer about how the process would have behaved at the time.
Hyperparameters and calibration are learned from inner training chronology, never the outer test block.
Five sessions separate training labels from the next test block so overlapping outcomes cannot cross the boundary.
A one-standard-error rule selects the simplest configuration still supported by inner log loss.
Outer minimum train504 sessionsTwo trading years before the first test
Inner minimum train252 sessionsOne year before tuning begins
Purge5 sessionsMatches the label horizon
Test block63 sessionsFinal block contains 47 sessions
Outer folds12 expandingOne score per chronological test date
TuningPurged inner walk-forwardOne-standard-error selection
CalibrationRegularized Platt mapFit only to inner OOF probabilities
Primary metricBrier improvementCompared with the fold-local prior
Bars show pooled Brier improvement versus the fold-local prior. Zero is the required boundary; every candidate remains negative.
Run 5 is the least weak candidate: AUC 0.4577, Brier 0.2529, and Brier improvement −0.00170. It beats the prior in 6 of 12 folds—but loses overall. This ranking was made after inspecting all five runs; it is not isolated selection evidence or a deployment recommendation.
When the inner evidence is allowed to choose conservatively, it selects the smoothed training prior for 677 of 740 test rows and elastic net once for the remaining 63.
The protocol can pass while the models fail. That separation is the point of an honest evaluation.
Every run starts from the same 1,249 labeled, feature-complete rows.
Each outer model trains on at least 504 earlier sessions.
A five-session purge separates every training window from its test block.
Twelve outer folds create 740 unique OOS predictions from August 2023 through July 2026.
Inner tuning begins only after 252 training sessions are available.
Scaling, hyperparameters, and calibration are fit inside training chronology.
The final outer block is kept at its available 47 sessions rather than discarded.
The smoothed training prior is recomputed independently inside each fold.
Earlier ASTS results were already inspected before these five candidates were compared. All five runs are therefore development evidence and still require a newly accumulated paper-forward period before any deployment consideration.
Phase A above is unchanged. Every Phase B result reuses the same 740 dates, so these are diagnostic development replays, not additional holdouts or independent validation.
It clears the Run 1 comparison and fold-consistency gates, but its probabilities remain tightly compressed and qcut ECE is 6.028%.
It is the closest diagnostic to addressing spread, qcut ECE, and fold consistency together, but it was not selected and misses qcut ECE by 0.5738 percentage points.
It avoids near-constant forecasts and passes qcut ECE, but loses to the prior overall and reaches only half the folds.
Only three folds find an eligible active configuration. Nine folds correctly emit the predeclared no-signal fallback.
| Candidate | Brier gain | AUC | Slope / intercept / qcut ECE | Within-fold flatness | Positive folds | Gate count |
|---|---|---|---|---|---|---|
| Prior-rank Ridge family output | +0.000050 | 0.444743 | -4.856800 / -0.349814 / 0.060283 | Median within-fold std. dev. 0.000569 | 9 / 12 | 2 / 9 / G4 / G8 |
| Empirical-bin family member | +0.000871 | 0.502218 | 0.4540 / 0.01547 / 0.055738 | Median within-fold std. dev. 0.027692 / maximum 0.039467 | 8 / 12 | 2 / 9 / G6 / G8 |
| Non-flat state model | -0.003385 | 0.486449 | -0.853846 / -0.181012 / 0.043606 | Median within-fold std. dev. 0.021151 | 6 / 12 | 1 / 9 / G7 |
| Neutral-anchor Core Ridge | +0.000040 | 0.441582 | -3.876583 / -0.270358 / 0.060468 | Median within-fold std. dev. 0.000000 across all folds / 0.006596 active-only | 2 / 12 | 1 / 9 / G4 |
Neutral-anchor Core Ridge activates in only 3 of 12 folds. The other 9 folds emit the frozen fold-local prior exactly because no inner configuration clears its predeclared eligibility rules. That fallback is a safeguard, not an improvement.
Calibration family run log / Non-flat state run log / Neutral-core run log / Corrected legacy direction-aware run log. The Step 4 workbook preserves Phase A and adds distinct Phase B sheets.
Seven hundred forty chronological dates with fold IDs, baseline probabilities, and five untouched candidate probabilities.
The post-hoc least-bad candidate now faces deployment gates.