Best inspected candidate Brier score, 0.2529. Pooled Brier improvement is still negative at −0.00170 and classification accuracy is 47.70%.
Summarize
the model.
Turn five research steps and a focused repair audit into one decision: preserve the process, suppress the weak signal, and protect the next test from hindsight.
These readings update independently. They are not model inputs, deployment evidence, or a trading recommendation.
- Latest price
- $62.31 · fallback
- 20D momentum
- -13.4% Close-to-close · 20 completed sessions
- RSI 14
- 47.2 neutral · Wilder method
- Relative volume
- 1.0× Versus the prior 20 completed sessions
TECHNICALS THROUGH · TIMESTAMPED FALLBACK
Improve the fallback. Do not manufacture an edge.
The research chain is auditable, but none of the five candidates or eight repair variants improves pooled probability error versus the fold-local prior. The safer controller therefore emits the prior and withholds directional tilt.
The fold-local prior scores 50.95% accuracy and the lower Brier score. It is a neutral fallback—not proof of predictability.
More complexity did not improve the forecast.
The audit added 24 trailing features and tested recency weighting, rolling history, signed calibration, regularized linear models, and shallow nonlinear learners under the same purged chronology.
Run 5, using a fixed 0.50 decision threshold.
Fold-local smoothed prior; +3.24 percentage points versus Run 5.
Still worse than its prior; 95% block interval −0.00135 to +0.00013.
Use the training prior unless a future frozen challenger clears all nine gates.
The controller is safer and historically more accurate because it refuses the weak tilt. It does not turn the training prior into a predictive stock signal, and it does not improve the one-of-nine deployment score.
A sound process can end with no.
Data and protocol checks answer whether the experiment is auditable. Deployment gates answer a different question: whether the observed model evidence is strong enough.
- 01Target contractPASS
Raw ASTS close, five-session horizon, ties in class 0.
- 02Market dataPASS
1,334 joined sessions, zero blocking failures, two documented warnings.
- 03Feature contractPASS
23 trailing-only features on one common eligibility window.
- 04Development comparisonCOMPLETE
Five candidates, 12 purged folds, and 740 OOS dates.
- 05Deployment evidenceFAIL
Only 1 of 9 predeclared gates passes; the expanded repair audit also loses to the prior.
- 06RecommendationGUARDED
Directional no-go. The benchmark-only controller stays active.
- Asset
- ASTS
- Target
- Raw close higher in five sessions
- Source period
- 07 Apr 2021 — 29 Jul 2026
- Model rows
- 1,249
- Feature families
- 23 frozen + 24 audited challenger features
- Development runs
- 5 frozen + 8 repair candidates + fold-local prior
- Validation
- 12 nested outer folds · 5-session purge
- OOS evidence
- 740 dates · 09 Aug 2023 — 22 Jul 2026
- Post-hoc development ranking
- Run 5 · least-bad candidate
- Primary result
- Brier improvement −0.00170
- Stress result
- 1 of 9 gates passed
- Deployment recommendation
- NO-GO · benchmark-only fallback
The intervals still include no improvement.
Run 5's observed values sit on the weak side of both reference boundaries.
20-session 95% interval: 0.3800–0.5379
20-session 95% interval: −0.00498 to +0.00131
Required frozen-forecast tail rate: below 5%
Folds with positive Brier improvement; required at least 60%
Two of nine is the best individual count.
The same 740 dates informed every repair. Phase B can diagnose failure modes, but it cannot become a fresh confirmation set.
Prior-rank Ridge family output: G4 / G8. It clears the Run 1 comparison and fold-consistency gates, but its probabilities remain tightly compressed and qcut ECE is 6.028%. Median within-fold std. dev. 0.000569.
Empirical-bin family member: G6 / G8. It is the closest diagnostic to addressing spread, qcut ECE, and fold consistency together, but it was not selected and misses qcut ECE by 0.5738 percentage points. Median within-fold std. dev. 0.027692 / maximum 0.039467.
Non-flat state model: G7. It avoids near-constant forecasts and passes qcut ECE, but loses to the prior overall and reaches only half the folds. Median within-fold std. dev. 0.021151.
Neutral-anchor Core Ridge: G4. Only three folds find an eligible active configuration. Nine folds correctly emit the predeclared no-signal fallback. Median within-fold std. dev. 0.000000 across all folds / 0.006596 active-only.
No repair clears all nine gates, so the controller remains benchmark-only with no directional tilt.
G4 and G8 from the selected calibration family cannot be added to G7 from the non-flat state model or G6 and G8 from the unselected empirical-bin member. Neutral-core's G4 pass also belongs only to its own mostly-fallback probability path.
It passes G4 and G8 across 9 of 12 positive folds, yet remains compressed inside each fold: median within-fold standard deviation is 0.000569. qcut ECE is 6.028%, above the 5% gate.
It spans 32.25% to 54.10% and passes qcut ECE at 4.361%, but wins only 6 of 12 folds and has negative pooled Brier improvement. Its median within-fold standard deviation is 0.021151.
Neutral-anchor Core Ridge activates in 3 of 12 folds and falls back to the frozen prior in 9. It passes G4 only. The fallback gives a 0.000000 median within-fold standard deviation across all folds; the active-only median is 0.006596. It behaved exactly as designed: when inner evidence is ineligible, emit no signal.
Calibration family run log / Calibration family gates / Calibration family summary / Non-flat state run log / Non-flat state gates / Non-flat state summary / Neutral-core run log / Neutral-core gates / Neutral-core summary / Corrected legacy direction-aware run log / Corrected legacy direction-aware gates. The Step 6 workbook preserves Phase A and adds distinct Phase B sheets.
G1 · BrierFail · Positive lower boundFresh paper-forward probability scores
G2 · AUCFail · AUC ≥ 0.52; lower > 0.50Fresh discrimination evidence
G3 · Family-wiseFail · Tail rate < 5%Predeclare one challenger before new data
G4 · Core comparisonFail · Paired lower bound > 0Beat the core model robustly
G5 · Calibration slopeFail · 0.70 to 1.30Calibrate only after discrimination exists
G6 · Calibration interceptPass · Absolute value ≤ 0.10Preserve the frozen definition
G7 · Calibration errorFail · ECE < 5%Validate a prior-only calibration map
G8 · Fold consistencyFail · At least 60% positiveShow stability across new folds
G9 · SensitivityFail · All conditions passSurvive block, event, and regime checks
Run neutral now. Earn the right to tilt later.
The next step is a frozen controller plus one predeclared challenger on fresh outcomes—not another search over the same test history.
- P0Keep the directional model switched off.
The benchmark-only controller improves observed classification accuracy over Run 5 by 3.24 percentage points without inventing an edge.
- P1Freeze the nine gates and the Model 1 benchmark.
Keep the target, snapshots, features, folds, forecasts, and acceptance rules unchanged so the next study has an honest comparator.
- P2Treat the repair search as development evidence.
Twenty-four added features and eight repair candidates were inspected. The 740 OOS dates cannot be reused as a fresh holdout.
- P3Predeclare one mechanism-led challenger.
A new model must state its feature mechanism, training window, calibration rule, and all nine acceptance gates before new outcomes arrive.
- P4Accumulate a paper-forward period.
Only newly arriving observations can change the deployment recommendation. Passing a subset of the gates is not enough.
Model 1 evaluates five-session direction probabilities for ASTS. It is not investment advice and does not evaluate a trading strategy, positions, execution, costs, or portfolio risk.
A decision-ready record of the target, data, repair audit, benchmark-only controller, unchanged no-go verdict, and the exact evidence needed to reconsider it.
Step 7 now uses a stronger prequential simulation audit and higher Monte Carlo precision without reopening the model decision.