Skip to main content
GioModel
Launch Portfolio
ASTSStep 5
MODEL 1 · STEP 5

Stress
the result.

Run 5 is the post-hoc least-bad development candidate. Nine predeclared gates decide whether that ranking is strong enough to consider deployment.

MARKET CONTEXT · OUTSIDE MODEL FREEZEASTS

These readings update independently. They are not model inputs, deployment evidence, or a trading recommendation.

Latest price
$62.31
· fallback
20D momentum
-13.4%
Close-to-close · 20 completed sessions
RSI 14
47.2
neutral · Wilder method
Relative volume
1.0×
Versus the prior 20 completed sessions

TECHNICALS THROUGH · TIMESTAMPED FALLBACK

THE QUESTION

Does the post-hoc least-bad candidate earn consideration?

The answer depends on evidence, not rank. Run 5 can be least bad among the five inspected candidates and still fail against the prior, uncertainty, calibration, stability, and null benchmarks.

REFERENCEFrozen Run 5 OOS path
CHALLENGEFour stress families
DECISION1 of 9 gates pass
FOUR STRESS FAMILIES

Try to disprove the post-hoc ranking.

The family-wise null circularly shifts the frozen OOS label path against frozen forecasts. It excludes shifts within 63 sessions of zero and protects against selecting the strongest of five runs.

01

Block uncertainty

Resample 20-session blocks 2,000 times so overlapping outcomes are not treated as independent rows.

02

Family-wise null

Circularly shift the frozen OOS label path and compare the strongest of five development runs.

03

Calibration

Audit slope, intercept, five-bin calibration error, and confidence coverage.

04

Sensitivity

Change block length, inspect regimes, remove extreme-event windows, and measure feature drift.

STRESS RESULTS

Run 5 remains weaker than the prior.

The primary 20-session bootstrap keeps zero inside the Brier improvement interval and 0.50 inside the AUC interval.

Run 5 ROC AUC0.4577

Run 5 · shallow boosting

Brier improvement−0.00170

Worse than the fold-local prior

20-session AUC 95% CI0.3800–0.5379

Lower bound remains below 0.50

Family-wise tail rate96.59%

Frozen-forecast circular-shift null

Five-bin ECE6.44%

Above the 5% evidence gate

20-SESSION BLOCK BOOTSTRAP · 95% INTERVALS2,000 repetitions
DEPLOYMENT RECOMMENDATIONNO-GO — DEVELOPMENT EVIDENCE IS INSUFFICIENT.

Run 5 remains the post-hoc least-bad candidate, but that label only ranks the five inspected runs. Eight of nine deployment evidence gates fail, so no model should move forward.

GATES PASSED1 / 9
STEP 5 EVIDENCE GATES

One pass cannot rescue eight failures.

These are conservative research guardrails, not nine independent hypothesis tests. The observed and required values stay visible.

Nine deployment evidence gates for the post-hoc Run 5 candidate
GateRuleObservedRequiredResult
G1Positive Brier improvement with 20-session lower bound above zero−0.00170; lower −0.00498Point > 0; lower > 0FAIL
G2ROC AUC and its 20-session lower bound clear no-skill0.4577; lower 0.3800AUC ≥ 0.52; lower > 0.50FAIL
G3Family-wise frozen-forecast circular-shift tail rate96.59%< 5%FAIL
G4Run 5 improves Brier score on Run 1 with paired lower bound+0.00197; lower −0.00222Point > 0; lower > 0FAIL
G5Calibration slope−1.3190.70 to 1.30FAIL
G6Absolute calibration intercept0.098≤ 0.10PASS
G7Five-bin expected calibration error6.44%< 5%FAIL
G8Outer folds with positive Brier improvement6 / 12 · 50%≥ 60%FAIL
G9Block-length, event-window, and supported-regime sensitivityLower −0.00503; event −0.00242All sensitivity conditions passFAIL
PHASE B / SEPARATE REPAIR AUDITS

No lawful candidate fixes all three weaknesses.

The repair program tests flatness, equal-frequency qcut calibration error, and fold consistency separately. Each row keeps its own nine-gate denominator.

CandidateProbability spreadqcut ECEPositive foldsIndividual result
Prior-rank Ridge family output0.4642 to 0.4976
Std. dev. 0.0108
Median within-fold std. dev. 0.000569
0.0602839 / 122 / 9 / G4 / G8
Empirical-bin family member0.4257 to 0.5433
Std. dev. 0.030792
Median within-fold std. dev. 0.027692 / maximum 0.039467
0.0557388 / 122 / 9 / G6 / G8
Non-flat state model0.3225 to 0.5410
Std. dev. 0.0370
Median within-fold std. dev. 0.021151
0.0436066 / 121 / 9 / G7
Neutral-anchor Core RidgePrior fallback in 9 / 12 folds
Std. dev. 0.0120
Median within-fold std. dev. 0.000000 across all folds / 0.006596 active-only
0.0604682 / 121 / 9 / G4
Legacy direction-aware Core RidgeCompressed legacy output0.0661264 / 121 / 9 / G4
DO NOT AGGREGATE GATES

The selected calibration family passes G4 and G8. The non-flat state model passes G7. The unselected empirical-bin member passes G6 and G8. Neutral-anchor and the corrected legacy direction-aware audit pass G4 only. These passes belong to different probability paths and cannot be combined.

FLATNESS VERSUS CALIBRATION

The selected prior-rank Ridge family reaches 9 of 12 positive folds, yet its median within-fold forecast standard deviation is only 0.000569 and qcut ECE is 6.028%. The non-flat model spans 32.25% to 54.10%, has a 0.021151 median within-fold standard deviation, and passes qcut ECE at 4.361%, but wins only 6 of 12 folds and loses to the prior overall.

CLOSEST UNSELECTED DIAGNOSTIC

The empirical-bin family member was not selected. It passes G6 and G8, posts +0.000871 Brier improvement and 0.502218 AUC, spans 42.57% to 54.33%, and reaches 8 of 12 positive folds. Its median within-fold standard deviation is 0.027692, but qcut ECE is 5.5738%, missing G7 by 0.5738 percentage points. It remains development-only and does not raise the best gate count above 2 of 9.

CORRECTED LEGACY AUDIT

This corrected audit uses the original frozen reference, qcut ECE, seed 20260730, and 5/20/63-session blocks. The superseded legacy gate file must not be used. Its only pass is G4: +0.003139 / lower +0.000281. qcut ECE is 0.066126, and the family-wise tail rate is 0.970779.

PHASE B EVIDENCE FILES

Calibration family gates / Non-flat state gates / Neutral-core gates / Corrected legacy direction-aware gates. The Phase A table above remains unchanged.

CLASSIFICATION BOUNDARY

This is a direction-probability exercise, not a trading backtest. It does not model positions, overlapping trades, execution, transaction costs, or portfolio risk.

STEP 5 OUTPUTmodel1_evidence_gate_v2

One frozen audit of uncertainty, family-wise null evidence, calibration, folds, regimes, drift, and event sensitivity.

NEXTStep 6 · Summarize the model

Turn the evidence into a durable recommendation.