Skip to main content
GioModel
Launch Portfolio
ASTSStep 4
MODEL 1 · STEP 4

Train
through time.

Compare five development candidates inside the same nested, purged chronology—and let the training prior compete too.

MARKET CONTEXT · OUTSIDE MODEL FREEZEASTS

These readings update independently. They are not model inputs, deployment evidence, or a trading recommendation.

Latest price
$62.31
· fallback
20D momentum
-13.4%
Close-to-close · 20 completed sessions
RSI 14
47.2
neutral · Wilder method
Relative volume
1.0×
Versus the prior 20 completed sessions

TECHNICALS THROUGH · TIMESTAMPED FALLBACK

THE QUESTION

Which candidate improves probability forecasts through time?

“Best model” means the lowest out-of-sample probability error among the five candidates—not permission to deploy. The smoothed fold-local prior remains the benchmark every run must beat.

INPUT1,249 model rows
REPLAY12 nested outer folds
OUTPUT740 OOS dates × 5 runs
THREE TRAINING RULES

Put every choice behind the boundary.

Nested tuning is slower than one global fit. That cost buys a cleaner answer about how the process would have behaved at the time.

01

Nest the choices

Hyperparameters and calibration are learned from inner training chronology, never the outer test block.

02

Purge five

Five sessions separate training labels from the next test block so overlapping outcomes cannot cross the boundary.

03

Prefer simplicity

A one-standard-error rule selects the simplest configuration still supported by inner log loss.

TRAINING CONTRACTfive_run_walkforward_v2
FROZEN / V2
Candidates5 model families
Outer folds12 chronological
OOS period09 Aug 2023 — 22 Jul 2026
Primary scoreBrier improvement
RULESETTINGWHY
Outer minimum train504 sessions

Two trading years before the first test

Inner minimum train252 sessions

One year before tuning begins

Purge5 sessions

Matches the label horizon

Test block63 sessions

Final block contains 47 sessions

Outer folds12 expanding

One score per chronological test date

TuningPurged inner walk-forward

One-standard-error selection

CalibrationRegularized Platt map

Fit only to inner OOF probabilities

Primary metricBrier improvement

Compared with the fold-local prior

FIVE-RUN COMPARISON

All five lose to the prior.

Bars show pooled Brier improvement versus the fold-local prior. Zero is the required boundary; every candidate remains negative.

POST-HOC DEVELOPMENT RANKINGRun 5 · Least-bad candidate

Run 5 is the least weak candidate: AUC 0.4577, Brier 0.2529, and Brier improvement −0.00170. It beats the prior in 6 of 12 folds—but loses overall. This ranking was made after inspecting all five runs; it is not isolated selection evidence or a deployment recommendation.

ISOLATED SELECTION EVIDENCE · NESTED ONE-SEPrior selected in 11 of 12 folds

When the inner evidence is allowed to choose conservatively, it selects the smoothed training prior for 677 of 740 test rows and elastic net once for the remaining 63.

STEP 4 PROCESS GATE

Eight checks protect the replay.

The protocol can pass while the models fail. That separation is the point of an honest evaluation.

  1. 01

    Every run starts from the same 1,249 labeled, feature-complete rows.

  2. 02

    Each outer model trains on at least 504 earlier sessions.

  3. 03

    A five-session purge separates every training window from its test block.

  4. 04

    Twelve outer folds create 740 unique OOS predictions from August 2023 through July 2026.

  5. 05

    Inner tuning begins only after 252 training sessions are available.

  6. 06

    Scaling, hyperparameters, and calibration are fit inside training chronology.

  7. 07

    The final outer block is kept at its available 47 sessions rather than discarded.

  8. 08

    The smoothed training prior is recomputed independently inside each fold.

DEVELOPMENT EVIDENCE

Earlier ASTS results were already inspected before these five candidates were compared. All five runs are therefore development evidence and still require a newly accumulated paper-forward period before any deployment consideration.

PHASE B / DEVELOPMENT-ONLY REPAIR PROGRAM

Four audited outputs expose different failure modes.

Phase A above is unchanged. Every Phase B result reuses the same 740 dates, so these are diagnostic development replays, not additional holdouts or independent validation.

01

Prior-rank Ridge family output

It clears the Run 1 comparison and fold-consistency gates, but its probabilities remain tightly compressed and qcut ECE is 6.028%.

02

Empirical-bin family member

It is the closest diagnostic to addressing spread, qcut ECE, and fold consistency together, but it was not selected and misses qcut ECE by 0.5738 percentage points.

03

Non-flat state model

It avoids near-constant forecasts and passes qcut ECE, but loses to the prior overall and reaches only half the folds.

04

Neutral-anchor Core Ridge

Only three folds find an eligible active configuration. Nine folds correctly emit the predeclared no-signal fallback.

Separate Phase B candidates on the same 740 development dates
CandidateBrier gainAUCSlope / intercept / qcut ECEWithin-fold flatnessPositive foldsGate count
Prior-rank Ridge family output+0.0000500.444743-4.856800 / -0.349814 / 0.060283Median within-fold std. dev. 0.0005699 / 122 / 9 / G4 / G8
Empirical-bin family member+0.0008710.5022180.4540 / 0.01547 / 0.055738Median within-fold std. dev. 0.027692 / maximum 0.0394678 / 122 / 9 / G6 / G8
Non-flat state model-0.0033850.486449-0.853846 / -0.181012 / 0.043606Median within-fold std. dev. 0.0211516 / 121 / 9 / G7
Neutral-anchor Core Ridge+0.0000400.441582-3.876583 / -0.270358 / 0.060468Median within-fold std. dev. 0.000000 across all folds / 0.006596 active-only2 / 121 / 9 / G4
NO-SIGNAL FALLBACK

Neutral-anchor Core Ridge activates in only 3 of 12 folds. The other 9 folds emit the frozen fold-local prior exactly because no inner configuration clears its predeclared eligibility rules. That fallback is a safeguard, not an improvement.

PHASE B AUDIT FILES

Calibration family run log / Non-flat state run log / Neutral-core run log / Corrected legacy direction-aware run log. The Step 4 workbook preserves Phase A and adds distinct Phase B sheets.

STEP 4 OUTPUTfive_run_oos_v2

Seven hundred forty chronological dates with fold IDs, baseline probabilities, and five untouched candidate probabilities.

NEXTStep 5 · Stress the result

The post-hoc least-bad candidate now faces deployment gates.