Where the models stand.
None of them is approved.
This note ranks models by how much evidence they have passed. That is not a view on any company, and it is not a ranking of what to own. Every model in this universe is NO-GO.
Download the underlying artifactNo model can pass today, and that is arithmetic.
G9 requires forward paper evidence — outcomes observed after a model is frozen. On release day there are none, so G9 fails by construction for every model in the universe. The only cure is elapsed time. Two other gates behave the same way: G1 sample size is fixed by a company’s listing date, and G4 cross-verification needs a comparison workbook that does not yet cover the newer names.
So the honest ceiling for the best model here is 7 of 8, not 8 of 8. A model failing only G1, G4 and G9 is not a weak model. It is a young one with a missing comparison artifact, and counting those failures the same way would rank recency as a defect.
Three at six of eight, strong for opposite reasons.
| Ticker | Gates | Ceiling | Still addressable | Coverage h5 / h10 / h21 | vs baseline |
|---|---|---|---|---|---|
| PYPL | 6 / 8 | 7 / 8 | G6 | 75% / 79% / 79% | beats / loses / loses |
| CRM | 6 / 8 | 7 / 8 | G6 | 75% / 78% / 77% | beats / loses / loses |
| ORCL | 6 / 8 | 7 / 8 | G5 | 69% / 64% / 59% | beats / beats / beats |
These are the only models at six, and the only ones whose ceiling is seven — one structural failure rather than two or three. They are strong for opposite reasons, which is the part worth reading.
PYPL and CRM have honest bands. Coverage lands inside the required [75%, 85%] at every horizon. They fail only G6, the comparison against a scaled random walk, and only at the longer horizons, by roughly 0.1% to 3%.
ORCL is the mirror image. It beats the baseline at all three horizons — the only one of the three that does — but its bands are far too narrow: 59% coverage at the 21-session horizon against a label claiming 80%. It earns its score by under-claiming uncertainty, which is why it fails G5 rather than G6. A band that covers 59% of outcomes while calling itself P10–P90 is not a more decisive forecast. It is a mislabelled one.
WMT is arguably the most complete, at five of eight.
WMT is the only model in the universe that both beats the baseline at every horizon and keeps its coverage inside the band. Its single addressable failure is G2, data quality — a fixable input problem rather than a modelling one. Ranked purely by gate count it sits in the second tier; ranked by whether its forecasts are calibrated and better than a naive alternative, it leads.
That gap between the count and the substance is why this note exists rather than a leaderboard.
The weakest three, and why they are weak.
These are the original nine-gate builds, and they are weak on their own terms rather than because they are new. ASTS is the clearest case: Phase A passed one gate of nine, and the best individual Phase B repair reached two. Its published boundary already says the simulation is for exploring conditional ranges, not for reading a direction.
Weakness here is not a statement about the companies. It is a statement about how much the data supports a forecast of them at these horizons, which is a different question and the only one this site answers.
Most failures are near misses, which is the danger.
Of 51 walk-forward horizons measured, 30 lose to the random-walk baseline — and 25 of those lose by 3% or less.
That proximity cuts both ways. A genuine improvement could plausibly cross a 3% gap. So could a change tuned against the same origins it is then measured on, for no reason at all. Any attempt to close these gaps has to be predeclared and evaluated on origins it has not seen, or the result means nothing.
G5 and G6 also pull against each other. Widening an interval raises coverage toward G5 and worsens the score against G6; narrowing does the reverse. Of the horizons measured, 9 are too narrow and 12 too wide. There is no free direction.