What we measured — and what we retired because of it.
Every other options analytics service shows you a backtest
that worked. The selection is the problem: you are only ever shown the
survivors. These are ours, including the tests that cost us
features.
Each finding below was produced by a walk-forward measurement in this
codebase, not a curve fit after the fact. Where a result is negative, it
is stated as a negative. Where a result is unmeasured, it is
stated as unmeasured rather than quietly implied to be good —
those are different claims and conflating them is how this category
misleads people.
We have deleted working code over these numbers: a scanner's
directional claim, a circular exhaustion claim, and an “edge”
claim on our own marketing page. The text below is rendered directly from
the product's explainer module, so what you read here is what a paying
subscriber reads inside the product.
01 · confluence.score
Measured
The confluence score has been measured
The score has now been scored, walk-forward, using the shipped compute_confluence_score over 34,297 bars and fifteen years. AT THE LIVE BOT'S OWN GATE (|score| >= 0.20, confidence >= 45) AND ITS OWN five-session hold, it carries no directional information the measurement can establish. On the bot tier (^GSPC): 327 onset-deduplicated entries, 55.0 % correct against an instrument-and-month matched base of 57.2 % — lift −2.1 pp, interval [−6.7, +2.1] pp, which straddles. Against a second, unstratified base of 53.8 % the lift is +1.2 pp [−5.6, +7.3] — also straddling. The wider index and wide tiers land the same way, and NOT ONE of the 4,113 cells in the shipped table survives the three-hurdle verdict gate. Two caveats travel with every number. First, only six of the score's ten components have fifteen years of history here — 62 % of the shipped weight — so this is a result about that 62 %, and any information in the live score would have to come from the option-derived remainder. Second, the same measurement run on a bar that is NOT an entry, carrying the entry's own direction claim, scores comparably, which is why nothing here is published as an inversion either.
02 · exhaustion.evidence
Measured
How the Buy Zone was measured
The question a Buy Zone answers is "does price come down here". So the measured outcome is exactly that: did the low of the window trade at or through the level. It is the same test detect_exhaustion_alerts already runs when it sets pb1_hit.
A rate on its own says nothing, because a shallow pullback from a two-day high is an ordinary thing for an index to do. So every rate is paired with the identical construction on bars belonging to no trigger streak, in the same instrument and the same month — matching the instrument matters, because a trigger needs two sessions under 0.60 % range and therefore fires on the quietest names, and a null pooled across everything would be a yardstick from a different market.
A rate and its base rate must describe the SAME events, or the difference between them is not a lift. An event whose instrument and month hold no signal-free bar has no base rate, so it leaves the cell entirely rather than counting toward the rate alone — the count of those is published on every cell as unmatched, and it is zero for every figure quoted here.
Three controls decide whether a lift is real. Intervals come from a calendar-month block bootstrap, because triggers arrive in clusters and on several instruments at once. The lift and its null are resampled TOGETHER, so the base rate is not treated as an exact constant. And the whole measurement is repeated on bars 35 and 70 sessions either side of each trigger, with each placebo removed from its own comparison group so it is not scored partly against itself: if those placebo bars showed the same lift, the finding would be about the era, not the trigger. They show roughly zero.
Every window is published, from day 8 out to day 60, so no window can be chosen after seeing which one flattered the model. So is the one specification choice that moved a headline number by more than a point — whether the signal-free arm keeps bars whose next five sessions contain a streak. Keeping them is the shipped and more conservative reading; dropping them reports +14.7 pp instead of +7.5 pp, and both ship under null_specification. Re-run the lot with python3 -m heatseeker.data.exhaustion_outcomes.
03 · exhaustion.validation-evidence
Measured
The day 8-9 close does carry information
A close below the trigger day's close on day 8 or day 9 happened on 44.2 % of all 480 triggers — this test needs only nine forward sessions, so none is censored — against 30.1 % on matched signal-free bars: a lift of +14.0 pp, interval +8.9 to +19.5 pp. Of everything the Q.E. model asserts, this and the shallow PB1 band are the two that separate from chance.
04 · exhaustion.arrival-evidence
Measured
The 16-19 window is not where the low lands
Of the 477 triggers whose day 8-40 hunt completed, the deepest low landed inside days 16-19 11.5 % of the time, against 10.6 % on matched signal-free bars — a lift of +0.9 pp with an interval of -3.4 to +5.1 pp. Four sessions out of thirty-three is about 12 % by chance alone. Treat the band as a calendar marker, not a forecast.
05 · exhaustion.regime-evidence
Measured
The VIX buckets are not separated by the data
CALM (n=368) shows a PB1 lift of +8.5 pp on an interval of +3.5 to +12.4 pp. NORMAL (n=98) shows +5.2 pp on -3.9 to +13.6 pp — its own interval contains zero, so at that sample size NORMAL is not distinguishable from no effect at all, which is a different statement from "a smaller effect". The two cannot be told apart either, which is why the CALM/NORMAL split in the target table is unsupported. ELEVATED has 11 triggers in fifteen years, below the 20 this platform will report on. CRISIS has zero: a trigger needs two quiet sessions, and VIX above 30 does not produce them. The CRISIS row of the target table is inherited, never fitted.
06 · exhaustion.deep-evidence
Measured
The Deep tier is untested, not conservative
The -16.18 % extension was reached on 1 of 477 triggers inside the day 8-19 window in fifteen years — 8 of 477 by day 40 — no more often than on bars with no trigger at all. It is a drawn extreme, not a measured expectation.
07 · exhaustion.lag
Measured
The alert is dated one session late
A trigger is dated on the LAST zero-succession day of its streak, and whether today is the last one is not known until tomorrow closes — if tomorrow is also quiet the streak extends and the date moves. So a Q.E. alert is a one-session-lagged label, not a same-day signal. Every measured number on this page starts its forward walk at day 6, well past that, so nothing here depends on the difference.
08 · exhaustion.pb1
Measured
PB1 · Fib 0.236
Fib 0.236 × 10 = 2.36 % below the Day 4-5 pivot high, identical in every VIX regime. That much is the framework's definition. What follows is measurement.
Two levels, not one, and the copy has to say which. The detector scores pb1_hit against the 2.124 % variance line — PB1 widened by the prep multiplier — so that is the cell the headline rate belongs to. Over the 477 triggers whose day 8-19 window resolved, out of 480 fired across SPX, SPY, QQQ and IWM from 2011 to 2026, price traded into it 43.2 % of the time against 35.7 % on bars where no trigger fired — same instrument, same calendar month. Lift +7.5 pp, 95 % month-block interval +3.5 to +11.2 pp. At the 2.36 % line the ladder actually draws, the same test reads 38.4 % against 31.6 %, lift +6.8 pp, interval +2.5 to +11.0 pp. Both clear their base; neither is the other.
The trigger is a COMPRESSION event, so the obvious alternative reading is that price simply moves further afterwards in both directions. The test for that is the mirror: the identical construction run UPWARD off the Day 4-5 pivot low. The gap between the two — `+11.1 pp`, interval `+4.0 to +18.0 pp` — is the directional finding, and it is the figure to read. The raw upside lift of -3.6 pp is NOT a standalone finding: the two tests hang off opposite ends of the Day 4-5 spread, and a trigger's spread is 1.22 % against 2.13 % on a signal-free bar, which makes BOTH of its tests mechanically harder. That headwind cancels in the difference and does not cancel in either raw lift — where it works AGAINST the downside finding, so the +7.5 pp is if anything conservative.
The control that matters most: repeating the whole measurement on bars 35 and 70 sessions either side of each trigger — same instruments, same era, no trigger — returns lifts of +2.7, +0.7, +0.2 and +0.8 pp, every interval spanning zero. What is being measured is the trigger, not the period.
What this does NOT say. It does not say the band is the bottom; reaching a level and the level holding are different claims and only the first is measured. It does not say 43 % is a forecast — on the 2020-2026 window alone the same cell reads 54.1 % and its lift interval -1.0 to +17.8 pp contains zero, which is exactly why a bare rate must never be quoted without its base. And 477 triggers is not 477 independent events: SPX and SPY are one index, and 42.5 % of the triggers are the same day counted twice.
09 · exhaustion.pb2
Measured
PB2 · the two deep lines
PB2 is drawn twice: the primary line is the p90 stretch-bottom fitted per VIX regime, the secondary is the Fib-pure 6.18 % (0.618 × 10). Showing both was meant to keep the empirical and the harmonic readings side by side.
Measured over 477 triggers, 2011-2026, in the detector's own day 8-40 hunt window: the 6.18 % secondary was reached 15.5 % of the time against 12.2 % on matched signal-free bars — lift +3.3 pp, interval +0.8 to +6.3 pp. That interval clears zero, so it is a separation, but a small one that nearly touches zero and is one cell among the many this table publishes; it is not the kind of margin the PB1 rung carries. The 7.86 % CALM primary reads +1.7 pp on -0.3 to +3.7 pp and the 10.00 % primary -0.5 pp on -1.5 to +0.6 pp. Both intervals contain zero: no measured edge.
In the tighter day 8-19 window — the one the ladder is hunted in first — none of the three separates at all: the 6.18 % line reads +0.2 pp on an interval of -1.2 to +1.7 pp.
So the honest reading of the deep ladder is mostly descriptive: it shows how far a selloff would have to run to reach each Fibonacci rung. Only the Fib-pure 6.18 % line, and only over the full day 8-40 hunt, does better than its own base — and the measured margin there is a few points, not a forecast.
The levels are arithmetic. The claims are what we keep honest.
Dealer gamma, exhaustion ladders and the succession grid compute what
they compute. What this page is for is the line between that and what
anyone can prove it predicts.