Xiphers Research

Xiphers Research publishes pre-registered quantitative tests of trading theses on ES and NQ tick data. Every thesis is fixed together with its criterion before the run starts, and is then measured against Control Arms — same geometry, same session, only the link to the price is cut. Each verdict is given per asset: a finding on NQ is not yet a finding on ES.

Quantitative Validation — Option Flow — Gexbot

cl1Does the SPX spot's side of Zero Gamma predict short-horizon direction?NQNo edgeESNo edge

NQ

Over 106 Dev sessions and 7,491 anchors the hit rate P(+160 before −160) reaches 51.9 % against a 50.7 % baseline (Δ +1.21 %, above every Control Arm), but the predicted sign shows up in only 3 of 7 hour buckets — the pre-registered breadth criterion asks for at least 4, so the thesis fails on this config. One unregistered cross-tab cell — spot below Zero Gamma while price is above VWAP — reached 55.6 % against a 48.2 % baseline (Δ +7.40 %, n 541 on 76 days) and was promoted to a thesis of its own with its own forward test; it does not change this verdict.

Full report →

ES

Over the same 106 Dev sessions and 6,402 anchors the hit rate P(+40 before −40) reaches 51.5 % against a 50.5 % baseline (Δ +1.01 %, above every Control Arm), again with the predicted sign in only 3 of 7 hour buckets — the pre-registered criterion fails on ES exactly as it does on NQ. The same contradiction cell against VWAP looks even larger here (56.7 % against 47.1 %) but stays exploration: at z = +1.56 it misses the z ≥ 2 entry ticket for promotion.

Full report →

cl1aDoes the position between Zero Gamma and the nearest wall drive the CL1 edge?Planned

CL1 broken down by the position parameter d = (spot − zero gamma) / (wall − zero gamma) — the hypothesis is weakest near the wall, strongest in the middle. Waits for the forward test of CL1.

cl1bMatteo — a market maker tested every retail belief about gamma exposure and found one that survived: not the side of the gamma flip, but the moment price crosses it. Does that reproduce here?NQOpenESOpen

NQ

Open, and deliberately not called a null result. Built on the author's own source (NDX 0DTE levels) and measured his way — a fixed 15-minute hold, no stop, no target — the crossing comes out at +2.09 ticks against the shifted-random baseline over 437 signals on 103 days, z +0.33. That is one sixth of what he reports, but this sample cannot tell the two apart: the control arms scatter far enough that a verdict would need ±12.7 ticks, and his own figure translates to about +11.5 ticks per trade. Had his effect been present here exactly as published, it would have landed at z ≈ 1.8 and failed our own bar. 108 sessions against his 820 is the whole story. Where the sample does have the power to see something, it agrees with him, and that is the finding worth keeping: without his cluster rule the crossing loses money — −9.24 ticks here and −1.51 ticks at z −2.83 on the largest cell of all — so repeated crossings inside one cluster are not merely noise, and his 35-minute lockout lifts that back to roughly zero. His wall claim points the other way here (following the break gives −11.4 and −22.3 ticks) but sits far below the ±36 and ±31 ticks a verdict would need, so neither direction is shown. One thing needs no sample size at all: the contrast the video is built on — side versus crossing — does not exist in the data. The disagreement cells are empty, because five minutes after an up-crossing the spot is above the level every single time. The crossing is not different information from the side, it is a selection of 5.6 % of the anchors. Source: Matteo Conti's video, https://www.youtube.com/watch?v=bMs9PzQzrqM

Full report →

ES

The same NDX signal series traded on MES, because every study on this site runs on both instruments — the author's claim is about NQ, so the sizes below do not translate to his numbers. The crossing gives −0.34 ticks against the baseline over the same 437 signals, z −0.11, hit rate 51.3 % against 49.6 % (z +0.75): nothing, and the total points the wrong way while the breadth count formally passes. The cluster finding is sharper here than on NQ and it is the one solid cell in the study: without a lockout the crossing sits at −1.51 ticks with z −2.83 over 1,085 signals, and the lockout lifts it to zero. Both walls lose when followed, again far short of a verdict. Held longer — 30 and 60 minutes, his 50-minute refinement — the sign turns positive on every arm (+1.46 and +1.83 ticks) but never leaves the control band. Only the 1-minute horizon clears |z| ≥ 2 on both instruments, which is microstructure immediately after the crossing and is gone before commission is paid.

Full report →

cl2Do stable major walls act as reversal zones?Planned

The event is the approach, not the touch — how deep price got is a reported dimension, never a filter condition. Needs a wall stability rule over a trailing window and two control arms.

cl3aDoes the option positioning before the bell predict the first hour and the direction of the day?Planned

A daily study on the first valid OI net GEX value of the day. Two separate verdicts from one data path: volatility of the first hour, and direction of the RTH day.

cl3bDoes the sign of classic volume net GEX predict short-horizon direction?NQPreconfirmed — forward test pendingESPreconfirmed — forward test pending

NQ

On the frozen threshold θ = 150,000 the hit rate P(+160 before −160) reaches 55.1 % against a 50.6 % baseline (Δ +4.53 %, n 3,111 over 106 Dev sessions), above every Control Arm and with the predicted sign in 5 of 6 hour buckets — the largest Δ any study in this chapter has produced. Two honest caveats: the forward test is still pending, and the threshold sits at the upper edge of the grid, so the grid does not bound the edge from above.

Full report →

ES

On the same frozen threshold θ = 150,000 the hit rate P(+40 before −40) reaches 53.4 % against a 50.8 % baseline (Δ +2.62 %, n 2,378 over 106 Dev sessions), above every Control Arm and with the predicted sign in 4 of 6 hour buckets — the Dev criterion is met, at roughly half the Δ seen on NQ. Forward test pending, and the same edge-of-grid caveat applies.

Full report →

cl4Can the max-change strikes be derived from the ladder, and do they yield candidates?Planned

Labelled exploration, no pre-registered thesis and therefore no verdict. The strikes are not a field in the stream; they have to be derived from the ladder difference over 1/5/15/30 minutes and checked against the frontend that displays them.

st1Does the sign of state net GEX predict short-horizon direction?NQPreconfirmed — forward test pendingESPreconfirmed — forward test pending

NQ

On the frozen threshold θ = 1,500 — the user's live setting, fixed before the run — the hit rate P(+160 before −160) reaches 53.2 % against a 50.3 % baseline (Δ +2.83 %, n 4,333 over 105 Dev sessions), with the predicted sign in all 7 hour buckets and above every one of the 8 Control Arms. That clears the Dev criterion; the forward test on data that did not exist at the freeze is still pending, so this is not yet a confirmed edge.

Full report →

ES

On the same frozen threshold θ = 1,500 the hit rate P(+40 before −40) reaches 52.4 % against a 50.0 % baseline (Δ +2.41 %, n 3,548 over 105 Dev sessions), above every Control Arm and with the predicted sign in 4 of 7 hour buckets — enough for the Dev criterion, but visibly less broad than on NQ. The forward test is still pending, so the finding stands as strong but unconfirmed.

Full report →

st2Does the same max-change screening show anything on the state ladder?Planned

Conditional by design: it only runs if the ladder screening above produces candidates. If it does not, this entry is closed unworked, and that is recorded as the result.

st3Does state net GEX carry more when several underlyings agree — SPX, SPY, NDX, QQQ, alone and combined?Planned

ST1 ran on one underlying per instrument. This one runs the same direction test on ES and NQ for all 15 subsets of {SPX, SPY, NDX, QQQ}: four singles, six pairs, four triples, the full set. A subset votes only when every member shows the same sign above its own threshold. Judged: the four singles, the natural pairs and the quad, Bonferroni-corrected, against the best single on the same anchors; every other subset is reported as exploration. Waits for frames of the four underlyings.

of1Do extremes in the GEX orderflow mark local turning points?Planned

Positive spikes are supposed to mark local tops, negative ones local bottoms. Definition, grid and debouncing are taken over from the quadrant study without change.

of2Fredy Signals: do the four convexity × GEX quadrants mark tradeable setups?Planned

The four frontend setups — Upside Reversion (++), Upside Continuation (−−), Bottom Reversion (+−), Downside Continuation (−+) — as an event study on the orderflow stream, each quadrant broken out on its own. This study also fixes the spike definition and the debouncing for the whole orderflow block — the studies below inherit it unchanged, otherwise they would be measuring different events.

of3Does the aggregate DEX quadrant at 10:00 set the bias for the day?Planned

A snapshot of aggregated call and put DEX classifies the day into four quadrants. Two separate verdicts — direction and volatility — plus mandatory side findings, because the quadrant is cumulative and can flip during the day.

of4Is net convexity a risk gauge, and does it carry direction as well?Planned

Two parts judged separately, both running along the day rather than from a snapshot: cumulative series are read at the anchor, never frozen once in the morning.

of5Do vanna and charm move the last hour into the close?Planned

A direction thesis for the last hour only. Before 15:00 New York none of this is registered as a direction thesis — the documentation of the feed itself describes front-run reversals there.

of6Do DEX orderflow spikes mark pivots, regardless of their own sign?Planned

The sign of the spike does not set the direction here: the spike marks the pivot, and the trade is against the move that preceded it. Own grid — the scale is roughly ten times smaller.

of7After a strong GEX orderflow spike on SPX (|gexoflow| > 1,000), which way does price go — and does it depend on whether net GEX is positive or negative?Planned

Exploratory by design: no sign is registered. Four cells — spike up / spike down × GEX+ / GEX− — each reported for both directions against two control arms, one of them same-regime, same-hour anchors without a spike, so that a finding belongs to the spike and not to the regime it lands in. Threshold 1,000 sits above the 99th percentile of the stream; 500 and 2,000 run alongside. Anything that stands out becomes its own pre-registered thesis and goes forward.

of8Does a large positive GEX orderflow spike early in the session push price into an upward drift?Planned

A continuation reading, and it contradicts the turning-point study above, where the same positive spike is registered as a local top. Two design decisions are settled before it runs. The window: across 143 sessions the same threshold produces 193 morning spikes and 9,989 in the closing quarter-hour — a factor of 52 in a twelfth of the time — so restricting the study to the first four and a half hours is part of the thesis, not a later clean-up, and the late block is reported alongside without a verdict. And the target: the mechanism claimed here is that size moves price, which predicts more than a hit or a miss — it predicts that a bigger spike moves price further. The primary arm is therefore the rank correlation between spike size and forward return across every anchor, with hit rates per threshold as the secondary reading. Those are judged by an exact binomial test rather than a minimum anchor count, so that an event which is rare but almost always right can still count as evidence; the rarest cell carries too few anchors to ever reach a verdict and is reported as such, which is not the same as finding nothing.

gv1Does the same volatility move mean two different markets under positive and negative gamma?Planned

An interaction test, not two separate ones: volatility metrics given the sign of net GEX. The claim under test is that dealers amplify a move under negative gamma and dampen it under positive gamma. Waits for the volatility metrics below; the cells will be small, so the likely outcome is exploration awaiting a forward test.

Quantitative Validation — Technical Analysis

px1Does the price's side of the RTH session VWAP predict short-horizon direction?NQNo edgeESNo edge

NQ

Pre-registered on Dev (362 sessions) and Val (62 sessions), and failed on both: the hit rate P(+160 before −160) is 49.7 % against a 49.7 % baseline on Dev (Δ −0.07 %) and 49.2 % against 49.5 % on Val (Δ −0.26 %), so the Val sign points the wrong way. The verdict covers this config and this window only — the yearly cut on the full history shows the effect was alive in 2021–2023 (Δ up to +3.24 %) and gone from 2025 on.

Full report →

ES

Same design on ES, same outcome: 49.5 % against a 49.9 % baseline on Dev (Δ −0.36 %, 362 sessions) and 49.4 % against 49.6 % on Val (Δ −0.20 %, 62 sessions) — the Val sign is wrong, the thesis fails. The yearly cut runs parallel to NQ: Δ +2.05 % in 2021, −0.53 % in 2025.

Full report →

px1sDoes the price's side of the 18:00-anchored session VWAP predict short-horizon direction?NQNo edgeESNo edge

NQ

The overnight-anchored variant clears Dev by a hair — 49.8 % against a 49.7 % baseline (Δ +0.13 %, 362 sessions) — and then fails the single Val run at 49.8 % against 49.9 % (Δ −0.07 %, 62 sessions). Registered as the secondary variant to PX1, it does not rescue the VWAP thesis: the yearly cut shows the same decay, Δ +2.68 % in 2021 down to −0.08 % in 2025.

Full report →

ES

On ES the variant fails both roles outright: 49.7 % against a 50.2 % baseline on Dev (Δ −0.43 %, predicted sign in only 3 of 7 hour buckets) and 50.0 % against 50.3 % on Val (Δ −0.32 %). The yearly cut matches the NQ picture — Δ +2.14 % in 2021, −0.60 % in 2025.

Full report →

px2Is a NQ breakout that ES does not confirm a fakeout?Planned

A replication with the full measurement protocol — control arms, hourly breakdown, costs — of an older result that was only ever recorded as a raw win rate. Its dev and val windows are burned for this question, so the verdict falls on the untouched holdout alone.

px3Does a large opening move continue?Planned

The opening move is measured after the bell; the overnight session only enters as the yardstick that says what counts as large. Same honesty clause as above — verdict on the holdout only.

px4Do the first three trades of the day work on the side of the session VWAP?Planned

Three independent tests at 09:30, 09:35 and 09:40, each judged on its own — the question is whether the first has an edge, the second, and/or the third, not whether the sequence works. Direction comes from the session VWAP, the only VWAP line that already exists at the bell.

px5Do the VWAP sigma bands work as a distance filter — and are they what the GEX filter was measuring?Planned

Two parts with the same new ingredient, the volume-weighted standard deviation around the RTH VWAP. The second part is a price control: high absolute net GEX marks trend days, and on trend days price is far from the VWAP — so the GEX filter may be an expensive detour for a plain distance rule. Currently on hold at the user's request.

px6Does the opening playbook hold — four scenarios from where the RTH opens relative to yesterday's value area?Planned

Four sub-theses, each with its own entry level, its own exit level and its own verdict; there is no overall "the playbook works". Only days whose previous session was a D profile count. Waits on two definitions that have to be fixed before the run: what counts as the outer edge of the value area, and what makes a profile a D.

rp1Matteo — a published VWAP strategy claims a 64–65 % win rate on NQ, 2020–2024. Does it reproduce here — and does the win rate mean anything at a 1 : 0.5 risk-reward?NQPreconfirmed — forward test pendingESOpen

NQ

The claimed win rate reproduces: 63.2 % over 3,168 trades in the author's window 2020–2024 and 64.7 % over 1,284 trades forward (2025 → 2026-08), in every single year between 61 and 66 %. But at 80 points of risk against 40 the geometry alone hands out roughly 60 % — the question is what the rule adds on top. Against the strictest null (same entries, side drawn once per day and arm) the rule earns +0.025 R per trade against −0.011 in the author's window, above all 8 Control Arms, z +3.4; forward +0.032 against +0.010, z +1.5 — the sign holds, the margin is thin. The information sits in the trend context (price above a rising VWAP, +0.1 % hourly drift), not in the pullback candle: entering the same side five minutes later does just as well. Slippage is not included (break-even slippage +4.1 / +5.1 ticks per side); the author's own claim — a prop-firm pass rate — is not measured here. A first version of this study read "random beats the rule by 10 points"; that null looked into the future and was withdrawn. Source: Matteo Conti's video, https://www.youtube.com/watch?v=wm4A6qo0g3I

Full report →

ES

The rule is built for NQ and runs here unchanged on MES — same 80/40/50 point brackets, which on ES are far wider in relative terms, so only 1,248 trades fill in the author's window and 489 forward. Win rate 53.3 % / 52.8 %, +0.005 / +0.006 R against a baseline of −0.006 / +0.004 (z +2.0 / +0.3). Not calibrated for ES, not judged; shown for completeness because every study on this site runs on both instruments.

Full report →

Regime & Volatility

Four questions asked at one fixed moment: 09:30 New York, the RTH open. Everything before it may be used — the previous days, the opening print, every overnight tick — and nothing after it. The studies are walk-forward: every test year (2023, 2024, 2025, 2026) is forecast by a model that has only seen the years before it, refit once a year. Volatility turns out to be readable at that moment, direction and day type do not.

rv1Can we tell at 09:30 whether today's RTH range will be a big one (relr >= 1)?NQPreconfirmed — forward test pendingESPreconfirmed — forward test pending

NQ

Yes, and clearly: the L2 logistic regression on the 15 preregistered set C features reaches a BSS of 13.8 % (block bootstrap 95 % [9.9 ; 17.7]), AUC 71.1 % and 65.0 % accuracy against a base rate of 51.8 %, over 905 walk-forward days from 2023 to 2026 — every test year positive, the weakest 2023 at BSS 7.0 %. It is sharp as well as calibrated: on 14.0 % of the days it calls p >= 0.7 and hits 84.3 % of them. The forward test on days that did not exist when the study was frozen is still pending.

Full report →

ES

The same model is even stronger on ES: BSS 17.8 % (95 % [12.4 ; 22.9]), AUC 74.3 % and 68.0 % accuracy against a base rate of 50.0 %, over 904 walk-forward days, with the best year 2025 at BSS 25.1 %. Sharpness runs both ways — p >= 0.7 on 16.0 % of the days hitting 84.1 %, p <= 0.3 on 16.9 % hitting 79.7 % — so the finding is strong but still waiting for its forward test.

Full report →

rv2Where does the RTH close relative to yesterday's value area, given where it opened at 09:30?NQPreconfirmed — forward test pendingESPreconfirmed — forward test pending

NQ

The open position carries the day: opening above the previous value area, the RTH closes above it on 71.6 % of the days (n 342); opening below, it closes below on 64.3 % (n 249) — against an unconditional 42 % / 30 % split. Opening inside the value area says almost nothing (42.5 % above, 30.8 % inside). This is a conditional frequency, not a model, and it holds in all four test years; the forward test is still pending.

Full report →

ES

ES shows the same shape, a touch more symmetric: opening below the previous value area, the RTH closes below on 69.1 % of the days (n 243); opening above, it closes above on 68.6 % (n 347). Opening inside stays a coin toss across three cells (36.2 % inside, 37.4 % above). A table, not a model — and one that still has to survive a forward test.

Full report →

rv3Can we tell at 09:30 whether today will be a trend day (RTH body ratio >= 0.6)?NQNo edgeESNo edge

NQ

No. The same model that reads the range cannot read the day type: BSS −1.2 % (95 % [−2.7 ; 0.2]), AUC 52.4 % and 62.2 % accuracy against a base rate of 36.4 % — where simply always saying "no trend day" already gets 63.6 %, over 905 walk-forward days. The model calls p >= 0.7 on 0.0 % of the days: it never has anything to say. Whether today trends is decided after 09:30.

Full report →

ES

The same on ES, if anything weaker: BSS −1.9 % (95 % [−3.6 ; −0.3]), AUC 50.3 % — a coin flip — and 61.6 % accuracy against a base rate of 36.2 %, below the 63.8 % that the majority baseline gets for free, over 904 walk-forward days. Overnight range, gap and volatility history know nothing about the body ratio of the coming RTH.

Full report →

rv4Can we tell at 09:30 whether the RTH closes above its open?NQNo edgeESNo edge

NQ

No. On the 15 preregistered opening features — gap, position in the value area, overnight delta — the model lands at BSS −0.8 % (95 % [−1.5 ; −0.0]), AUC 47.4 % and 52.3 % accuracy against a base rate of 54.4 %, over 916 walk-forward days. Always going long would have been better than the model at 54.4 %. The direction of the RTH is not readable at the open; the range is.

Full report →

ES

The same on ES and a little worse: BSS −1.3 % (95 % [−2.2 ; −0.6]), AUC 46.7 % — below a coin flip — and 49.2 % accuracy against a base rate of 53.5 %, over 916 walk-forward days. The AUC sitting under 50 in both instruments is what a model without information looks like when it overfits its training window, not a signal to trade inverted.

Full report →

vx1Does implied volatility add what the ticks do not have?Planned

VIX and SPX metrics per trading day — change, percentile, term structure, the premium over realised volatility, divergence days — computed strictly causally and then put into the machines above, both as a reading axis and as a model feature. The direct successor to the null results in this chapter: everything there came from price alone.

vx2Is the VIX-based expected move calibrated — does ±1σ actually cover 68 % of days?Planned

The expected-move cone in the frontend is calibrated on ~110 sessions of option frames. Daily VIX since 1990 puts the same question on more than 8,000 days: how often does the realised ES/NQ range land inside ±1σ and ±2σ, which factor k makes the cone honest, and is that k a constant or a regime variable? Fitted 2020–2024, tested 2025 onwards, with a permuted-VIX control arm.

Methodology — Testing Against Chance

A hit rate near 50 % says nothing on its own — geometry, drift and time of day already move it for an entry that carries no information at all. Every study therefore runs against a Control Arm: the same anchors, the same direction mix, the same session, only the link to the price is cut. What counts is the Δ to that baseline, and PX0 is the proof that the machine reports nothing where there is nothing.

px0Does a signal with no market information show an edge?NQControl passedESControl passed

NQ

Over the full history — 1,467 sessions, 79,778 anchors — a seeded random direction hits 50.3 % against a 50.1 % baseline (Δ +0.15 %) and stays inside the Control Arm band in all 7 hour buckets. Finding nothing is the pass condition here: PX0 calibrates the measuring machine, it claims nothing about the market.

Full report →

ES

Same run on ES, on its own independent random series: 50.0 % against a 49.9 % baseline (Δ +0.04 %, 1,467 sessions, 71,228 anchors), inside the Control Arm band in all 7 hour buckets. The calibration passes on both instruments, which is what every other verdict on this site rests on.

Full report →