Xiphers Research·← Back to indexDEEN

ST1 — state net gex: regime threshold and direction · NQ

NQ · dev 2026-03-16 to 2026-08-13 · 105 sessions · 8 control arms

Verdict — NQ

This report judges NQ only. Every thesis gets two separate verdicts, one per instrument, each on its own signal and with its own arbiter width (P(+160 before −160) here). The other instrument sits in a report of its own; neither decides anything about the other.

Slot A — Dev passed on θ = 1500 (displayed 1,5) — forward val running, verdict on 2026-09-28

Slot A is the user's live setting from thesis.md §4 — the only outcome-independent threshold in the grid: it was the default in the code before the run. The choice rule originally intended (largest ST1a separation) is void because its input failed: ST1a does not separate. Frozen on 2026-08-15.

Slot B — strong, unconfirmed on θ = 2000 (displayed 2) — frozen for forward val, verdict on 2026-09-28

Slot B is the mechanical plateau pick from the dev grid: the smallest threshold of the contiguous range in which Δ > 0, z ≥ 2 and the direct neighbours show the same sign. That is θ = 2,000 — not θ = 4,000/5,000, where Δ is largest. That is exactly the purpose of the rule: smallest threshold means largest n, and a verdict is never passed by maximum. The two instruments land on the same threshold independently of each other. Frozen on 2026-08-15.

Val is pending (forward, from ≥ 30 new trading days). The val data role does not exist for options theses yet: after the addendum of 2026-08-15, val is the window of the first 30 trading days that accumulate after the freeze of this thesis (2026-08-15) — data that did not exist at freeze time. This run computes dev only; the box therefore only shows whether the dev criterion is met. An overall verdict of passed/failed falls only with the one val run.

ST1b, per instrument separately and on the 0DTE variant only: dev — the predicted sign in at least half of the hour buckets; val — one run, the same sign New York total. ST1a (NQ, registered): dev — the difference with the predicted sign in at least half of the hour buckets; val — the same sign New York total. Val does not exist for either part yet (forward, from ≥ 30 new trading days after the freeze).

Slot · configRoleSessionsnP(+160 before −160)BaselineΔ %Control spanBuckets with the signResult
Slot A · θ = 1500 (displayed 1,5)dev105433353.2 %50.3 %+2.8345.8 … 51.8 % ◆7 of 7Dev criterion met
Slot A · θ = 1500 (displayed 1,5)valpending (forward, from ≥ 30 new trading days, verdict on 2026-09-28)
Slot B · θ = 2000 (displayed 2)dev105373254.0 %50.3 %+3.6946.8 … 51.9 % ◆7 of 7Dev criterion met
Slot B · θ = 2000 (displayed 2)valpending (forward, from ≥ 30 new trading days, verdict on 2026-09-28)

Two slots, two verdicts, no race. Slot A and slot B are frozen separately and judged separately — the better one does not overwrite the worse one, and there are never more than two (multiplicity protection, thesis.md §2, rule 9). A “⟂ edge” behind the config means: the choice sits at the edge of the grid, the edge could lie beyond it.

Promotion box

An exploratory finding may only stand as strong, unconfirmed if it meets all four criteria from thesis.md §2, rule 10 — it is then registered the same day as a thesis of its own with a freeze date, and the forward clock starts. Whatever fails them stays exploration with a note. The criteria are recomputed mechanically here, not asserted.

Rule 10 — ST1b plateau, slot B (θ = 2000 (displayed 2))

The monotone rise of Δ across the ST1b grid is the most striking finding of the run and was not preregistered — slot A (θ = 1,500) does not test it. Rule 10 turns it into a thesis of its own as soon as all four criteria are mechanically met; a forward clock of its own then runs from the freeze.

strong, unconfirmed on ST1b plateau (thesis.md §4) — frozen for forward val, verdict on 2026-09-28

Finding: real 54.0 % · baseline 50.3 % · +3.69 % · n 3732 · carrying days 105. Freeze 2026-08-15 · verdict on 2026-09-28 · ST1b plateau (thesis.md §4)

Criterionrequiredmeasuredmet
1 — neighbour consistencygrid neighbours with the same sign, no lone spike2 direct grid neighbours checkedyes
2 — Breadth≥ 4 of 7 hour buckets with the predicted sign, or one contiguous block of ≥ 3 hours7 of 7 buckets, contiguous 09:30–10:00 to 15–16 · scope: New York totalyes
3 — z (entry ticket, not a verdict)z ≥ 2, New York totalz = +2.48yes
4 — carrying days≥ 30 distinct trading days105 daysyes

How the “verdict on” date is estimated: 30 trading days are six trading weeks, so freeze + 42 calendar days, then rolled to the next weekday. Holidays are not subtracted — the date is an estimate, and the run happens once 30 genuinely new trading days sit in the store.

Preregistration
ClaimThe user's live working hypothesis: |net gex| below ~1–2 (displayed; raw 1000–2000) means range-bound, above it trend and the sign gives the direction. Two separately testable parts — ST1a regime, ST1b direction.
Conditionsum_gex_vol from the state stream (classified order flow; state has no zero_gamma and no OI fields). ST1b: only anchors with |sum_gex_vol| ≥ θ, the direction is the sign — below that flat. ST1a: |sum_gex_vol| < θ against ≥ θ, the target quantity is the 30-min forward range. Warm-up frames (value 0) carry neither a sign nor a regime side. Per anchor the last frame with a feed timestamp ≤ anchor applies; a frame after the anchor is never read. Both instruments read the same SPX series — the declared cross-instrument exception from §2, rule 1.
PredictionST1b: hit rate P(+X before −X) above the baseline in the signal direction, with the predicted sign in the majority of the hour buckets. ST1a: a smaller 30-min forward range and lower breakout persistence below θ than above it, so the difference negative.
Target metric (arbiter)P(+160 before −160) — For ST1b the default per instrument: hit rate P(+X before −X) minus baseline at the arbiter width (NQ ±160 T, ES ±40 T), New York total, gross and over resolved anchors only, computed at the frozen threshold. For ST1a the declared deviation: difference of the median 30-min forward range (NQ, ticks), below minus above θ — a volatility quantity, therefore not signed. Both parts are judged separately and both reported.
Baseline (control arm)ST1b: the same anchors, the sign series shifted cyclically across the days. ST1a: the same anchors and the same ranges, but the θ assignment shifted cyclically across the days — the day keeps its movement and loses only the assignment to its net gex. Plus the hour breakdown in both parts, because net gex and time of day are confounded.
Success criterionST1b, per instrument separately and on the 0DTE variant only: dev — the predicted sign in at least half of the hour buckets; val — one run, the same sign New York total. ST1a (NQ, registered): dev — the difference with the predicted sign in at least half of the hour buckets; val — the same sign New York total. Val does not exist for either part yet (forward, from ≥ 30 new trading days after the freeze).
SignalSign of sum_gex_vol from feed ES_SPX state (0DTE), only anchors with |sum_gex_vol| ≥ 1500 — 107 trading days, 2315895 frames, 0.2 % of them warm-up (value 0, flat) — read on the options feed ES_SPX state (0DTE) — the same sign series for NQ and ES, the declared cross-instrument exception (thesis.md §2, rule 1). No futures ticks enter the condition, no SPX → futures basis mapping
Variant0DTE (gex_zero) — primary variant, this is where the verdict falls — The variant the user trades off; declared the primary variant in advance (thesis.md §1).
Threshold gridST1b — θ grid on state — ST1b across the whole θ grid: only anchors with |sum_gex_vol| ≥ θ, the direction is the sign. All seven registered thresholds are named in advance and are reported in full; θ = 0 (sign only, without a magnitude filter) runs along as a baseline column. The verdict above falls at the frozen threshold (★). The one threshold on which the verdict falls is chosen in the dev window by a rule fixed before the run: the threshold with the largest regime separation in ST1a (difference of the median 30-min forward range, NQ) whose direct neighbours point in the same direction — a threshold that jumps out alone between inconspicuous neighbours is noise per Learnings 10 and is not chosen. The chosen threshold is marked with ★, stands with its reason in the measurement and is frozen for the forward val. Addendum 2026-08-15: this choice rule is void because its input failed — ST1a does not separate (difference with the wrong sign, and the control difference reproduces it almost entirely: 96 T against 94 T on NQ). A rule that maximises on noise picks noise. What is frozen is therefore the only threshold that does not come out of the numbers: θ = 1,500 (displayed 1.5), the user's live threshold from thesis.md §4, which was in the code before the run. That Δ in ST1b rises monotonically across the grid up to θ = 4,000/5,000 is thereby explicitly dev exploration and carries no verdict. Thresholds named in advance: θ = 0 (sign only) · θ = 500 (displayed 0,5) · θ = 1000 (displayed 1) · θ = 1500 (displayed 1,5) · θ = 2000 (displayed 2) · θ = 3000 (displayed 3) · θ = 4000 (displayed 4) · θ = 5000 (displayed 5).
Reference instead of a duplicate tableThe classic grid lives in the CL3b report — Addendum 2026-08-15, at the user's instruction: the classic grid stood here a second time as mandatory by-catch — the same computation CL3b already shows in full in its own report. Shown twice means read twice, so only the reference remains here. The question stays the same: does the classification engine (classic) carry more than the raw volume (state)? In a direct comparison the window difference has to be kept in mind — ST1 computes 105 sessions, CL3b 106. ../cl3b-net-gex-sign/
Regime partST1a — regime: is there less range below θ? — The second registered question of this ticket, and the only one with a different target quantity: below θ the market should be calmer. What is measured is the 30-min forward range per anchor in ticks, what is compared are the medians of the two groups, and the target quantity is their difference (below minus above). A negative sign is predicted. Two zero points: the cyclically shifted θ assignment and the hour table below it. Target metric: difference of the median 30-min forward range, |net gex| < θ minus |net gex| ≥ θ. A volatility metric, unsigned — a declared departure from the default.
ValVal is pending (forward, from ≥ 30 new trading days). The val data role does not exist for options theses yet: after the addendum of 2026-08-15, val is the window of the first 30 trading days that accumulate after the freeze of this thesis (2026-08-15) — data that did not exist at freeze time. This run computes dev only; the box therefore only shows whether the dev criterion is met. An overall verdict of passed/failed falls only with the one val run.
Data rolesdev: 2026-03-16 to 2026-08-13
Holdoutfrom 2026-08-14 — not computed in this run, not looked at

NQ · 0DTE (gex_zero) — primary variant, this is where the verdict falls

Primary variant — The variant the user trades off; declared the primary variant in advance (thesis.md §1).

Sign of sum_gex_vol from feed ES_SPX state (0DTE), only anchors with |sum_gex_vol| ≥ 1500 — 107 trading days, 2315895 frames, 0.2 % of them warm-up (value 0, flat) · brackets 40 / 80 / 120 / 160 / 200 / 280 / 320 / 400 / 560 ticks · fill phase spread 1.85/1.62/1.14/1.50

NQ · Dev — full state coverage, digging is allowed here

Arbiter of this role: +2.83 % away from the baseline, predicted sign in 7 of 7 hour buckets.

105 sessions, 2026-03-16 to 2026-08-13 · 4832 anchors: 2947 long (61.0 %) / 1885 short / 3358 flat · 0 without ticks · 0 horizons running past the session end · 8 control arms, shifts +78, +14, +24, +101, +65, +98, +36, +32

Hit rate probe: is the direction right? · the arbiter

HourP(+160 before −160)BaselineΔ %Control spanzresolvednDistribution of the daily hit ratesDays above baseline
09:30–10:0054.9 %43.1 %+11.7637.3 … 54.9 %+1.6100.0 %5116 of 23 (69.6 %)
10–1151.6 %50.7 %+0.8446.4 … 53.0 %+0.799.3 %41533 of 69 (47.8 %)
11–1255.5 %48.5 %+7.0043.9 … 52.7 % ◆+2.6 ●96.9 %70848 of 89 (53.9 %)
12–1350.6 %50.4 %+0.2344.0 … 53.7 %+0.491.3 %67835 of 90 (38.9 %)
13–1452.9 %51.4 %+1.5147.1 … 56.9 %+0.484.5 %69839 of 92 (42.4 %)
14–1556.1 %49.8 %+6.2844.5 … 53.8 % ◆+2.4 ●85.6 %83655 of 100 (55.0 %)
15–1651.5 %51.0 %+0.5343.3 … 54.2 %+0.587.2 %94746 of 104 (44.2 %)
New York total53.2 %50.3 %+2.8345.8 … 51.8 % ◆+1.989.7 %433358 of 105 (55.2 %)
Other bracket widths
HourΔ ±40 TΔ ±80 TΔ ±120 TΔ ±200 TΔ ±280 TΔ ±320 TΔ ±400 TΔ ±560 T
09:30–10:00+3.92+3.92+11.76+17.65+14.17+8.56+21.14+15.62
10–11+0.72-3.35 ●-2.63+3.59-0.13+0.57-4.07-0.16
11–12+2.60+4.24 ●+3.98+6.42+0.27-1.54-4.77-9.26
12–13+2.02 ●+1.21+1.11-0.33+4.96+7.69+4.65+5.53
13–14+3.75 ●+2.31-0.88-1.24-6.01-9.53-4.66-7.05
14–15+2.76 ●+1.39+3.00+9.58 ●+13.38 ●+11.76+5.82-7.18
15–16+5.71 ●+2.86+1.17-0.12+0.88+0.13-3.51-5.74
New York total+3.25 ●+1.69+1.82+3.22+3.09+2.10+0.95-1.00

Breadth (7 hours × 9 bracket widths): 69.8 %, control arms 20.6 % to 61.9 %.

Bracket up to the close (16:00 New York) — context

The same calculation as above, except the bracket does not run for 60 minutes but until 16:00 New York — the real arm and all 8 control arms under the same cap. Context, not a verdict: the arbiter stays the 60-min row above.

HourP(+160 before −160)BaselineΔ %Control spanzresolvedn Δ ±40 TΔ ±80 TΔ ±120 TΔ ±200 TΔ ±280 TΔ ±320 TΔ ±400 TΔ ±560 T
09:30–10:0054.9 %43.1 %+11.7637.3 … 54.9 %+1.6100.0 %51+3.92+3.92+11.76+17.65+13.73+12.78+20.90+16.90
10–1151.7 %50.7 %+0.9646.4 … 52.9 %+0.8100.0 %418+0.72-3.35 ●-2.63+3.83+1.43+1.93-1.89+1.28
11–1254.7 %47.7 %+6.9844.0 … 52.9 % ◆+2.3 ●100.0 %731+2.60+4.24 ●+4.24+4.79+2.20+1.48+0.87-1.92
12–1349.7 %49.6 %+0.1346.0 … 52.1 %+0.199.3 %738+2.02 ●+1.21+0.54-0.90+2.28+4.67+2.37+2.34
13–1451.6 %51.0 %+0.5647.6 … 55.6 %+0.198.3 %812+3.75 ●+2.42-1.39-0.56-0.56-1.26+3.43-1.12
14–1555.5 %49.4 %+6.1445.0 … 54.9 % ◆+2.0 ●94.1 %919+2.76 ●+1.13+2.65+9.56 ●+13.00 ●+13.65 ●+12.51 ●-0.21
15–1651.9 %50.5 %+1.3243.7 … 54.8 %+0.576.5 %831+5.71 ●+2.78+2.48-0.49-2.58-2.64-1.39+3.07
New York total52.7 %49.9 %+2.8046.2 … 51.9 % ◆+1.793.1 %4500+3.25 ●+1.75+1.47+2.47+3.55+3.46+3.85+3.03

The resolved column is the interesting one here: it rises against 60 minutes because the slow path is given time. Only the last row runs the other way — the 15:55 anchor has 5 minutes left until the close, so the 15–16 bucket is capped shorter than above, not longer.

ST1a — regime: is there less range below θ?

The second registered question of this ticket, and the only one with a different target quantity: below θ the market should be calmer. What is measured is the 30-min forward range per anchor in ticks, what is compared are the medians of the two groups, and the target quantity is their difference (below minus above). A negative sign is predicted. Two zero points: the cyclically shifted θ assignment and the hour table below it.

Adjacent thresholds share their signals (Learnings 10): an anchor with |net gex| = 4,000 sits in every threshold below it. The columns are therefore not independent tests — a single threshold that jumps out of the row while its neighbours show nothing is noise, not a finding. What is read is the course across the row; the verdict falls on exactly one, chosen in dev and frozen for the forward val (★).

Artefact: Time of day (the control arm reproduces +94 of +96 T on NQ, θ = 1,500) — the difference points against the prediction, and the cyclically shifted θ assignment produces it almost entirely a second time. |net gex| is small in the morning and grows over the day, and in the morning the forward range is large; within one hour the difference shrinks to −9 to +23 T. No regime label, no signal — difference +96.0 T at the frozen threshold θ = 1500 (displayed 1,5), predicted sign in 2 of 7 hour buckets. The difference points the opposite way from the prediction.

Thresholdn belown aboveMedian belowMedian aboveDifference (T)Control differenceControl spanzBreakout persistence below / aboveBuckets with the sign
θ = 500 (displayed 0,5)12756790450.0339.0+111.0+99.0+79.00 … +129.00+0.677.1 % / 79.8 %3 of 7
θ = 1000 (displayed 1)23195746429.0328.0+101.0+99.0+70.00 … +109.00+0.678.5 % / 79.7 %2 of 7
θ = 1500 (displayed 1,5) ★32334832415.0319.0+96.0+94.0+67.00 … +108.00+0.778.4 % / 80.0 %2 of 7
θ = 2000 (displayed 2)38894176409.0309.0+100.0+89.0+60.00 … +106.00+1.278.7 % / 79.9 %1 of 7
θ = 3000 (displayed 3)48843181398.0295.0+103.0+80.0+51.00 … +102.00+1.678.9 % / 80.1 %0 of 7
θ = 4000 (displayed 4)55952470389.0286.0+103.0+75.0+45.00 … +86.00+2.3 ●79.0 % / 80.1 %1 of 7
θ = 5000 (displayed 5)61301935381.0281.0+100.0+72.0+35.00 … +84.00+2.2 ●79.0 % / 80.3 %0 of 7
The same difference per hour — the time-of-day control

net gex and time of day are confounded — the midday hours are quiet anyway, and net gex accumulates over the day. A single total would therefore measure the clock; only this breakdown shows whether below θ there is less range in every hour.

Hourθ = 500 (displayed 0,5)θ = 1000 (displayed 1)θ = 1500 (displayed 1,5)θ = 2000 (displayed 2)θ = 3000 (displayed 3)θ = 4000 (displayed 4)θ = 5000 (displayed 5)
09:30–10:00+58.0+47.0+85.0+61.0+165.0-14.0+162.0
10–11+101.0 ●+73.0 ●+94.0 ●+104.0 ●+155.0 ●+173.0 ●+170.0
11–12-30.0-17.0-3.0+24.0+64.0+101.0 ●+117.0 ●
12–13-2.0+14.0+18.0+32.0+39.0+54.0+76.0 ●
13–14+44.0+37.0+22.0+17.0+3.0+10.0+12.0
14–15+28.0+34.0 ●+23.0+12.0+25.0+30.0+29.0
15–16-32.0-2.0-9.0-2.0+15.0+6.0+6.0
New York total+111.0+101.0+96.0+100.0+103.0+103.0 ●+100.0 ●

ST1b — θ grid on state

ST1b across the whole θ grid: only anchors with |sum_gex_vol| ≥ θ, the direction is the sign. All seven registered thresholds are named in advance and are reported in full; θ = 0 (sign only, without a magnitude filter) runs along as a baseline column. The verdict above falls at the frozen threshold (★). The one threshold on which the verdict falls is chosen in the dev window by a rule fixed before the run: the threshold with the largest regime separation in ST1a (difference of the median 30-min forward range, NQ) whose direct neighbours point in the same direction — a threshold that jumps out alone between inconspicuous neighbours is noise per Learnings 10 and is not chosen. The chosen threshold is marked with ★, stands with its reason in the measurement and is frozen for the forward val. Addendum 2026-08-15: this choice rule is void because its input failed — ST1a does not separate (difference with the wrong sign, and the control difference reproduces it almost entirely: 96 T against 94 T on NQ). A rule that maximises on noise picks noise. What is frozen is therefore the only threshold that does not come out of the numbers: θ = 1,500 (displayed 1.5), the user's live threshold from thesis.md §4, which was in the code before the run. That Δ in ST1b rises monotonically across the grid up to θ = 4,000/5,000 is thereby explicitly dev exploration and carries no verdict.

Adjacent thresholds share their signals (Learnings 10): an anchor with |net gex| = 4,000 sits in every threshold below it. The columns are therefore not independent tests — a single threshold that jumps out of the row while its neighbours show nothing is noise, not a finding. What is read is the course across the row; the verdict falls on exactly one, chosen in dev and frozen for the forward val (★).

Threshold θ on |sum_gex_vol|AnchorsP(+160 before −160)BaselineΔ %Control spanzresolvedncarrying daysBuckets with the sign
θ = 0 (sign only)806551.5 %50.0 %+1.4746.5 … 51.7 %+1.392.1 %74251055 of 7
θ = 500 (displayed 0,5)679051.8 %49.8 %+1.9446.3 … 51.3 % ◆+1.591.1 %61871056 of 7
θ = 1000 (displayed 1)574652.4 %50.3 %+2.1346.1 … 51.4 % ◆+1.790.1 %51801056 of 7
θ = 1500 (displayed 1,5) ★ [Slot A]483253.2 %50.3 %+2.8345.8 … 51.8 % ◆+1.989.7 %43331057 of 7
θ = 2000 (displayed 2) [Slot B]417654.0 %50.3 %+3.6946.8 … 51.9 % ◆+2.5 ●89.4 %37321057 of 7
θ = 3000 (displayed 3)318155.5 %49.9 %+5.6246.7 … 52.3 % ◆+3.0 ●88.6 %28191057 of 7
θ = 4000 (displayed 4)247056.7 %48.9 %+7.8046.5 … 51.6 % ◆+4.3 ●88.0 %21731006 of 7
θ = 5000 (displayed 5)193557.1 %49.4 %+7.6846.3 … 50.4 % ◆+5.4 ●87.7 %1697916 of 7

The plateau rule picks the smallest threshold, not the largest number (thesis.md §2, rule 9): what is sought is the contiguous range in which Δ > 0, z ≥ 2 and the direct neighbours show the same sign — of that range the smallest threshold is taken, because it has the largest n and is thus the most conservative choice. A verdict is never passed by maximum or by z ranking. A “⟂ edge” behind a slot mark means: the choice sits at the edge of the grid, the edge could lie beyond it.

The same table per hour — Δ % per threshold
Hourθ = 0 (sign only)θ = 500 (displayed 0,5)θ = 1000 (displayed 1)θ = 1500 (displayed 1,5)θ = 2000 (displayed 2)θ = 3000 (displayed 3)θ = 4000 (displayed 4)θ = 5000 (displayed 5)
09:30–10:00+1.77+3.98+10.68+11.76+3.45+10.00-66.67-100.00
10–11-0.56+0.60+0.47+0.84+0.99+6.45+10.28 ●+10.61 ●
11–12+3.47+4.62 ●+6.64 ●+7.00 ●+7.07 ●+11.71 ●+16.67 ●+14.65 ●
12–13+2.14+1.69+0.65+0.23+2.76+4.82+10.96 ●+16.25
13–14-3.49-2.97-1.37+1.51+1.12+2.58+3.90+1.35
14–15+6.23 ●+5.56+4.78+6.28 ●+6.73 ●+8.20 ●+7.46 ●+8.65 ●
15–16+0.63+0.19+0.15+0.53+1.71+1.52+4.69+5.29
New York total+1.47+1.94+2.13+2.83+3.69 ●+5.62 ●+7.80 ●+7.68 ●

The classic grid lives in the CL3b report

Addendum 2026-08-15, at the user's instruction: the classic grid stood here a second time as mandatory by-catch — the same computation CL3b already shows in full in its own report. Shown twice means read twice, so only the reference remains here. The question stays the same: does the classification engine (classic) carry more than the raw volume (state)? In a direct comparison the window difference has to be kept in mind — ST1 computes 105 sessions, CL3b 106. Open the CL3b report

Year by year: when did the effect live?

The same metric as above, only read by calendar year — reading cuts through the same 8 control arms, edge years labelled with their span; z is screening, not a verdict.

YearP(+160 before −160)BaselineΔ %Control spanzresolvednSessions
2026 (16.03.–13.08.)53.2 %50.3 %+2.8345.8 … 51.8 % ◆+1.989.7 %4333105
+3.23+18.7+2.32+13.5+1.41+8.2+0.51+3.0-0.40-2.3Δ %z2026 (16.03.–13.08.) · Δ +2.83 % · z +1.88 · n 4333z +1.882026 ** Teiljahr

Bars = Δ % per year against the baseline (left axis), ◆ = z against the control arms of the same year (right axis), filled from |z| ≥ 2. * marks a partial year.

Time probe: how many ticks were there to take?

Hour1 min5 min15 min30 min60 minnØ z
09:30–10:00+3.33+55.61 ●+34.24-44.27+129.2951+1.0
10–11-2.51-4.89+2.28+10.72+76.54418+0.2
11–12+1.73+3.75+18.82+13.48-5.64731+0.6
12–13+4.00-0.04-2.45+15.42+20.55743+0.8
13–14+4.27+0.05-8.46-21.65-44.86826-0.3
14–15+3.96+0.50+2.81+21.02+39.40977+1.1
15–16+2.56+3.28+2.76-13.16-46.811086+0.6
New York total+2.95+0.97+2.42+5.89+13.954832+1.5
All horizons per hour, with MFE and MAE
HourHorizonnØ FRBaselineΔControl spanShare positiveØ MFEØ MAE
09:30–10:001 min51-1.00-4.33+3.33-36.73 … +13.9243.1 %+72.47-71.80
09:30–10:005 min51+45.02-10.59+55.61-49.92 … +11.59 ◆58.8 %+155.10-140.84
09:30–10:0015 min51+44.29+10.06+34.24-176.04 … +85.9266.7 %+260.22-236.51
09:30–10:0030 min51-18.35+25.92-44.27-193.78 … +164.2562.7 %+337.86-355.16
09:30–10:0060 min51+110.35-18.94+129.29-221.33 … +162.6168.6 %+422.75-402.27
10–111 min418-3.69-1.18-2.51-5.87 … +6.6848.3 %+50.66-57.14
10–115 min418-6.06-1.16-4.89-11.56 … +7.9249.3 %+106.40-119.33
10–1115 min418+3.27+0.99+2.28-27.28 … +25.0452.4 %+179.66-191.43
10–1130 min418+15.00+4.28+10.72-46.53 … +32.8656.2 %+243.30-255.28
10–1160 min418+66.59-9.95+76.54-58.83 … +72.7757.2 %+344.63-323.49
11–121 min731+2.01+0.28+1.73-4.47 … +4.0451.6 %+46.65-44.94
11–125 min731+5.01+1.26+3.75-9.20 … +9.0152.1 %+94.76-91.11
11–1215 min731+19.98+1.15+18.82-29.60 … +25.3356.0 %+166.44-151.32
11–1230 min731+27.78+14.30+13.48-53.88 … +44.4256.6 %+234.89-204.63
11–1260 min731+20.48+26.12-5.64-97.71 … +99.6752.7 %+320.53-291.19
12–131 min743+3.10-0.91+4.00-5.09 … +2.49 ◆52.6 %+38.76-35.09
12–135 min743-0.61-0.57-0.04-7.08 … +6.1348.5 %+76.41-75.05
12–1315 min743-1.17+1.28-2.45-20.51 … +16.0451.1 %+132.69-134.68
12–1330 min743+10.25-5.17+15.42-32.66 … +24.6853.3 %+188.99-186.77
12–1360 min743+25.31+4.75+20.55-45.38 … +11.34 ◆56.4 %+265.06-245.87
13–141 min826+4.65+0.38+4.27-3.32 … +3.43 ◆54.2 %+37.49-30.31
13–145 min826+0.79+0.74+0.05-3.65 … +4.9050.2 %+73.19-69.16
13–1415 min826-5.82+2.64-8.46-8.49 … +12.4951.7 %+121.58-122.84
13–1430 min826-14.98+6.66-21.65-20.69 … +27.1649.6 %+164.37-174.00
13–1460 min826-27.71+17.14-44.86-44.35 … +32.9846.4 %+217.97-238.97
14–151 min977+3.38-0.58+3.96-3.10 … +3.26 ◆52.4 %+34.04-29.37
14–155 min977-1.03-1.53+0.50-7.57 … +3.7948.9 %+64.98-63.52
14–1515 min977+0.42-2.38+2.81-14.46 … +5.4652.4 %+112.38-112.10
14–1530 min977+12.65-8.36+21.02-21.25 … +20.7853.1 %+163.86-158.98
14–1560 min977+33.25-6.15+39.40-45.57 … +32.63 ◆54.6 %+248.01-217.72
15–161 min1086+3.90+1.34+2.56-3.15 … +1.75 ◆53.3 %+42.00-35.89
15–165 min1086+5.63+2.36+3.28-6.76 … +5.33 ◆52.8 %+80.30-70.31
15–1615 min1086+6.49+3.73+2.76-19.71 … +12.8750.8 %+142.21-130.44
15–1630 min1086-1.13+12.03-13.16-32.89 … +19.3648.2 %+199.56-192.76
15–1660 min1086-17.64+29.17-46.81-64.97 … +42.5149.3 %+263.11-277.78
New York total1 min4832+2.81-0.15+2.95-2.06 … +0.49 ◆52.4 %+40.89-37.08
New York total5 min4832+1.81+0.83+0.97-4.72 … +2.0850.6 %+80.62-77.60
New York total15 min4832+4.14+1.72+2.42-10.56 … +6.8052.4 %+139.34-135.64
New York total30 min4832+6.62+0.73+5.89-16.30 … +13.1152.3 %+195.29-190.72
New York total60 min4832+11.93-2.01+13.95-20.64 … +27.6752.3 %+270.06-261.39

Breadth (7 hours × 5 horizons): 68.6 %, control arms 14.3 % to 60.0 %.

Both halves of the window, separately
HalfBracketn resolvedP(+X before −X)BaselineΔ %
1 (2026-03-16 to 2026-05-29)±40 T234553.4 %50.3 %+3.17
1 (2026-03-16 to 2026-05-29)±80 T233851.9 %50.3 %+1.61
1 (2026-03-16 to 2026-05-29)±120 T223552.4 %50.8 %+1.60
1 (2026-03-16 to 2026-05-29)±160 T203753.7 %50.7 %+2.95
1 (2026-03-16 to 2026-05-29)±200 T178454.3 %50.6 %+3.74
1 (2026-03-16 to 2026-05-29)±280 T134651.6 %49.6 %+1.93
1 (2026-03-16 to 2026-05-29)±320 T113950.8 %49.1 %+1.78
1 (2026-03-16 to 2026-05-29)±400 T75249.5 %47.0 %+2.47
1 (2026-03-16 to 2026-05-29)±560 T30746.3 %50.7 %-4.44
2 (2026-06-01 to 2026-08-13)±40 T248752.6 %49.3 %+3.31
2 (2026-06-01 to 2026-08-13)±80 T248251.1 %49.2 %+1.93
2 (2026-06-01 to 2026-08-13)±120 T243051.1 %50.0 %+1.07
2 (2026-06-01 to 2026-08-13)±160 T229652.7 %49.7 %+3.02
2 (2026-06-01 to 2026-08-13)±200 T209752.2 %49.6 %+2.57
2 (2026-06-01 to 2026-08-13)±280 T169051.4 %49.1 %+2.30
2 (2026-06-01 to 2026-08-13)±320 T150849.8 %48.3 %+1.51
2 (2026-06-01 to 2026-08-13)±400 T115048.3 %49.2 %-0.90
2 (2026-06-01 to 2026-08-13)±560 T66949.2 %48.0 %+1.17
HalfHorizonnØ FRBaselineΔ (ticks)
1 (2026-03-16 to 2026-05-29)1 min2345+1.75-0.02+1.77
1 (2026-03-16 to 2026-05-29)5 min2345+3.23+0.76+2.47
1 (2026-03-16 to 2026-05-29)15 min2345+3.36+1.82+1.54
1 (2026-03-16 to 2026-05-29)30 min2345+4.44+1.11+3.33
1 (2026-03-16 to 2026-05-29)60 min2345+14.25+2.42+11.83
2 (2026-06-01 to 2026-08-13)1 min2487+3.80-0.22+4.02
2 (2026-06-01 to 2026-08-13)5 min2487+0.46+0.01+0.46
2 (2026-06-01 to 2026-08-13)15 min2487+4.88-0.37+5.25
2 (2026-06-01 to 2026-08-13)30 min2487+8.68+0.31+8.38
2 (2026-06-01 to 2026-08-13)60 min2487+9.76-1.07+10.83

NQ · full (gex_full, all expiries ≤ 90 days) — secondary variant

Secondary variant, no verdict — Mandatory secondary variant, the same days and the same control arms, only the feed is a different one. It does not judge: an effect that shows up only here is an exploratory finding and must run as a thesis of its own on untouched data (thesis.md §1).

Sign of sum_gex_vol from feed ES_SPX state (full, all expiries ≤ 90 days), only anchors with |sum_gex_vol| ≥ 1500 — 107 trading days, 2315897 frames, 0.2 % of them warm-up (value 0, flat) · brackets 40 / 80 / 120 / 160 / 200 / 280 / 320 / 400 / 560 ticks · fill phase spread 1.85/1.62/1.14/1.50

NQ · Dev — full state coverage, digging is allowed here

Metric (no verdict) of this role: +3.49 % away from the baseline, predicted sign in 6 of 7 hour buckets.

105 sessions, 2026-03-16 to 2026-08-13 · 5714 anchors: 3443 long (60.3 %) / 2271 short / 2476 flat · 0 without ticks · 0 horizons running past the session end · 8 control arms, shifts +78, +14, +24, +101, +65, +98, +36, +32

Hit rate probe: is the direction right?

HourP(+160 before −160)BaselineΔ %Control spanzresolvednDistribution of the daily hit ratesDays above baseline
09:30–10:0057.3 %46.9 %+10.4236.5 … 56.2 % ◆+1.9100.0 %9628 of 44 (63.6 %)
10–1148.9 %49.6 %-0.7448.8 … 54.3 %-0.899.3 %61052 of 92 (56.5 %)
11–1254.4 %48.8 %+5.5644.5 … 51.3 % ◆+2.9 ●96.9 %87557 of 100 (57.0 %)
12–1353.0 %49.5 %+3.5244.7 … 52.9 % ◆+1.692.2 %83756 of 99 (56.6 %)
13–1451.5 %50.1 %+1.4247.6 … 57.2 %+0.185.3 %84948 of 103 (46.6 %)
14–1556.9 %49.7 %+7.2144.2 … 53.1 % ◆+2.8 ●85.9 %90858 of 104 (55.8 %)
15–1652.8 %51.5 %+1.2942.8 … 53.8 %+0.887.6 %99947 of 104 (45.2 %)
New York total53.2 %49.7 %+3.4946.3 … 52.1 % ◆+2.2 ●90.5 %517460 of 105 (57.1 %)
Other bracket widths
HourΔ ±40 TΔ ±80 TΔ ±120 TΔ ±200 TΔ ±280 TΔ ±320 TΔ ±400 TΔ ±560 T
09:30–10:00+3.12+6.25+9.38 ●+13.54+10.87+8.50+12.54+1.87
10–11-1.79 ●-3.58 ●-2.93+0.26-0.03-0.48-2.65+0.66
11–12+1.33+0.44+3.34+4.55-1.91-2.74-5.22-13.62
12–13+2.64+3.74+3.49+5.43+7.63 ●+7.28+6.58+5.30
13–14+5.23 ●+2.63-0.00+0.54-4.82-7.35-5.51-1.70
14–15+3.22 ●+1.19+3.91+10.65 ●+12.82 ●+12.55 ●+8.96-5.35
15–16+5.52 ●+3.16 ●+2.12-0.11+0.34-0.11-3.70-4.70
New York total+3.06 ●+1.80 ●+2.27+3.71 ●+3.13+2.00+0.11-1.83

Breadth (7 hours × 9 bracket widths): 66.7 %, control arms 17.5 % to 61.9 %.

Bracket up to the close (16:00 New York) — context

The same calculation as above, except the bracket does not run for 60 minutes but until 16:00 New York — the real arm and all 8 control arms under the same cap. Context, not a verdict: the arbiter stays the 60-min row above.

HourP(+160 before −160)BaselineΔ %Control spanzresolvedn Δ ±40 TΔ ±80 TΔ ±120 TΔ ±200 TΔ ±280 TΔ ±320 TΔ ±400 TΔ ±560 T
09:30–10:0057.3 %46.9 %+10.4236.5 … 56.2 % ◆+1.9100.0 %96+3.12+6.25+9.38 ●+13.54+11.46+11.04+13.30 ●+10.84
10–1148.9 %49.7 %-0.8148.9 … 54.4 %-0.8100.0 %614-1.79 ●-3.58 ●-2.93+0.81+1.79+1.65+2.26+7.38
11–1254.2 %49.4 %+4.7644.5 … 51.5 % ◆+2.7 ●100.0 %903+1.33+0.44+3.43+3.21+0.17-0.29+0.37+2.59
12–1352.3 %48.4 %+3.9346.9 … 52.1 % ◆+1.599.3 %902+2.64+3.74+3.30+4.51+6.74+7.28+4.49+6.91
13–1450.6 %51.0 %-0.4646.9 … 57.0 %-0.198.6 %981+5.23 ●+2.51-0.55+0.06+0.17+0.80+3.80+1.03
14–1556.1 %49.3 %+6.8344.6 … 53.9 % ◆+2.6 ●94.1 %995+3.22 ●+0.95+3.36+10.57 ●+14.55 ●+14.21 ●+14.89 ●+5.39
15–1653.1 %51.0 %+2.0843.2 … 54.7 %+0.877.4 %883+5.52 ●+3.27 ●+3.39-0.07-2.18-2.16-2.72+0.39
New York total52.8 %49.4 %+3.4546.7 … 52.3 % ◆+2.094.0 %5374+3.06 ●+1.92 ●+1.97+3.98+4.45+4.40+4.54+5.91

The resolved column is the interesting one here: it rises against 60 minutes because the slow path is given time. Only the last row runs the other way — the 15:55 anchor has 5 minutes left until the close, so the 15–16 bucket is capped shorter than above, not longer.

ST1a — regime: is there less range below θ?

The same computation as in the NQ report, context here: the target quantity of ST1a is registered explicitly on NQ in thesis.md §4. The ES numbers run along in full but carry no dev criterion.

Adjacent thresholds share their signals (Learnings 10): an anchor with |net gex| = 4,000 sits in every threshold below it. The columns are therefore not independent tests — a single threshold that jumps out of the row while its neighbours show nothing is noise, not a finding. What is read is the course across the row; the verdict falls on exactly one, chosen in dev and frozen for the forward val (★).

Thresholdn belown aboveMedian belowMedian aboveDifference (T)Control differenceControl spanzBreakout persistence below / aboveBuckets with the sign
θ = 500 (displayed 0,5)8857180468.0341.0+127.0+112.0+90.00 … +119.00+1.978.8 % / 79.4 %1 of 7
θ = 1000 (displayed 1)16696396448.0333.0+115.0+102.0+90.00 … +118.00+1.178.5 % / 79.6 %1 of 7
θ = 1500 (displayed 1,5)23515714436.0327.0+109.0+99.0+79.00 … +113.00+1.178.4 % / 79.7 %1 of 7
θ = 2000 (displayed 2)29795086422.0320.0+102.0+98.0+76.00 … +109.00+0.878.7 % / 79.7 %1 of 7
θ = 3000 (displayed 3)40524013408.0307.0+101.0+84.0+56.00 … +110.00+1.179.0 % / 79.7 %1 of 7
θ = 4000 (displayed 4)46963369402.0298.0+104.0+79.0+53.00 … +110.00+1.379.0 % / 79.8 %2 of 7
θ = 5000 (displayed 5)52212844393.0291.0+102.0+75.0+48.00 … +103.00+1.679.1 % / 79.8 %3 of 7
The same difference per hour — the time-of-day control

net gex and time of day are confounded — the midday hours are quiet anyway, and net gex accumulates over the day. A single total would therefore measure the clock; only this breakdown shows whether below θ there is less range in every hour.

Hourθ = 500 (displayed 0,5)θ = 1000 (displayed 1)θ = 1500 (displayed 1,5)θ = 2000 (displayed 2)θ = 3000 (displayed 3)θ = 4000 (displayed 4)θ = 5000 (displayed 5)
09:30–10:00+66.0+85.0+75.0+111.0+58.0+162.0-13.0
10–11+108.0 ●+125.0 ●+118.0 ●+102.0 ●+124.0 ●+153.0 ●+169.0 ●
11–12+74.0+2.0+8.0+9.0+21.0+54.0+93.0 ●
12–13+2.0+7.0+12.0+20.0+29.0+43.0+53.0
13–14+34.0+36.0+3.0+3.0+3.0-1.0-1.0
14–15+16.0+32.0 ●+37.0 ●+29.0+28.0+25.0+32.0
15–16-14.0-11.0-18.0-17.0-14.0-12.0-25.0
New York total+127.0+115.0+109.0+102.0+101.0+104.0+102.0

ST1b — θ grid on state

ST1b across the whole θ grid: only anchors with |sum_gex_vol| ≥ θ, the direction is the sign. All seven registered thresholds are named in advance and are reported in full; θ = 0 (sign only, without a magnitude filter) runs along as a baseline column. The verdict above falls at the frozen threshold (★). The one threshold on which the verdict falls is chosen in the dev window by a rule fixed before the run: the threshold with the largest regime separation in ST1a (difference of the median 30-min forward range, NQ) whose direct neighbours point in the same direction — a threshold that jumps out alone between inconspicuous neighbours is noise per Learnings 10 and is not chosen. The chosen threshold is marked with ★, stands with its reason in the measurement and is frozen for the forward val. Addendum 2026-08-15: this choice rule is void because its input failed — ST1a does not separate (difference with the wrong sign, and the control difference reproduces it almost entirely: 96 T against 94 T on NQ). A rule that maximises on noise picks noise. What is frozen is therefore the only threshold that does not come out of the numbers: θ = 1,500 (displayed 1.5), the user's live threshold from thesis.md §4, which was in the code before the run. That Δ in ST1b rises monotonically across the grid up to θ = 4,000/5,000 is thereby explicitly dev exploration and carries no verdict.

Adjacent thresholds share their signals (Learnings 10): an anchor with |net gex| = 4,000 sits in every threshold below it. The columns are therefore not independent tests — a single threshold that jumps out of the row while its neighbours show nothing is noise, not a finding. What is read is the course across the row; the verdict falls on exactly one, chosen in dev and frozen for the forward val (★).

Threshold θ on |sum_gex_vol|AnchorsP(+160 before −160)BaselineΔ %Control spanzresolvedncarrying daysBuckets with the sign
θ = 0 (sign only)806552.6 %49.7 %+2.9547.3 … 51.7 % ◆+2.3 ●92.1 %74251057 of 7
θ = 500 (displayed 0,5)718053.0 %49.6 %+3.4147.2 … 51.8 % ◆+2.6 ●91.5 %65691056 of 7
θ = 1000 (displayed 1)639653.0 %49.8 %+3.2546.5 … 51.9 % ◆+2.3 ●90.9 %58161056 of 7
θ = 1500 (displayed 1,5) ★ [Slot A]571453.2 %49.7 %+3.4946.3 … 52.1 % ◆+2.2 ●90.5 %51741056 of 7
θ = 2000 (displayed 2) [Slot B]508653.5 %49.4 %+4.0946.2 … 52.3 % ◆+2.1 ●90.1 %45821056 of 7
θ = 3000 (displayed 3)401355.2 %49.3 %+5.8647.0 … 52.6 % ◆+2.8 ●89.4 %35861057 of 7
θ = 4000 (displayed 4)336956.4 %49.1 %+7.2547.2 … 52.5 % ◆+3.6 ●88.9 %2994997 of 7
θ = 5000 (displayed 5)284456.6 %49.0 %+7.6747.1 … 52.7 % ◆+3.6 ●88.2 %2507976 of 7

The plateau rule picks the smallest threshold, not the largest number (thesis.md §2, rule 9): what is sought is the contiguous range in which Δ > 0, z ≥ 2 and the direct neighbours show the same sign — of that range the smallest threshold is taken, because it has the largest n and is thus the most conservative choice. A verdict is never passed by maximum or by z ranking. A “⟂ edge” behind a slot mark means: the choice sits at the edge of the grid, the edge could lie beyond it.

The same table per hour — Δ % per threshold
Hourθ = 0 (sign only)θ = 500 (displayed 0,5)θ = 1000 (displayed 1)θ = 1500 (displayed 1,5)θ = 2000 (displayed 2)θ = 3000 (displayed 3)θ = 4000 (displayed 4)θ = 5000 (displayed 5)
09:30–10:00+2.17+6.71 ●+11.46 ●+10.42+5.26+4.35+18.18-75.00
10–11+0.72-0.10+0.00-0.74+0.00+3.45+10.41 ●+7.75 ●
11–12+5.38 ●+5.64 ●+4.77 ●+5.56 ●+5.20 ●+8.32 ●+10.27 ●+14.34 ●
12–13+3.25+3.73+3.57+3.52+2.56+4.00+7.00+8.29
13–14+1.38+1.60+0.11+1.42+1.96+3.36+4.57+3.90
14–15+5.22+6.04 ●+6.78 ●+7.21 ●+8.05 ●+9.66 ●+8.14 ●+8.51 ●
15–16+0.54+1.35+1.31+1.29+1.96+2.74+3.65+4.58
New York total+2.95 ●+3.41 ●+3.25 ●+3.49 ●+4.09 ●+5.86 ●+7.25 ●+7.67 ●

The classic grid lives in the CL3b report

Addendum 2026-08-15, at the user's instruction: the classic grid stood here a second time as mandatory by-catch — the same computation CL3b already shows in full in its own report. Shown twice means read twice, so only the reference remains here. The question stays the same: does the classification engine (classic) carry more than the raw volume (state)? In a direct comparison the window difference has to be kept in mind — ST1 computes 105 sessions, CL3b 106. Open the CL3b report

Year by year: when did the effect live?

The same metric as above, only read by calendar year — reading cuts through the same 8 control arms, edge years labelled with their span; z is screening, not a verdict.

YearP(+160 before −160)BaselineΔ %Control spanzresolvednSessions
2026 (16.03.–13.08.)53.2 %49.7 %+3.4946.3 … 52.1 % ◆+2.2 ●90.5 %5174105
+3.98+18.7+2.86+13.5+1.74+8.2+0.63+3.0-0.49-2.3Δ %z2026 (16.03.–13.08.) · Δ +3.49 % · z +2.18 · n 5174z +2.182026 ** Teiljahr

Bars = Δ % per year against the baseline (left axis), ◆ = z against the control arms of the same year (right axis), filled from |z| ≥ 2. * marks a partial year.

Time probe: how many ticks were there to take?

Hour1 min5 min15 min30 min60 minnØ z
09:30–10:00+10.94+54.07 ●+14.18-67.21-7.2096+0.5
10–11-6.60 ●-10.17 ●-9.06+15.50+65.61 ●614-0.4
11–12+2.52+3.72+5.69-6.22-35.26903+0.6
12–13+3.28-1.30+6.95+25.77+28.66908+1.2
13–14+3.42-0.65-6.36-15.49-25.13995-0.2
14–15+3.53-0.59-0.45+23.28+43.87 ●1057+1.2
15–16+1.87+4.21+3.78-16.71-46.601141+0.4
New York total+2.49+0.98+3.24+5.47+11.655714+1.4
All horizons per hour, with MFE and MAE
HourHorizonnØ FRBaselineΔControl spanShare positiveØ MFEØ MAE
09:30–10:001 min96+7.54-3.40+10.94-26.48 … +16.6150.0 %+77.14-71.16
09:30–10:005 min96+46.64-7.44+54.07-42.08 … +33.50 ◆59.4 %+168.96-138.47
09:30–10:0015 min96+19.22+5.04+14.18-125.64 … +91.1961.5 %+272.56-255.14
09:30–10:0030 min96-64.59+2.61-67.21-134.82 … +75.6757.3 %+337.48-375.97
09:30–10:0060 min96-42.35-35.16-7.20-161.00 … +106.9762.5 %+425.79-507.17
10–111 min614-5.78+0.82-6.60-3.71 … +4.3747.1 %+52.89-60.56
10–115 min614-9.10+1.07-10.17-6.35 … +3.2549.5 %+106.81-124.02
10–1115 min614-9.69-0.63-9.06-27.81 … +12.1750.7 %+176.59-203.02
10–1130 min614+13.29-2.21+15.50-37.62 … +23.8056.2 %+245.33-272.58
10–1160 min614+67.89+2.27+65.61-46.04 … +41.44 ◆58.3 %+353.16-340.82
11–121 min903+3.70+1.18+2.52-5.39 … +2.75 ◆51.7 %+47.48-44.31
11–125 min903+4.48+0.76+3.72-7.94 … +6.9952.6 %+94.30-91.87
11–1215 min903+12.98+7.29+5.69-28.59 … +21.7354.7 %+161.36-156.03
11–1230 min903+10.53+16.75-6.22-49.41 … +38.9154.0 %+224.43-213.49
11–1260 min903+1.25+36.50-35.26-88.58 … +56.1552.2 %+302.26-303.74
12–131 min908+3.30+0.03+3.28-3.31 … +2.85 ◆52.5 %+39.41-35.61
12–135 min908+0.11+1.41-1.30-5.09 … +4.8050.0 %+77.91-75.23
12–1315 min908+3.13-3.82+6.95-16.29 … +7.9752.1 %+135.60-131.51
12–1330 min908+14.68-11.09+25.77-27.41 … +17.7855.9 %+191.44-183.02
12–1360 min908+20.15-8.51+28.66-45.52 … +10.91 ◆58.7 %+266.51-247.97
13–141 min995+4.21+0.79+3.42-2.86 … +3.43 ◆54.1 %+37.44-30.84
13–145 min995+0.64+1.29-0.65-1.79 … +4.5950.8 %+73.00-69.17
13–1415 min995-2.46+3.91-6.36-7.64 … +13.3452.0 %+122.08-123.27
13–1430 min995-8.66+6.83-15.49-15.98 … +31.7051.2 %+169.17-175.56
13–1460 min995-16.15+8.97-25.13-34.91 … +45.7247.8 %+231.87-241.31
14–151 min1057+2.93-0.60+3.53-3.61 … +2.72 ◆52.6 %+33.99-30.23
14–155 min1057-1.00-0.41-0.59-7.94 … +3.3049.9 %+65.88-65.03
14–1515 min1057+0.32+0.76-0.45-15.59 … +6.2653.0 %+114.14-114.81
14–1530 min1057+14.86-8.42+23.28-25.70 … +18.9754.5 %+168.17-162.59
14–1560 min1057+36.69-7.18+43.87-46.10 … +20.29 ◆56.7 %+254.35-224.75
15–161 min1141+2.25+0.38+1.87-3.01 … +2.01 ◆52.1 %+41.23-37.18
15–165 min1141+5.63+1.42+4.21-7.29 … +4.87 ◆53.2 %+80.16-70.97
15–1615 min1141+6.45+2.67+3.78-19.85 … +16.4052.1 %+142.67-131.30
15–1630 min1141-2.62+14.09-16.71-36.74 … +23.9848.9 %+199.09-192.59
15–1660 min1141-23.10+23.50-46.60-68.24 … +39.9449.2 %+261.30-278.13
New York total1 min5714+2.34-0.15+2.49-1.51 … +0.58 ◆52.0 %+41.78-38.75
New York total5 min5714+1.58+0.60+0.98-3.97 … +1.7051.3 %+82.51-80.37
New York total15 min5714+2.75-0.49+3.24-11.42 … +6.1852.6 %+141.46-140.58
New York total30 min5714+5.06-0.41+5.47-19.00 … +10.2953.2 %+198.24-197.53
New York total60 min5714+9.35-2.31+11.65-27.21 … +21.6653.5 %+274.83-271.68

Breadth (7 hours × 5 horizons): 54.3 %, control arms 17.1 % to 54.3 %.

Both halves of the window, separately
HalfBracketn resolvedP(+X before −X)BaselineΔ %
1 (2026-03-16 to 2026-05-29)±40 T283453.3 %49.8 %+3.51
1 (2026-03-16 to 2026-05-29)±80 T282752.2 %50.0 %+2.18
1 (2026-03-16 to 2026-05-29)±120 T270953.2 %50.2 %+2.97
1 (2026-03-16 to 2026-05-29)±160 T250054.4 %50.8 %+3.62
1 (2026-03-16 to 2026-05-29)±200 T220654.9 %51.2 %+3.79
1 (2026-03-16 to 2026-05-29)±280 T164952.5 %50.3 %+2.19
1 (2026-03-16 to 2026-05-29)±320 T138651.4 %50.2 %+1.28
1 (2026-03-16 to 2026-05-29)±400 T91649.2 %47.5 %+1.69
1 (2026-03-16 to 2026-05-29)±560 T36543.0 %50.9 %-7.93
2 (2026-06-01 to 2026-08-13)±40 T288051.8 %49.3 %+2.58
2 (2026-06-01 to 2026-08-13)±80 T287450.5 %49.5 %+0.98
2 (2026-06-01 to 2026-08-13)±120 T281650.6 %49.6 %+1.07
2 (2026-06-01 to 2026-08-13)±160 T267452.1 %48.8 %+3.38
2 (2026-06-01 to 2026-08-13)±200 T246751.8 %48.9 %+2.87
2 (2026-06-01 to 2026-08-13)±280 T201851.3 %47.6 %+3.70
2 (2026-06-01 to 2026-08-13)±320 T181150.0 %47.1 %+2.84
2 (2026-06-01 to 2026-08-13)±400 T141648.9 %47.4 %+1.56
2 (2026-06-01 to 2026-08-13)±560 T84048.7 %49.9 %-1.19
HalfHorizonnØ FRBaselineΔ (ticks)
1 (2026-03-16 to 2026-05-29)1 min2834+1.72-0.01+1.73
1 (2026-03-16 to 2026-05-29)5 min2834+3.17+0.39+2.79
1 (2026-03-16 to 2026-05-29)15 min2834+5.60+0.07+5.53
1 (2026-03-16 to 2026-05-29)30 min2834+10.13-0.22+10.35
1 (2026-03-16 to 2026-05-29)60 min2834+21.99+4.37+17.61
2 (2026-06-01 to 2026-08-13)1 min2880+2.95-0.03+2.97
2 (2026-06-01 to 2026-08-13)5 min2880+0.02-0.76+0.78
2 (2026-06-01 to 2026-08-13)15 min2880-0.06-1.11+1.05
2 (2026-06-01 to 2026-08-13)30 min2880+0.06-4.63+4.69
2 (2026-06-01 to 2026-08-13)60 min2880-3.09-9.76+6.67

Sample days

10 sessions, drawn by seed (20260813), not hand-picked — illustration, not evidence (Learnings 11). They are the same days as in the other instrument's report, so the two can be laid side by side. Price NQ as a 1-minute path in points; there is no indicator line here — the condition lives in a different price space and does not fit on this axis, only the anchors are visible.
Every marker is an anchor: long, short; green = +160 T reached first, red = −160 T first, grey = neither side within 60 minutes.

2026-03-24

78 anchors · 24 long / 54 short · 25 resolved, 5 of them with +160 T first (20.0 %)

09:3010:0011:0012:0013:0014:0015:0016:002412724195242632433124399

2026-04-09

78 anchors · 62 long / 16 short · 57 resolved, 38 of them with +160 T first (66.7 %)

09:3010:0011:0012:0013:0014:0015:0016:002494225028251142520125287

2026-04-30

78 anchors · 40 long / 38 short · 31 resolved, 20 of them with +160 T first (64.5 %)

09:3010:0011:0012:0013:0014:0015:0016:002714027266273932751927645

2026-05-21

78 anchors · 41 long / 37 short · 41 resolved, 23 of them with +160 T first (56.1 %)

09:3010:0011:0012:0013:0014:0015:0016:002912029232293432945429565

2026-06-01

78 anchors · 48 long / 30 short · 48 resolved, 22 of them with +160 T first (45.8 %)

09:3010:0011:0012:0013:0014:0015:0016:003027730386304953060430714

2026-06-12

78 anchors · 15 long / 63 short · 17 resolved, 2 of them with +160 T first (11.8 %)

09:3010:0011:0012:0013:0014:0015:0016:002920829353294972964229786

2026-06-16

78 anchors · 0 long / 78 short · 59 resolved, 45 of them with +160 T first (76.3 %)

09:3010:0011:0012:0013:0014:0015:0016:003027630436305963075730917

2026-07-27

78 anchors · 4 long / 74 short · 61 resolved, 31 of them with +160 T first (50.8 %)

09:3010:0011:0012:0013:0014:0015:0016:002792028100282812846128642

2026-08-04

78 anchors · 72 long / 6 short · 72 resolved, 59 of them with +160 T first (81.9 %)

09:3010:0011:0012:0013:0014:0015:0016:002919229392295922979229993

2026-08-05

78 anchors · 14 long / 64 short · 31 resolved, 12 of them with +160 T first (38.7 %)

09:3010:0011:0012:0013:0014:0015:0016:002958029705298312995730082
How to read this report

Layout

This report belongs to one instrument (NQ) and judges nothing else. First the verdict, then per data role (dev, then val) the same set of sections, always in the same order: hit rate probe including the distribution across days (is the direction right? — the arbiter) · time probe (how many ticks were there to take) · both halves of the window separately. Which section carries the arbiter depends on the preregistration of this study; the order is fixed. Nothing collapsed is second-rate, only second-asked — the other bracket widths, the full MFE/MAE matrix and the halves.

Money is no longer reported here. The probe asks about direction; the tradability section (money per sizing model and the metrics beside it) was cut from this report without replacement on 2026-08-15. Whether a direction can be traded after costs is checked later by a test of its own — once a thesis holds up.

Two instruments, two verdicts. Every thesis is tested on NQ and on ES separately, each instrument with the signal from its own ticks and its own arbiter width (NQ ±160 T ≈ 40 points ≈ 0.17 % of price, ES ±40 T ≈ 10 points — the same relative move on the ES grid). One instrument alone does not make a thesis pass; the other's report sits in the same folder.

Arbiter

The arbiter is the one number declared the referee before the run: For ST1b the default per instrument: hit rate P(+X before −X) minus baseline at the arbiter width (NQ ±160 T, ES ±40 T), New York total, gross and over resolved anchors only, computed at the frozen threshold. For ST1a the declared deviation: difference of the median 30-min forward range (NQ, ticks), below minus above θ — a volatility quantity, therefore not signed. Both parts are judged separately and both reported.. It decides alone, and it decides in exactly the metric shown in the verdict table. Everything else in this report is exploration — it is shown in full because what works and what does not should be visible, but it does not turn this verdict. Whatever stands out here is a new thesis and re-runs on untouched data.

The verdict box: six phrases, and no more

Since the protocol reform of 2026-08-15 (thesis.md §2, rules 8–12) the verdict box carries exactly one of six stages, and each names the configuration in the same sentence: confirmed on ⟨config⟩ (dev and forward val passed) · confirmed (weak) on ⟨config⟩ (the val sign is right, but the val Δ is smaller than 50 % of the dev Δ) · dev passed on ⟨config⟩ — forward val running, verdict on ⟨date⟩ · no evidence on ⟨config⟩ · strong, unconfirmed on ⟨config⟩ — frozen for forward val, verdict on ⟨date⟩ · artefact: ⟨what the control arm reproduces⟩.

“The metric has no edge” is never written here. A verdict applies to the triple of condition, configuration and time scope, not to a metric in general — which is why it reads no evidence on ⟨config⟩. A null finding across a whole metric would only be justified once not a single configuration of the grid named in advance shows a neighbour-consistent Δ > 0 with z ≥ 2. And a confirmed stands: every further forward window after it runs as a labelled replication and builds the series, it does not turn the verdict.

Slot A and Slot B

Per thesis and instrument at most 2 frozen configurations enter the forward val — there are no more, and that is the guard against multiplicity. Slot A is the primary config: the user's live setting or the selection rule named in advance. Slot B is optional and is drawn mechanically from the dev grid under the plateau rule: the smallest threshold of the contiguous range in which Δ carries the predicted sign, |z| ≥ 2 holds, and the direct neighbours show the same sign. Smallest threshold means largest n and therefore the most conservative choice — not the z maximum, or the choice would be a ranking again.

Both slots are frozen with a date and get a verdict of their own; they are never netted against each other and never played off against each other. If the choice lands on the edge threshold of the grid, the report says ⟂ edge — the edge may lie beyond the grid. A grid may be extended in dev (dev is the digging window, the freeze is the guard); the extension is declared beforehand and the selection rule is then recomputed over the full grid.

Promotion: when an exploratory finding moves on

A finding outside the preregistration may stand as strong, unconfirmed if it meets all four criteria: (1) neighbour consistency — no lone spike, the grid neighbours show the same sign; (2) breadth — the predicted sign in ≥ 4 of 7 hour buckets, or in one contiguous block of ≥ 3 hours, which is then registered as the scope; (3) z ≥ 2 New York total — an entry ticket, not a verdict; (4) ≥ 30 carrying days. If the finding is not a grid cell (a cross-table cell, say), criterion 1 does not apply and criterion 4 tightens to ≥ 40 days.

Whatever passes is registered the same day as a thesis of its own with a freeze date and is then validated on forward data only — price theses included, because the tick store grows daily; the PX holdout stays reserved for the originally registered PX theses. Whatever fails stays exploration with the note too thin for promotion, n days = X — no vague “must re-run”. The criteria appear in the report as a mechanically recomputed table, not as an assertion.

Scoped resubmission. If a thesis fails on breadth even though the total Δ sits above every control arm with z ≥ 2, one check follows: do the positive buckets form a contiguous block? If they do, the hour-scoped version (“⟨thesis⟩, only ⟨block⟩”) attempts the promotion. If they do not, that is stated just as plainly — scattered hours are not a scope.

Carrying days and the date behind “verdict on”

A carrying day is a trading day that contributes at least one resolved anchor at the arbiter width to exactly this configuration — resolved meaning one of the two barriers was reached. The count appears as its own column in the grid and cross tables. Why it matters: anchors of the same day are correlated — 78 anchors from 12 days are not 78 observations, and a large n from few days is a day effect, not an edge.

The date behind verdict on is an estimate: the forward val runs over 30 trading days, so roughly six trading weeks — computed as freeze + 42 calendar days, then rolled to the next weekday. This calculation knows nothing about holidays; the real verdict falls once the 30 trading days have actually accumulated, not on the calendar date.

Δ, baseline, control span and ◆

Δ is real minus baseline, and the baseline is the median over the 8 control arms — the same run with the same condition, but with the sign series shifted cyclically across the days. The control span runs from the worst to the best arm. Whatever the baseline also achieves is no edge — only the distance counts. A means: the real arm sits above every single control arm, not just above their median. A ◆ is a marker, not significance: if the signal carries no information, the rank of the real arm among the 9 is uniform — a ◆ falls purely by chance in 1 of 9 cases (11 %). Likewise the min…max control span of 8 arms is only a 8/9 interval (88.9 %), not a 95 % interval (Learnings 27). And the hour buckets are 09:30–10:00 plus the full hours after it; because the anchors overlap, n is not a count of independent observations — how uncertain a cell is is read off the control span, not off n.

z and Ø z — the screening number per hour

z is the distance of the real arm from the mean of the 8 control arms, measured in their standard deviation: z = (real − mean of the arms) / sd of the arms. z = +2.0 means two standard deviations above what the shifted arms achieve on the same anchors. The column appears in the hit rate table per hour row and in the total row, at the arbiter width; in the time probe matrix every cell carries its z in the tooltip, as do the cells under “other bracket widths” and in the halves table — there against the control arms of the respective width. A marks |z| ≥ 2. Without spread (sd = 0) or without resolved anchors the cell reads “—”.

The ● does not mark the largest number but the largest distance in units of noise. The spread of the control arms differs in width per horizon and per bracket width — which is why a +1.3 can be marked at 30 min and a +4.6 not at 60 min: the 60-min path simply scatters far more, and a large tick value is nothing special there. Raw tick values are not comparable across horizons; that is exactly what Ø z beside them is for.

Ø z on the right of the time probe matrix is the mean of the z values of an hour row across the five horizons — the single number answering whether this hour sits outside the noise across all horizons. The z are averaged, not the tick Δ: the horizons have different scales. Cells without sd or without n drop out of the mean.

Two warnings, both meant seriously. First: the sd comes from only 8 arms and is therefore itself roughly estimated — z is not a clean test statistic but an order of magnitude. Second: with 7 hours × many cells |z| ≥ 2 also falls purely by chance; look long enough and you will always find a cell. And the five horizons overlap (the 1-min path sits inside the 5-min path and so on) — a high Ø z is therefore not a fivefold independent confirmation. z and Ø z are screening markers like the ◆, not a verdict: the verdict is passed by the preregistered arbiter alone.

Anchors

An anchor is a fixed sampling point: every 5 minutes between 09:30 and 15:55 New York the condition is read and turned into long, short or flat — 78 anchors per session. It deliberately does not hang on an entry signal; the question is whether the condition carries direction at all.

Hit rate probe

P(+X before −X) is the share of anchors at which price reaches X ticks in the signal direction before it reaches X ticks against it, within 60 minutes — gross on the raw price, without spread and commission: the barriers are symmetric, costs hit both sides equally and cancel out in the question about direction. The resolved column is the share of anchors at which either side was reached at all within 60 minutes; only those enter the rate, and whatever stays open drops out identically in both arms. Δ % is the difference in percentage points: hit rate minus baseline. +2 means 2 points more hits than the control, not 2 % relative — 51.3 % against 49.9 % is +1.4.

The highlighted columns are ±160 T — the arbiter of this study. The same quantity as Δ for the remaining widths sits collapsed under other bracket widths. It is read against the baseline, never against 50 %: a symmetric barrier does not land at 50 % by itself, drift and the mix of long and short move the zero point.

The width means ticks of this instrument: ±160 T are points on 40. The arbiter width is chosen per instrument so that it measures the same relative move (NQ ±160 T ≈ 0.17 % of price, ES ±40 T likewise); the resolved column shows how much of that is decided within 60 minutes at all.

Bracket to the close

Strictly speaking the 60-min rate measures P(+X before −X), given that one side is reached within 60 minutes. What resolves are preferentially the fast, volatile moments; the slow trend day drops out as open, the more often the wider the bracket. If an effect lives exactly there, the 60-min measurement systematically underrates it. The collapsed table bracket to the close therefore computes the same widths and the same booking a second time, capped at the RTH close 16:00 New York — the user is flat there, and does not trade what happens afterwards.

It does not judge. Baseline and z come from the same 8 control arms, computed with the same close cap — the real arm gets no special path. The arbiter nevertheless stays the 60-min row (NQ ±160 T, ES ±40 T): it is preregistered and frozen. An effect that shows up only in the close variant is exploration and must run as a new thesis on untouched data.

The 15:55 anchor has only 5 minutes left. The anchors run to 15:55, the cap sits fixed at 16:00 — from 15:05 onward the window is therefore shorter than 60 minutes. The 15–16 bucket is thus the only row in this table that gets less time than above; its resolved column and its n stand visibly beside it so it is not confused with the other hours.

Year by year

The year by year section reads the same arbiter metric a second time, split by calendar year — the year of the anchor in New York time, and because the anchors sit between 09:30 and 15:55 that is simply the calendar day of the session. Why at all: the market has quirks on an hourly basis and macroeconomic cycles above that; one number over six years hides that something carries in some years and not at all in others. The same logic as “total is irrelevant” across the day, only across the years.

The years are read-off slices, not runs of their own. The cyclic shift of the 8 control arms still runs over the full day list of the role; only afterwards are real arm and arms split by year. The baseline of a year is therefore the median of the same arms on the same anchors, not a freshly shuffled run. Edge years are partial years and carry their span in the name (“2020 (from 29.05.)”), in the chart a * — their n is smaller, their control span wider, and a year with few sessions swings accordingly.

The chart below shows Δ per year as bars (left axis) and z as a diamond (right axis). z stays screening, here even more than usual: with seven years × two instruments × several studies |z| ≥ 2 also falls purely by chance. A striking year is a new thesis and must run again on untouched data — it does not overturn a preregistered verdict.

The regime part: range instead of direction

This section asks not about direction but about volatility — a declared deviation from the default arbiter. What is measured is the 30-min forward range: highest minus lowest price in the 30 minutes after the anchor, in ticks. Per threshold two groups stand side by side (below θ and from θ), their medians are compared, and the target quantity is their difference, below minus above. The prediction is a negative value: below θ the market should be calmer. Green therefore means “as predicted” here, not “large”.

Two zero points, both mandatory. First the control difference: the same computation with the θ assignment shifted cyclically across the days — the day keeps its ranges but loses the assignment to its net gex; z measures the distance to that in their spread. Second the hour table: net gex and time of day are confounded (around midday the market is quiet anyway, and net gex accumulates over the day). Read only the total and you end up measuring the clock — which is why the dev criterion is decided by the number of hour buckets with the predicted sign.

Breakout persistence is the secondary quantity and purely descriptive: the share of 5-min closes outside the last 30-min range that are still outside five minutes later. It should be smaller below θ — but it decides nothing.

The threshold grid

The thresholds are named before the run and are all computed and reported — which one separates best is part of what should be visible here. The verdict is nevertheless passed on exactly one: it is chosen in the dev window, marked with ★ and frozen for the forward val; the choice and its reason are recorded in the measurement.

Seven columns are not seven tests (Learnings 10). Adjacent thresholds share their anchors — an anchor with large |net gex| sits in every threshold below it. The course across the row is therefore what is read: a smooth trend over several thresholds is a hint, a single spike between inconspicuous neighbours is noise.

Primary and secondary variant

The same thesis runs in two feed variants within the same run: 0DTE (nearest expiry) first and as the primary variant — that is where the verdict falls, because the user trades off that picture —, below it full across all expiries ≤ 90 days. Both compute the same days, the same anchors and the same control arms. An effect that shows up only in full is an exploratory finding and must run as a thesis of its own on untouched data (thesis.md §1).

Why only dev appears here

Since the addendum of 2026-08-15 options theses have a forward val: dev is the full coverage of the feed, and val is the first 30 trading days that accumulate after the freeze of the thesis — data that did not exist at freeze time and that therefore cannot have been looked at by construction. That window does not exist yet today; the verdict box therefore stands at the stage dev passed on ⟨config⟩ — forward val running, verdict on ⟨date⟩, provided the dev criterion is met (predicted sign in at least half of the hour buckets), otherwise at no evidence on ⟨config⟩. Only the one val run says confirmed.

The hours are all there is here. Options theses are New-York-only by construction — the feed covers RTH 09:30–16:00 alone. A split by Asia and London is therefore absent not out of convenience but because there is no data there; the session breakdown from the learnings is the hour breakdown here.

Distribution across the days

One hit rate per session, turned into a histogram. The question is not whether the rate is right but whether it is carried by many days or by few. The bars are the ten deciles from 0 % to 100 %, the height is the share of days in that decile; days without a resolved anchor in that hour do not count. The tooltip above the histogram names the median day, the band from 10 % to 90 % and the number of days, the tooltip per bar the days in that decile.

The hour rows hold only about a dozen anchors per day — the daily rate is coarsely stepped there. The New York total row is the robust one: there every day has all 78 anchors.

Time probe

The signed forward return (FR) is the price path in ticks, turned into signal direction: signal short and price falls 10 ticks means +10. Bracket-free — no stop, no target, nothing that truncates the result — and gross, because costs move every cell equally and cancel out in the difference to the baseline. What is shown is Δ to the baseline; green means above. This is the second question — direction is answered by the hit rate probe.

MFE / MAEmaximum favorable / adverse excursion: how far price ran at most for and against the signal within the horizon. A good MFE with a poor forward return means: the move was there, it just did not hold to the horizon.

Both halves

A setting that carries in only one half is not a setting but a period (Learnings 26). First the hit rate, then the time probe — both as Δ to the baseline.

Breadth

Breadth is the share of cells above the baseline — 7 hours × bracket widths or × horizons; beside it stands what the best and the worst control arm achieve on the same computation. For the hit rate a control arm sits at around 50 % by construction. The verdict is never passed on the best cell.

Run
Instrument of this reportNQ
SignalSign of sum_gex_vol from feed ES_SPX state (0DTE), only anchors with |sum_gex_vol| ≥ 1500 — 107 trading days, 2315895 frames, 0.2 % of them warm-up (value 0, flat)
Sessions read105
Anchors per session78 (09:30–15:55, every 5 min)
Horizons1 / 5 / 15 / 30 / 60 min, bracket horizon 60 min; second bracket run capped at the RTH close 16:00 (context)
Bracket widths40 / 80 / 120 / 160 / 200 / 280 / 320 / 400 / 560 ticks
ArbiterP(+160 before −160)
Control arms8, seed 20260813, cyclic shift of the sign series across the days
Half-spread at entry, measured1.31 ticks
Reading ticks3.1 s
Run total10.1 s
Commitf1a70e3c099644522bfadfea043aa7c2b6d7cefb