This report judges ES only. Every thesis gets two separate verdicts, one per instrument, each on its own signal and with its own arbiter width (P(+40 before −40) here). The other instrument sits in a report of its own; neither decides anything about the other.
Slot A — confirmed on Seeded randomness per anchor (seed 20260813, own series for ES): long or short at 50 % each, deterministic from seed, instrument, day and slot, with no market input at all
No configuration of its own declared — the condition itself is the primary config. Frozen on 2026-05-01.
Inverted — passing means: nothing found. The real arm sits, New York total, inside the band mean of the control arms ± 2·sd and likewise in at least 5 of 7 hour buckets. If the random arm falls outside, that is a finding about the machine, not about the market. The criterion holds per instrument separately, on its own random series and its own arbiter width. ADDENDUM AFTER THE FIRST RUN (2026-08-15): what was preregistered was the min…max span of the 8 control arms. Under exchangeability that is a 7/9 interval and rejects perfect randomness in about a third of the runs — which is exactly what happened in the first run. The band was therefore changed to mean ± 2·sd; both verdicts stand side by side in the measurement.
| Slot · config | Role | Sessions | n | P(+40 before −40) | Baseline | Δ % | Control span | Buckets inside the band | Result |
|---|---|---|---|---|---|---|---|---|---|
| Slot A · Seeded randomness per anchor (seed 20260813, own series for ES): long or short at 50 % each, deterministic from seed, instrument, day and slot, with no market input at all | full | 1467 | 71228 | 50.0 % | 49.9 % | +0.04 | 49.8 … 50.1 % | 7 of 7 | passed |
| Claim | A signal without information shows no edge. This study calibrates the machine, it claims nothing about the market. |
| Condition | One random direction per anchor, long or short at 50 % each, mixed deterministically from seed, instrument, day and slot (splitmix64). No market input, no state, no ticks read. Since 2026-08-15 every instrument draws its own independent series — the instrument ordinal (NQ 0, ES 1) enters the hash as a splitmix64 step of its own so that NQ and ES do not trade the same signs. |
| Prediction | No edge — the real arm sits inside the spread of the control arms, the distance of the hit rate is zero, and the raw hit rate P(+X before −X) sits close to 50 %, because drift cancels out under random direction choice. That holds per instrument separately. |
| Target metric (arbiter) | P(+40 before −40) — Hit rate P(+X before −X) minus baseline at the arbiter width of the instrument (NQ ±160 T, ES ±40 T), New York total, gross and over resolved anchors only. Two separate verdicts, one per instrument. |
| Baseline (control arm) | The same anchors, direction by a cyclic shift of the sign series of the respective instrument across the days. Shifted randomness stays randomness — that is exactly the point here. |
| Success criterion | Inverted — passing means: nothing found. The real arm sits, New York total, inside the band mean of the control arms ± 2·sd and likewise in at least 5 of 7 hour buckets. If the random arm falls outside, that is a finding about the machine, not about the market. The criterion holds per instrument separately, on its own random series and its own arbiter width. ADDENDUM AFTER THE FIRST RUN (2026-08-15): what was preregistered was the min…max span of the 8 control arms. Under exchangeability that is a 7/9 interval and rejects perfect randomness in about a third of the runs — which is exactly what happened in the first run. The band was therefore changed to mean ± 2·sd; both verdicts stand side by side in the measurement. |
| Signal | Seeded randomness per anchor (seed 20260813, own series for ES): long or short at 50 % each, deterministic from seed, instrument, day and slot, with no market input at all — read on no ticks |
| Data roles | full: 2020-05-29 to 2026-04-30 |
| Holdout | from 2026-05-01 — not computed in this run, not looked at |
Seeded randomness per anchor (seed 20260813, own series for ES): long or short at 50 % each, deterministic from seed, instrument, day and slot, with no market input at all · brackets 20 / 40 / 80 / 120 / 160 / 200 / 300 / 400 ticks · fill phase spread 0.56/0.54/0.51/0.54
Arbiter of this role: +0.04 % away from the baseline, inside the spread of the control arms in 7 of 7 hour buckets.
1467 sessions, 2020-06-11 to 2026-04-30 · 114426 anchors: 57477 long (50.2 %) / 56949 short / 0 flat · 0 without ticks · 0 horizons running past the session end · 8 control arms, shifts +1185, +327, +836, +206, +954, +1315, +962, +1174
| Hour | P(+40 before −40) | Baseline | Δ % | Control span | z | resolved | n | Distribution of the daily hit rates | Days above baseline |
|---|---|---|---|---|---|---|---|---|---|
| 09:30–10:00 | 50.4 % | 50.2 % | +0.17 | 48.9 … 51.2 % | +0.5 | 82.0 % | 7220 | 518 of 1383 (37.5 %) | |
| 10–11 | 50.4 % | 49.7 % | +0.74 | 49.3 … 50.8 % | +0.8 | 72.0 % | 12680 | 840 of 1361 (61.7 %) | |
| 11–12 | 49.6 % | 49.9 % | -0.26 | 49.5 … 50.5 % | -0.8 | 60.3 % | 10615 | 728 of 1251 (58.2 %) | |
| 12–13 | 50.0 % | 50.0 % | +0.02 | 49.2 … 50.5 % | +0.2 | 55.2 % | 9720 | 700 of 1195 (58.6 %) | |
| 13–14 | 49.8 % | 49.8 % | -0.04 | 49.1 … 50.3 % | -0.2 | 56.9 % | 10021 | 670 of 1180 (56.8 %) | |
| 14–15 | 50.0 % | 50.0 % | +0.05 | 49.5 … 50.2 % | +0.3 | 58.3 % | 10265 | 698 of 1214 (57.5 %) | |
| 15–16 | 49.6 % | 50.2 % | -0.64 | 49.2 … 50.4 % | -1.0 | 60.8 % | 10707 | 516 of 1243 (41.5 %) | |
| New York total | 50.0 % | 49.9 % | +0.04 | 49.8 … 50.1 % | +0.3 | 62.2 % | 71228 | 765 of 1452 (52.7 %) |
| Hour | Δ ±20 T | Δ ±80 T | Δ ±120 T | Δ ±160 T | Δ ±200 T | Δ ±300 T | Δ ±400 T |
|---|---|---|---|---|---|---|---|
| 09:30–10:00 | -0.21 | +0.42 | +0.23 | -0.88 | +1.60 | -5.00 | -13.33 |
| 10–11 | -0.66 | -0.31 | -0.56 | +1.42 | -1.64 | +2.13 | -3.23 |
| 11–12 | +0.12 | -0.25 | -1.80 | -2.78 | +1.22 | +3.39 | +16.00 ● |
| 12–13 | +0.36 | +0.34 | -1.83 | -3.00 | +1.90 | -5.77 | -10.00 |
| 13–14 | -0.08 | -0.65 | -2.11 | -3.99 | -4.17 | +14.00 | +5.56 |
| 14–15 | +0.69 | -0.88 | -2.26 | -3.44 | -4.39 | -6.45 | +15.38 |
| 15–16 | -0.59 | +0.21 | +1.38 ● | +3.72 ● | +2.37 | +10.26 | +16.67 |
| New York total | +0.15 | -0.07 | -0.78 | -0.37 | -0.13 | +3.65 | +6.92 |
Breadth (7 hours × 8 bracket widths): 46.4 %, control arms 21.4 % to 53.6 %.
The same calculation as above, except the bracket does not run for 60 minutes but until 16:00 New York — the real arm and all 8 control arms under the same cap. Context, not a verdict: the arbiter stays the 60-min row above.
| Hour | P(+40 before −40) | Baseline | Δ % | Control span | z | resolved | n | Δ ±20 T | Δ ±80 T | Δ ±120 T | Δ ±160 T | Δ ±200 T | Δ ±300 T | Δ ±400 T |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 09:30–10:00 | 50.4 % | 50.0 % | +0.40 | 49.0 … 50.9 % | +0.8 | 98.5 % | 8667 | -0.14 | +0.16 | -0.86 | -0.97 | -1.16 | -2.67 | -3.75 ● |
| 10–11 | 50.3 % | 49.8 % | +0.47 | 49.3 … 50.9 % | +0.7 | 97.4 % | 17140 | -0.71 | -0.09 | -0.40 | +0.14 | +0.44 | -0.49 | +0.19 |
| 11–12 | 50.0 % | 49.9 % | +0.02 | 49.7 … 50.3 % | +0.1 | 95.1 % | 16745 | +0.15 | -0.21 | -0.66 | +0.27 | -0.18 | +1.35 | +0.00 |
| 12–13 | 50.2 % | 50.2 % | -0.02 | 49.7 … 50.4 % | +0.5 | 91.6 % | 16131 | +0.44 | -0.47 | -0.63 | -1.86 ● | -1.57 | -0.86 | -4.82 |
| 13–14 | 50.4 % | 49.8 % | +0.58 | 49.3 … 50.3 % ◆ | +2.0 ● | 86.0 % | 15143 | +0.13 | -0.60 | -0.47 | -0.84 | -1.21 | +1.41 | +3.95 |
| 14–15 | 50.0 % | 50.0 % | -0.06 | 49.8 … 50.6 % | -0.4 | 76.8 % | 13515 | +0.67 ● | -0.14 | -0.56 | -0.96 | -1.85 | -1.65 | +1.12 |
| 15–16 | 49.4 % | 50.2 % | -0.78 | 49.3 … 50.6 % | -1.6 | 51.8 % | 9118 | -0.44 | -0.70 | +0.58 | +4.34 ● | +3.29 | +10.42 | +27.27 ● |
| New York total | 50.1 % | 50.0 % | +0.14 | 49.7 … 50.2 % | +1.2 | 84.3 % | 96459 | +0.15 | -0.18 | -0.50 | -0.46 | -0.52 | -0.44 | +0.11 |
The resolved column is the interesting one here: it rises against 60 minutes because the slow path is given time. Only the last row runs the other way — the 15:55 anchor has 5 minutes left until the close, so the 15–16 bucket is capped shorter than above, not longer.
The same metric as above, only read by calendar year — reading cuts through the same 8 control arms, edge years labelled with their span; z is screening, not a verdict.
| Year | P(+40 before −40) | Baseline | Δ % | Control span | z | resolved | n | Sessions |
|---|---|---|---|---|---|---|---|---|
| 2020 (from 11.06.) | 50.2 % | 50.7 % | -0.51 | 49.0 … 51.0 % | -0.1 | 53.6 % | 5854 | 140 |
| 2021 | 49.6 % | 49.9 % | -0.29 | 49.5 … 50.6 % | -0.8 | 40.5 % | 7926 | 251 |
| 2022 | 49.6 % | 50.0 % | -0.37 | 49.7 … 50.6 % | -1.6 | 84.0 % | 16374 | 250 |
| 2023 | 50.1 % | 49.5 % | +0.55 | 49.0 … 50.6 % | +0.7 | 54.5 % | 10536 | 248 |
| 2024 | 49.7 % | 50.0 % | -0.33 | 49.2 … 50.6 % | -0.8 | 59.2 % | 11493 | 249 |
| 2025 | 51.0 % | 49.8 % | +1.20 | 49.5 … 50.2 % ◆ | +5.9 ● | 72.2 % | 13903 | 247 |
| 2026 (until 30.04.) | 49.2 % | 50.5 % | -1.36 | 49.3 … 50.8 % | -2.0 ● | 80.4 % | 5142 | 82 |
Bars = Δ % per year against the baseline (left axis), ◆ = z against the control arms of the same year (right axis), filled from |z| ≥ 2. * marks a partial year.
| Hour | 1 min | 5 min | 15 min | 30 min | 60 min | n | Ø z |
|---|---|---|---|---|---|---|---|
| 09:30–10:00 | +0.04 | +0.26 | +0.40 | +0.92 | +0.19 | 8802 | +0.9 |
| 10–11 | +0.06 | -0.09 | +0.05 | -0.00 | +0.10 | 17604 | +0.4 |
| 11–12 | +0.05 | +0.10 | +0.05 | +0.13 | +0.04 | 17604 | +0.5 |
| 12–13 | +0.07 | +0.02 | +0.18 | -0.11 | -0.55 | 17604 | +0.4 |
| 13–14 | -0.02 | -0.01 | +0.03 | +0.19 | +0.34 ● | 17604 | +0.8 |
| 14–15 | +0.06 ● | +0.15 | -0.20 | -0.43 | -0.30 | 17604 | +0.1 |
| 15–16 | -0.01 | +0.08 | -0.26 | +0.06 | +0.19 | 17604 | +0.3 |
| New York total | +0.04 | +0.10 ● | +0.04 | +0.06 | +0.01 | 114426 | +1.0 |
| Hour | Horizon | n | Ø FR | Baseline | Δ | Control span | Share positive | Ø MFE | Ø MAE |
|---|---|---|---|---|---|---|---|---|---|
| 09:30–10:00 | 1 min | 8802 | +0.01 | -0.02 | +0.04 | -0.12 … +0.23 | 47.6 % | +8.55 | -8.53 |
| 09:30–10:00 | 5 min | 8802 | +0.38 | +0.12 | +0.26 | -0.43 … +0.44 | 49.0 % | +17.23 | -16.84 |
| 09:30–10:00 | 15 min | 8802 | +0.55 | +0.15 | +0.40 | -1.02 … +0.69 | 49.9 % | +29.52 | -28.91 |
| 09:30–10:00 | 30 min | 8802 | +1.14 | +0.22 | +0.92 | -0.74 … +1.20 | 50.2 % | +41.24 | -39.96 |
| 09:30–10:00 | 60 min | 8802 | +0.53 | +0.34 | +0.19 | -0.51 … +1.09 | 50.0 % | +54.81 | -53.64 |
| 10–11 | 1 min | 17604 | +0.09 | +0.02 | +0.06 | -0.17 … +0.06 ◆ | 47.4 % | +7.02 | -6.94 |
| 10–11 | 5 min | 17604 | -0.03 | +0.06 | -0.09 | -0.19 … +0.10 | 48.0 % | +14.48 | -14.42 |
| 10–11 | 15 min | 17604 | +0.10 | +0.05 | +0.05 | -0.42 … +0.47 | 49.4 % | +24.29 | -24.33 |
| 10–11 | 30 min | 17604 | +0.03 | +0.03 | -0.00 | -0.83 … +0.68 | 49.4 % | +33.35 | -33.24 |
| 10–11 | 60 min | 17604 | +0.14 | +0.04 | +0.10 | -1.11 … +0.98 | 49.7 % | +44.55 | -44.40 |
| 11–12 | 1 min | 17604 | +0.05 | -0.00 | +0.05 | -0.08 … +0.20 | 46.4 % | +5.50 | -5.45 |
| 11–12 | 5 min | 17604 | +0.09 | -0.01 | +0.10 | -0.19 … +0.09 | 48.3 % | +11.58 | -11.41 |
| 11–12 | 15 min | 17604 | +0.01 | -0.04 | +0.05 | -0.19 … +0.29 | 49.3 % | +19.53 | -19.50 |
| 11–12 | 30 min | 17604 | +0.03 | -0.11 | +0.13 | -0.29 … +0.37 | 49.3 % | +26.97 | -27.08 |
| 11–12 | 60 min | 17604 | +0.15 | +0.11 | +0.04 | -0.60 … +0.58 | 49.8 % | +36.97 | -36.86 |
| 12–13 | 1 min | 17604 | +0.04 | -0.03 | +0.07 | -0.09 … +0.07 | 45.8 % | +4.73 | -4.69 |
| 12–13 | 5 min | 17604 | -0.01 | -0.03 | +0.02 | -0.25 … +0.05 | 47.9 % | +9.97 | -9.93 |
| 12–13 | 15 min | 17604 | +0.16 | -0.02 | +0.18 | -0.19 … +0.12 ◆ | 49.2 % | +17.26 | -17.13 |
| 12–13 | 30 min | 17604 | -0.09 | +0.02 | -0.11 | -0.14 … +0.13 | 49.2 % | +24.35 | -24.27 |
| 12–13 | 60 min | 17604 | -0.44 | +0.10 | -0.55 | -0.51 … +0.55 | 49.4 % | +34.48 | -35.00 |
| 13–14 | 1 min | 17604 | -0.03 | -0.01 | -0.02 | -0.15 … +0.06 | 45.3 % | +4.62 | -4.65 |
| 13–14 | 5 min | 17604 | -0.05 | -0.04 | -0.01 | -0.18 … +0.20 | 47.6 % | +9.81 | -9.87 |
| 13–14 | 15 min | 17604 | -0.04 | -0.07 | +0.03 | -0.13 … -0.02 | 48.6 % | +17.35 | -17.41 |
| 13–14 | 30 min | 17604 | +0.18 | -0.00 | +0.19 | -0.64 … +0.25 | 49.0 % | +24.96 | -24.97 |
| 13–14 | 60 min | 17604 | +0.27 | -0.07 | +0.34 | -0.44 … +0.07 ◆ | 49.6 % | +36.01 | -36.05 |
| 14–15 | 1 min | 17604 | +0.08 | +0.02 | +0.06 | -0.05 … +0.05 ◆ | 46.0 % | +5.19 | -5.08 |
| 14–15 | 5 min | 17604 | +0.17 | +0.01 | +0.15 | -0.16 … +0.24 | 48.3 % | +10.73 | -10.50 |
| 14–15 | 15 min | 17604 | -0.08 | +0.12 | -0.20 | -0.24 … +0.27 | 48.7 % | +18.34 | -18.18 |
| 14–15 | 30 min | 17604 | -0.45 | -0.02 | -0.43 | -0.36 … +0.51 | 49.2 % | +25.80 | -26.12 |
| 14–15 | 60 min | 17604 | -0.21 | +0.09 | -0.30 | -0.60 … +0.42 | 49.9 % | +37.41 | -37.74 |
| 15–16 | 1 min | 17604 | -0.01 | +0.00 | -0.01 | -0.07 … +0.09 | 46.3 % | +6.05 | -6.07 |
| 15–16 | 5 min | 17604 | +0.12 | +0.04 | +0.08 | -0.28 … +0.27 | 48.5 % | +12.24 | -12.16 |
| 15–16 | 15 min | 17604 | -0.08 | +0.19 | -0.26 | -0.19 … +0.30 | 48.8 % | +21.15 | -21.24 |
| 15–16 | 30 min | 17604 | +0.25 | +0.20 | +0.06 | -0.03 … +0.39 | 49.0 % | +29.56 | -29.30 |
| 15–16 | 60 min | 17604 | +0.30 | +0.11 | +0.19 | -0.24 … +0.42 | 49.3 % | +38.42 | -38.11 |
| New York total | 1 min | 114426 | +0.03 | -0.01 | +0.04 | -0.03 … +0.07 | 46.3 % | +5.75 | -5.72 |
| New York total | 5 min | 114426 | +0.07 | -0.03 | +0.10 | -0.08 … +0.06 ◆ | 48.2 % | +11.91 | -11.80 |
| New York total | 15 min | 114426 | +0.05 | +0.02 | +0.04 | -0.12 … +0.17 | 49.1 % | +20.41 | -20.34 |
| New York total | 30 min | 114426 | +0.08 | +0.03 | +0.06 | -0.19 … +0.19 | 49.2 % | +28.56 | -28.45 |
| New York total | 60 min | 114426 | +0.07 | +0.06 | +0.01 | -0.26 … +0.23 | 49.6 % | +39.27 | -39.23 |
Breadth (7 hours × 5 horizons): 68.6 %, control arms 20.0 % to 68.6 %.
| Half | Bracket | n resolved | P(+X before −X) | Baseline | Δ % |
|---|---|---|---|---|---|
| 1 (2020-06-11 to 2023-05-16) | ±20 T | 51770 | 50.2 % | 49.9 % | +0.23 |
| 1 (2020-06-11 to 2023-05-16) | ±40 T | 34787 | 49.8 % | 50.0 % | -0.24 |
| 1 (2020-06-11 to 2023-05-16) | ±80 T | 12978 | 49.8 % | 50.0 % | -0.28 |
| 1 (2020-06-11 to 2023-05-16) | ±120 T | 4925 | 48.6 % | 50.4 % | -1.75 ● |
| 1 (2020-06-11 to 2023-05-16) | ±160 T | 1999 | 48.3 % | 50.7 % | -2.40 ● |
| 1 (2020-06-11 to 2023-05-16) | ±200 T | 846 | 48.0 % | 50.9 % | -2.96 ● |
| 1 (2020-06-11 to 2023-05-16) | ±300 T | 114 | 53.5 % | 50.0 % | +3.51 |
| 1 (2020-06-11 to 2023-05-16) | ±400 T | 26 | 53.8 % | 50.0 % | +3.85 |
| 2 (2023-05-17 to 2026-04-30) | ±20 T | 54001 | 50.0 % | 49.9 % | +0.01 |
| 2 (2023-05-17 to 2026-04-30) | ±40 T | 36441 | 50.1 % | 50.0 % | +0.19 |
| 2 (2023-05-17 to 2026-04-30) | ±80 T | 13901 | 50.0 % | 50.0 % | -0.04 |
| 2 (2023-05-17 to 2026-04-30) | ±120 T | 5700 | 50.1 % | 50.2 % | -0.16 |
| 2 (2023-05-17 to 2026-04-30) | ±160 T | 2600 | 50.7 % | 50.3 % | +0.42 |
| 2 (2023-05-17 to 2026-04-30) | ±200 T | 1380 | 51.7 % | 50.1 % | +1.59 |
| 2 (2023-05-17 to 2026-04-30) | ±300 T | 352 | 53.4 % | 50.6 % | +2.84 |
| 2 (2023-05-17 to 2026-04-30) | ±400 T | 133 | 58.6 % | 51.1 % | +7.52 |
| Half | Horizon | n | Ø FR | Baseline | Δ (ticks) |
|---|---|---|---|---|---|
| 1 (2020-06-11 to 2023-05-16) | 1 min | 57252 | +0.01 | -0.01 | +0.02 |
| 1 (2020-06-11 to 2023-05-16) | 5 min | 57252 | +0.04 | -0.00 | +0.04 |
| 1 (2020-06-11 to 2023-05-16) | 15 min | 57252 | -0.12 | +0.04 | -0.15 |
| 1 (2020-06-11 to 2023-05-16) | 30 min | 57252 | -0.10 | +0.10 | -0.21 |
| 1 (2020-06-11 to 2023-05-16) | 60 min | 57252 | -0.16 | +0.11 | -0.27 |
| 2 (2023-05-17 to 2026-04-30) | 1 min | 57174 | +0.06 | -0.00 | +0.06 |
| 2 (2023-05-17 to 2026-04-30) | 5 min | 57174 | +0.11 | -0.04 | +0.14 |
| 2 (2023-05-17 to 2026-04-30) | 15 min | 57174 | +0.23 | -0.00 | +0.23 |
| 2 (2023-05-17 to 2026-04-30) | 30 min | 57174 | +0.27 | -0.05 | +0.32 |
| 2 (2023-05-17 to 2026-04-30) | 60 min | 57174 | +0.30 | -0.07 | +0.38 |
10 sessions, drawn by seed (20260813), not hand-picked — illustration, not evidence (Learnings 11). They are the same days as in the other instrument's report, so the two can be laid side by side. Price ES as a 1-minute path in points; there is no indicator line here — the condition lives in a different price space and does not fit on this axis, only the anchors are visible.
Every marker is an anchor: ▲ long, ▼ short; green = +40 T reached first, red = −40 T first, grey = neither side within 60 minutes.
78 anchors · 45 long / 33 short · 33 resolved, 11 of them with +40 T first (33.3 %)
78 anchors · 45 long / 33 short · 62 resolved, 25 of them with +40 T first (40.3 %)
78 anchors · 36 long / 42 short · 51 resolved, 28 of them with +40 T first (54.9 %)
78 anchors · 38 long / 40 short · 72 resolved, 37 of them with +40 T first (51.4 %)
78 anchors · 36 long / 42 short · 24 resolved, 14 of them with +40 T first (58.3 %)
78 anchors · 37 long / 41 short · 42 resolved, 16 of them with +40 T first (38.1 %)
78 anchors · 45 long / 33 short · 61 resolved, 33 of them with +40 T first (54.1 %)
78 anchors · 39 long / 39 short · 0 resolved, 0 of them with +40 T first (—)
78 anchors · 46 long / 32 short · 40 resolved, 17 of them with +40 T first (42.5 %)
78 anchors · 38 long / 40 short · 10 resolved, 6 of them with +40 T first (60.0 %)
This report belongs to one instrument (ES) and judges nothing else. First the verdict, then per data role (dev, then val) the same set of sections, always in the same order: hit rate probe including the distribution across days (is the direction right? — the arbiter) · time probe (how many ticks were there to take) · both halves of the window separately. Which section carries the arbiter depends on the preregistration of this study; the order is fixed. Nothing collapsed is second-rate, only second-asked — the other bracket widths, the full MFE/MAE matrix and the halves.
Money is no longer reported here. The probe asks about direction; the tradability section (money per sizing model and the metrics beside it) was cut from this report without replacement on 2026-08-15. Whether a direction can be traded after costs is checked later by a test of its own — once a thesis holds up.
Two instruments, two verdicts. Every thesis is tested on NQ and on ES separately, each instrument with the signal from its own ticks and its own arbiter width (NQ ±160 T ≈ 40 points ≈ 0.17 % of price, ES ±40 T ≈ 10 points — the same relative move on the ES grid). One instrument alone does not make a thesis pass; the other's report sits in the same folder.
The arbiter is the one number declared the referee before the run: Hit rate P(+X before −X) minus baseline at the arbiter width of the instrument (NQ ±160 T, ES ±40 T), New York total, gross and over resolved anchors only. Two separate verdicts, one per instrument.. It decides alone, and it decides in exactly the metric shown in the verdict table. Everything else in this report is exploration — it is shown in full because what works and what does not should be visible, but it does not turn this verdict. Whatever stands out here is a new thesis and re-runs on untouched data.
Since the protocol reform of 2026-08-15 (thesis.md §2, rules 8–12) the verdict box carries exactly one of six stages, and each names the configuration in the same sentence: confirmed on ⟨config⟩ (dev and forward val passed) · confirmed (weak) on ⟨config⟩ (the val sign is right, but the val Δ is smaller than 50 % of the dev Δ) · dev passed on ⟨config⟩ — forward val running, verdict on ⟨date⟩ · no evidence on ⟨config⟩ · strong, unconfirmed on ⟨config⟩ — frozen for forward val, verdict on ⟨date⟩ · artefact: ⟨what the control arm reproduces⟩.
“The metric has no edge” is never written here. A verdict applies to the triple of condition, configuration and time scope, not to a metric in general — which is why it reads no evidence on ⟨config⟩. A null finding across a whole metric would only be justified once not a single configuration of the grid named in advance shows a neighbour-consistent Δ > 0 with z ≥ 2. And a confirmed stands: every further forward window after it runs as a labelled replication and builds the series, it does not turn the verdict.
Per thesis and instrument at most 2 frozen configurations enter the forward val — there are no more, and that is the guard against multiplicity. Slot A is the primary config: the user's live setting or the selection rule named in advance. Slot B is optional and is drawn mechanically from the dev grid under the plateau rule: the smallest threshold of the contiguous range in which Δ carries the predicted sign, |z| ≥ 2 holds, and the direct neighbours show the same sign. Smallest threshold means largest n and therefore the most conservative choice — not the z maximum, or the choice would be a ranking again.
Both slots are frozen with a date and get a verdict of their own; they are never netted against each other and never played off against each other. If the choice lands on the edge threshold of the grid, the report says ⟂ edge — the edge may lie beyond the grid. A grid may be extended in dev (dev is the digging window, the freeze is the guard); the extension is declared beforehand and the selection rule is then recomputed over the full grid.
A finding outside the preregistration may stand as strong, unconfirmed if it meets all four criteria: (1) neighbour consistency — no lone spike, the grid neighbours show the same sign; (2) breadth — the predicted sign in ≥ 4 of 7 hour buckets, or in one contiguous block of ≥ 3 hours, which is then registered as the scope; (3) z ≥ 2 New York total — an entry ticket, not a verdict; (4) ≥ 30 carrying days. If the finding is not a grid cell (a cross-table cell, say), criterion 1 does not apply and criterion 4 tightens to ≥ 40 days.
Whatever passes is registered the same day as a thesis of its own with a freeze date and is then validated on forward data only — price theses included, because the tick store grows daily; the PX holdout stays reserved for the originally registered PX theses. Whatever fails stays exploration with the note too thin for promotion, n days = X — no vague “must re-run”. The criteria appear in the report as a mechanically recomputed table, not as an assertion.
Scoped resubmission. If a thesis fails on breadth even though the total Δ sits above every control arm with z ≥ 2, one check follows: do the positive buckets form a contiguous block? If they do, the hour-scoped version (“⟨thesis⟩, only ⟨block⟩”) attempts the promotion. If they do not, that is stated just as plainly — scattered hours are not a scope.
A carrying day is a trading day that contributes at least one resolved anchor at the arbiter width to exactly this configuration — resolved meaning one of the two barriers was reached. The count appears as its own column in the grid and cross tables. Why it matters: anchors of the same day are correlated — 78 anchors from 12 days are not 78 observations, and a large n from few days is a day effect, not an edge.
The date behind verdict on is an estimate: the forward val runs over 30 trading days, so roughly six trading weeks — computed as freeze + 42 calendar days, then rolled to the next weekday. This calculation knows nothing about holidays; the real verdict falls once the 30 trading days have actually accumulated, not on the calendar date.
Δ is real minus baseline, and the baseline is the median over the 8 control arms — the same run with the same condition, but with the sign series shifted cyclically across the days. The control span runs from the worst to the best arm. Whatever the baseline also achieves is no edge — only the distance counts. A ◆ means: the real arm sits above every single control arm, not just above their median. A ◆ is a marker, not significance: if the signal carries no information, the rank of the real arm among the 9 is uniform — a ◆ falls purely by chance in 1 of 9 cases (11 %). Likewise the min…max control span of 8 arms is only a 8/9 interval (88.9 %), not a 95 % interval (Learnings 27). And the hour buckets are 09:30–10:00 plus the full hours after it; because the anchors overlap, n is not a count of independent observations — how uncertain a cell is is read off the control span, not off n.
z is the distance of the real arm from the mean of the 8 control arms, measured in their standard deviation: z = (real − mean of the arms) / sd of the arms. z = +2.0 means two standard deviations above what the shifted arms achieve on the same anchors. The column appears in the hit rate table per hour row and in the total row, at the arbiter width; in the time probe matrix every cell carries its z in the tooltip, as do the cells under “other bracket widths” and in the halves table — there against the control arms of the respective width. A ● marks |z| ≥ 2. Without spread (sd = 0) or without resolved anchors the cell reads “—”.
The ● does not mark the largest number but the largest distance in units of noise. The spread of the control arms differs in width per horizon and per bracket width — which is why a +1.3 can be marked at 30 min and a +4.6 not at 60 min: the 60-min path simply scatters far more, and a large tick value is nothing special there. Raw tick values are not comparable across horizons; that is exactly what Ø z beside them is for.
Ø z on the right of the time probe matrix is the mean of the z values of an hour row across the five horizons — the single number answering whether this hour sits outside the noise across all horizons. The z are averaged, not the tick Δ: the horizons have different scales. Cells without sd or without n drop out of the mean.
Two warnings, both meant seriously. First: the sd comes from only 8 arms and is therefore itself roughly estimated — z is not a clean test statistic but an order of magnitude. Second: with 7 hours × many cells |z| ≥ 2 also falls purely by chance; look long enough and you will always find a cell. And the five horizons overlap (the 1-min path sits inside the 5-min path and so on) — a high Ø z is therefore not a fivefold independent confirmation. z and Ø z are screening markers like the ◆, not a verdict: the verdict is passed by the preregistered arbiter alone.
An anchor is a fixed sampling point: every 5 minutes between 09:30 and 15:55 New York the condition is read and turned into long, short or flat — 78 anchors per session. It deliberately does not hang on an entry signal; the question is whether the condition carries direction at all.
P(+X before −X) is the share of anchors at which price reaches X ticks in the signal direction before it reaches X ticks against it, within 60 minutes — gross on the raw price, without spread and commission: the barriers are symmetric, costs hit both sides equally and cancel out in the question about direction. The resolved column is the share of anchors at which either side was reached at all within 60 minutes; only those enter the rate, and whatever stays open drops out identically in both arms. Δ % is the difference in percentage points: hit rate minus baseline. +2 means 2 points more hits than the control, not 2 % relative — 51.3 % against 49.9 % is +1.4.
The highlighted columns are ±40 T — the arbiter of this study. The same quantity as Δ for the remaining widths sits collapsed under other bracket widths. It is read against the baseline, never against 50 %: a symmetric barrier does not land at 50 % by itself, drift and the mix of long and short move the zero point.
The width means ticks of this instrument: ±40 T are points on 10. The arbiter width is chosen per instrument so that it measures the same relative move (NQ ±160 T ≈ 0.17 % of price, ES ±40 T likewise); the resolved column shows how much of that is decided within 60 minutes at all.
Strictly speaking the 60-min rate measures P(+X before −X), given that one side is reached within 60 minutes. What resolves are preferentially the fast, volatile moments; the slow trend day drops out as open, the more often the wider the bracket. If an effect lives exactly there, the 60-min measurement systematically underrates it. The collapsed table bracket to the close therefore computes the same widths and the same booking a second time, capped at the RTH close 16:00 New York — the user is flat there, and does not trade what happens afterwards.
It does not judge. Baseline and z come from the same 8 control arms, computed with the same close cap — the real arm gets no special path. The arbiter nevertheless stays the 60-min row (NQ ±160 T, ES ±40 T): it is preregistered and frozen. An effect that shows up only in the close variant is exploration and must run as a new thesis on untouched data.
The 15:55 anchor has only 5 minutes left. The anchors run to 15:55, the cap sits fixed at 16:00 — from 15:05 onward the window is therefore shorter than 60 minutes. The 15–16 bucket is thus the only row in this table that gets less time than above; its resolved column and its n stand visibly beside it so it is not confused with the other hours.
The year by year section reads the same arbiter metric a second time, split by calendar year — the year of the anchor in New York time, and because the anchors sit between 09:30 and 15:55 that is simply the calendar day of the session. Why at all: the market has quirks on an hourly basis and macroeconomic cycles above that; one number over six years hides that something carries in some years and not at all in others. The same logic as “total is irrelevant” across the day, only across the years.
The years are read-off slices, not runs of their own. The cyclic shift of the 8 control arms still runs over the full day list of the role; only afterwards are real arm and arms split by year. The baseline of a year is therefore the median of the same arms on the same anchors, not a freshly shuffled run. Edge years are partial years and carry their span in the name (“2020 (from 29.05.)”), in the chart a * — their n is smaller, their control span wider, and a year with few sessions swings accordingly.
The chart below shows Δ per year as bars (left axis) and z as a diamond (right axis). z stays screening, here even more than usual: with seven years × two instruments × several studies |z| ≥ 2 also falls purely by chance. A striking year is a new thesis and must run again on untouched data — it does not overturn a preregistered verdict.
One hit rate per session, turned into a histogram. The question is not whether the rate is right but whether it is carried by many days or by few. The bars are the ten deciles from 0 % to 100 %, the height is the share of days in that decile; days without a resolved anchor in that hour do not count. The tooltip above the histogram names the median day, the band from 10 % to 90 % and the number of days, the tooltip per bar the days in that decile.
The hour rows hold only about a dozen anchors per day — the daily rate is coarsely stepped there. The New York total row is the robust one: there every day has all 78 anchors.
The signed forward return (FR) is the price path in ticks, turned into signal direction: signal short and price falls 10 ticks means +10. Bracket-free — no stop, no target, nothing that truncates the result — and gross, because costs move every cell equally and cancel out in the difference to the baseline. What is shown is Δ to the baseline; green means above. This is the second question — direction is answered by the hit rate probe.
MFE / MAE — maximum favorable / adverse excursion: how far price ran at most for and against the signal within the horizon. A good MFE with a poor forward return means: the move was there, it just did not hold to the horizon.
A setting that carries in only one half is not a setting but a period (Learnings 26). First the hit rate, then the time probe — both as Δ to the baseline.
Breadth is the share of cells above the baseline — 7 hours × bracket widths or × horizons; beside it stands what the best and the worst control arm achieve on the same computation. For the hit rate a control arm sits at around 50 % by construction. The verdict is never passed on the best cell.
| Instrument of this report | ES |
| Signal | Seeded randomness per anchor (seed 20260813, own series for ES): long or short at 50 % each, deterministic from seed, instrument, day and slot, with no market input at all |
| Sessions read | 1467 |
| Anchors per session | 78 (09:30–15:55, every 5 min) |
| Horizons | 1 / 5 / 15 / 30 / 60 min, bracket horizon 60 min; second bracket run capped at the RTH close 16:00 (context) |
| Bracket widths | 20 / 40 / 80 / 120 / 160 / 200 / 300 / 400 ticks |
| Arbiter | P(+40 before −40) |
| Control arms | 8, seed 20260813, cyclic shift of the sign series across the days |
| Half-spread at entry, measured | 0.51 ticks |
| Reading ticks | 20.9 s |
| Run total | 21.6 s |
| Commit | f1a70e3c099644522bfadfea043aa7c2b6d7cefb |