Can we tell at 09:30 whether the RTH closes above its open?
Criterion not met — BSS -0.8 % [-1.5 ; -0.0], AUC 47.4 %, accuracy 52.3 % against a base rate of 54.4 %, n = 916 days pooled forward.
Criterion: BSS > 0 and accuracy ≥ 55 %, pooled forward across 2023–2026
| Decision time | 09:30 New York — bars up to t−1 complete, the opening print and every tick of session t before 09:30 |
| Target | rc − ro > 0 |
| Data window | 2020-05-29 to 2026-08-14 · 1607 daily bars · last day 2026-08-14 |
| Walk-forward | Dev ≤ 2022 (definitions only), test 2023 / 2024 / 2025 / 2026, expanding with a yearly refit |
| Model | L2 logistic regression on 15 preregistered set C features, standardised on the training window; C from inner validation (the last training year). Features: open_below_va, open_above_va, open_pos_va, open_dist_vah_atr, open_dist_val_atr, gap_atr, on_range_atr, on_delta_norm, open_pos_on, on_high_vs_rh1_atr, on_low_vs_rl1_atr, open_vs_on_vwap_atr, on_ext_recency, prev_ln_relr_rth, prev_box20_pos |
| Holdout | The PX holdout from 2026-05-01 concerns the PX theses on the tick store. This study is a walk-forward over 2023 to 2026 on daily data — explicitly approved by the user for regime rounds 1 and 2. Every test year is computed from data before it only; no fit ever sees its own test slice. |
Every row is a forecast over the same days. Base rate, majority and persistence are the baselines; the model has to beat them, otherwise it does not count.
| Row | Year | n | Base Rate % | Acc % | AUC % | Brier | BSS % |
|---|---|---|---|---|---|---|---|
| Base Rate | 2023 | 252 | 57.5 | 57.5 | 50.0 | 0.2466 | 0.0 |
| Majority | 2023 | 252 | 57.5 | 57.5 | — | — | — |
| Persistence | 2023 | 252 | 57.5 | 57.5 | 52.1 | 0.2463 | 0.1 |
| Always long | 2023 | 252 | 57.5 | 57.5 | — | — | — |
| Model (LogReg) | 2023 | 252 | 57.5 | 51.2 | 48.1 | 0.2491 | -1.0 |
| Base Rate | 2024 | 255 | 52.5 | 52.5 | 50.0 | 0.2496 | 0.0 |
| Majority | 2024 | 255 | 52.5 | 52.5 | — | — | — |
| Persistence | 2024 | 255 | 52.5 | 52.5 | 49.1 | 0.2498 | -0.1 |
| Always long | 2024 | 255 | 52.5 | 52.5 | — | — | — |
| Model (LogReg) | 2024 | 255 | 52.5 | 51.8 | 49.0 | 0.2507 | -0.4 |
| Base Rate | 2025 | 252 | 55.6 | 55.6 | 50.0 | 0.2472 | 0.0 |
| Majority | 2025 | 252 | 55.6 | 55.6 | — | — | — |
| Persistence | 2025 | 252 | 55.6 | 55.6 | 49.8 | 0.2472 | 0.0 |
| Always long | 2025 | 252 | 55.6 | 55.6 | — | — | — |
| Model (LogReg) | 2025 | 252 | 55.6 | 56.0 | 48.5 | 0.2494 | -0.9 |
| Base Rate | 2026 | 157 | 50.3 | 50.3 | 50.0 | 0.2514 | 0.0 |
| Majority | 2026 | 157 | 50.3 | 50.3 | — | — | — |
| Persistence | 2026 | 157 | 50.3 | 50.3 | 46.3 | 0.2516 | -0.1 |
| Always long | 2026 | 157 | 50.3 | 50.3 | — | — | — |
| Model (LogReg) | 2026 | 157 | 50.3 | 49.0 | 46.5 | 0.2538 | -0.9 |
| Base Rate | pooled | 916 | 54.4 | 54.4 | 47.5 | 0.2484 | 0.0 |
| Majority | pooled | 916 | 54.4 | 54.4 | — | — | — |
| Persistence | pooled | 916 | 54.4 | 54.4 | 47.1 | 0.2484 | 0.0 |
| Always long | pooled | 916 | 54.4 | 54.4 | — | — | — |
| Model (LogReg) | pooled | 916 | 54.4 | 52.3 | 47.4 | 0.2504 | -0.8 |
Block bootstrap (20-day blocks, 1000 runs, pooled forward): BSS 95 % interval [-1.5 ; -0.0].
Confusion matrix pooled: true 0 correct 40 of 418, true 1 correct 439 of 498 (recall 0 = 9.6 %, recall 1 = 88.2 %).
Regularisation chosen per refit — 2023: C = 0.01, n_train = 588 · 2024: C = 0.01, n_train = 840 · 2025: C = 0.01, n_train = 1095 · 2026: C = 0.01, n_train = 1347.
| Decile | n | mean p % | observed % | difference |
|---|---|---|---|---|
| 1 | 92 | 47.6 | 59.8 | 12.2 |
| 2 | 92 | 50.7 | 48.9 | -1.8 |
| 3 | 92 | 51.6 | 52.2 | 0.6 |
| 4 | 92 | 52.4 | 55.4 | 3.0 |
| 5 | 92 | 53.3 | 72.8 | 19.5 |
| 6 | 92 | 54.1 | 50.0 | -4.1 |
| 7 | 91 | 54.8 | 57.1 | 2.3 |
| 8 | 91 | 55.6 | 48.4 | -7.3 |
| 9 | 91 | 56.9 | 42.9 | -14.0 |
| 10 | 91 | 59.7 | 56.0 | -3.6 |
Sharpness: p ≥ 0.7 on 0.0 % of the days (n = 0), hitting — % there · p ≤ 0.3 on 0.0 % (n = 0), hitting — % there.
| Base Rate | the unconditional frequency in the training window, and at the same time the constant comparison forecast. |
| Baseline | the number a model has to beat. Three of them here: base rate, majority (always the more frequent class) and persistence. |
| Persistence | the forecast “today like yesterday” — the value of the same target on the previous day, turned into a rate. |
| BSS | Brier skill score: what percentage of the base rate constant's Brier score the model saves. 0 means equally good, negative means worse. |
| AUC | the probability that a random positive day is scored above a random negative one. 50 is a coin flip. |
| Brier | the mean squared error of the probability. Smaller is better. |
| Sharpness | on what percentage of the days the model says something clear (p ≥ 0.7 or ≤ 0.3) and how often it hits there. A calibrated model without sharpness is useless. |
| Value Area | the price band [rval, rvah] in which 70 % of the RTH volume traded. Always the previous day's here. |
| Trend Day | an RTH with a body ratio ≥ 0.6 — the body of the daily candle fills at least 60 % of its range. |
| Body Ratio | br = |rc − ro| / (rh − rl), computed on the RTH session. |
| relr | RTH range divided by the median of the 20 previous RTH ranges. relr ≥ 1 means: today is bigger than the typical one of the last four weeks. |
| Walk-Forward | every test year is computed from a model that has only seen data before it; the training window grows with every year. No fit sees its own test. |
The numbers in this report are recomputed on every publish — from data/daily-bars-*.json and data/open-features-*.json, which the extractor builds from the tick store.
Run: 1607 daily bars, 0.5 s.