Honesty doctrine. Every candidate anomaly is an artifact until proven otherwise; in-sample results are never findings; past statistical regularity does not imply future returns. This is research on statistical properties of market data — not investment advice, not a trading system.

P001 · version 1 · Tue Aug 11 2026 20:00:00 GMT-0400 (heure avancée de l’Est)

The Artifact Frontier: What Survives Honest Testing in Open High-Frequency Market Data?

Part I — Methods, artifact taxonomy, and train-period results (2000–2016)

Simon-Pierre Boucher · contact@spboucher.ai · data: hfmarketdata.io published

Abstract

We systematically scan open 1-minute-to-daily market data (hfmarketdata.io, sole source) for short-horizon mean-reversion, lead-lag, and calendar anomalies, under a pre-registered protocol: every detector must first pass a synthetic-data gate; every scan runs against measured artifact nulls; hypothesis budgets are declared before testing; and validation layers (multiple-testing correction, transaction costs) are applied in sequence. On the 2000–2016 train split, 62% of the 372 searched rules are naively 'significant' and 18% survive Hansen's SPA — yet the survivors carry physically implausible paper Sharpes, and a deliberately included known artifact (the SPX→SPY 'lead') survives statistical correction unharmed. The cost layer then eliminates essentially everything: the median surviving rule breaks even at 1.1% of one half-spread per trade, and zero intraday rules survive paying the full half-spread. The calendar family — including turn-of-month, the last published survivor — dies against a permuted-calendar null with an 8-test budget. Our principal positive contributions are a measured artifact taxonomy for this dataset and a demonstrated three-layer validation doctrine: artifact nulls, search correction, and costs are independent filters, and no one of them substitutes for another. Out-of-sample confirmation on the untouched validation split is reported in Part II.

The Artifact Frontier — Part I#

Every number in this document regenerates from committed results.json files, and every figure below is rendered live from them by the site. The research log is the audit trail; the charter is the pre-registered protocol.

1. Question and doctrine#

Given only open high-frequency market data — 1-minute to daily bars, no quotes — which statistical regularities are real (reproducible out-of-sample, robust to artifacts, economically nonzero after costs), and which are plumbing? The project's doctrine, stated before any data was touched: every candidate anomaly is an artifact until proven otherwise; in-sample results are never findings; negative results are first-class.

Three design rules operationalize this:

  1. Test the tests. No detector touches real data before passing a synthetic gate: it must find nothing in a pure random walk, recover planted effects, and flag a pure bid-ask-bounce series as artifact. The gate is not ceremonial — it caught two real bugs before they could contaminate results (a variance-ratio estimator reading ~1/q on random walks, and an end-freshness-only synchronization mask that silently attenuated correlations via the Epps mechanism).
  2. Pre-registration. Universes, time splits (train 2000–2016 / validation 2016–2021 / sealed holdout 2022→), hypothesis budgets (22, declared in research_gaps), and every experiment's falsification criterion were frozen in writing before the corresponding run.
  3. A frozen cache as the reproducibility anchor. All data flows through one client that never silently refetches; every response is indexed in a committed data manifest. Two of the seven experiments below ran with zero network requests, entirely from the frozen cache.

2. The data, measured#

Experiment A established empirically what the source actually provides: true 1-minute bars for ~7,700 stocks and ~5,200 ETFs from January 2000 (futures and indices from 2008, FX from 2010, crypto from 2013), plus daily options chains with quotes and Greeks over 67 quarters. Load-bearing facts that the documentation does not state: timestamps are US-Eastern wall-clock, bar-start labeled; bars exist only where trades occurred (no zero-volume placeholders — an illiquid name printed 38 bars in a full session); daily bars carry the official auction close and consolidated volume, both absent from the 1-minute series; and dividend-adjusted prices are re-based to the vendor's build date, so adjusted series are not point-in-time stable. Responses are hard-capped at 50,000 rows. Full profile: data_source_profile.

3. The artifact taxonomy, with magnitudes#

The project's first deliverable is a catalogue of the mechanisms in this dataset that manufacture fake anomalies — each with detection code and a measured magnitude (artifact_taxonomy). Experiment B measured the key ones on a pre-specified 42-ticker universe (Q1 2024, RTH 1-minute):

1min return AC1-0.347000.51staleness (share of RTH minutes without a fresh print)AAPL — staleness 0, 1min return AC1 0.00162MSFT — staleness 0, 1min return AC1 -0.01245NVDA — staleness 0, 1min return AC1 0.01298AMZN — staleness 0, 1min return AC1 -0.00916GOOGL — staleness 0, 1min return AC1 -0.02624META — staleness 0, 1min return AC1 -0.02418TSLA — staleness 0, 1min return AC1 0.01645JPM — staleness 0, 1min return AC1 -0.03379XOM — staleness 0, 1min return AC1 -0.01076UNH — staleness 0.0024, 1min return AC1 -0.05529SPY — staleness 0, 1min return AC1 0.00628QQQ — staleness 0, 1min return AC1 0.01507ATRO — staleness 0.6919, 1min return AC1 -0.05122AUVI — staleness 0.6646, 1min return AC1 -0.21019AXDX — staleness 0.8335, 1min return AC1 -0.32299BKE — staleness 0.288, 1min return AC1 -0.04574CECO — staleness 0.5311, 1min return AC1 -0.16489HPE — staleness 0.0003, 1min return AC1 -0.06183HTD — staleness 0.6906, 1min return AC1 -0.31279HWM — staleness 0.0052, 1min return AC1 -0.02708ICUI — staleness 0.4941, 1min return AC1 -0.10807KEY.K — staleness 0.8674, 1min return AC1 -0.3469KSPI — staleness 0.471, 1min return AC1 -0.20949LENZ — staleness 0.7067, 1min return AC1 -0.27902PBYI — staleness 0.3251, 1min return AC1 -0.11171PLNT — staleness 0.0305, 1min return AC1 -0.01553RITM.B — staleness 0.8984, 1min return AC1 -0.25345RPM — staleness 0.2196, 1min return AC1 -0.0903RUM — staleness 0.0291, 1min return AC1 -0.05627SLF — staleness 0.1826, 1min return AC1 -0.01447SPB — staleness 0.3332, 1min return AC1 -0.04111KEY.K SPY leads +1 min (LOCF join)00.11900.51staleness (share of RTH minutes without a fresh print)AAPL — staleness 0, SPY leads +1 min (LOCF join) 0.01237MSFT — staleness 0, SPY leads +1 min (LOCF join) 0.00645NVDA — staleness 0, SPY leads +1 min (LOCF join) 0.00859AMZN — staleness 0, SPY leads +1 min (LOCF join) 0.0037GOOGL — staleness 0, SPY leads +1 min (LOCF join) -0.00674META — staleness 0, SPY leads +1 min (LOCF join) 0.00367TSLA — staleness 0, SPY leads +1 min (LOCF join) 0.00585JPM — staleness 0, SPY leads +1 min (LOCF join) 0.00377XOM — staleness 0, SPY leads +1 min (LOCF join) -0.00728UNH — staleness 0.0024, SPY leads +1 min (LOCF join) 0.02334QQQ — staleness 0, SPY leads +1 min (LOCF join) 0.01068ATRO — staleness 0.6919, SPY leads +1 min (LOCF join) 0.06282AUVI — staleness 0.6646, SPY leads +1 min (LOCF join) 0.0057AXDX — staleness 0.8335, SPY leads +1 min (LOCF join) -0.00701BKE — staleness 0.288, SPY leads +1 min (LOCF join) 0.08981CECO — staleness 0.5311, SPY leads +1 min (LOCF join) 0.06433HPE — staleness 0.0003, SPY leads +1 min (LOCF join) 0.00666HTD — staleness 0.6906, SPY leads +1 min (LOCF join) 0.05151HWM — staleness 0.0052, SPY leads +1 min (LOCF join) 0.02045ICUI — staleness 0.4941, SPY leads +1 min (LOCF join) 0.0769KEY.K — staleness 0.8674, SPY leads +1 min (LOCF join) 0.02132KSPI — staleness 0.471, SPY leads +1 min (LOCF join) 0.02003LENZ — staleness 0.7067, SPY leads +1 min (LOCF join) 0.0143PBYI — staleness 0.3251, SPY leads +1 min (LOCF join) 0.03636PLNT — staleness 0.0305, SPY leads +1 min (LOCF join) 0.03961RITM.B — staleness 0.8984, SPY leads +1 min (LOCF join) 0.01628RPM — staleness 0.2196, SPY leads +1 min (LOCF join) 0.09142RUM — staleness 0.0291, SPY leads +1 min (LOCF join) 0.05423SLF — staleness 0.1826, SPY leads +1 min (LOCF join) 0.11852SPB — staleness 0.3332, SPY leads +1 min (LOCF join) 0.07764SLF
expB — artifact null levels, one dot per ticker (Q1 2024, RTH 1min). Bounce pushes AC1 negative and LOCF joins make SPY spuriously lead, both in proportion to staleness. Run 20260812T055602Z, regenerated from results.json.

Highlights: bid-ask bounce alone produces AC1 of −0.23 and VR(30) of 0.55 in the stalest liquidity tercile with zero planted economics; LOCF joins make SPY spuriously "lead" mid-staleness names (+0.047 at +1 min, Spearman vs staleness +0.43) while diluting the lead of ultra-stale names — the artifact is non-monotone; and the SPX index print lags SPY by one minute (+0.065 at 0.965 contemporaneous correlation) — a Fisher (1966) effect, measured live. The intraday profile (Experiment E) adds the time-of-day dimension: volatility is U-shaped (6.6 bp at the open, 2.4 midday, 2.9 at the close) while the effective spread declines monotonically (2.8 → 1.2 bp) — any "first-30-minutes" return claim fights 2–3× the midday artifact level.

4. The scans (train split only, Level 0 by construction)#

Mean-reversion (Experiment C). 127 cells (ticker × timeframe × sub-period), each tested against a bounce null and FDR-corrected. The scan's most valuable output was about the null itself: a daily effective spread combined with pure Roll alternation predicts impossible intraday autocorrelations (−3 to −27), because consecutive intraday closes rarely flip sides. The corrected, variance-consistent triage — an MA(1) null that absorbs all lag-1 effects — leaves 14 cells of genuine multi-lag reversion, concentrated in a daily 2008–2015 mega-cap/index family (XOM excess AC1 −0.13, SPY −0.055, both FDR) and a few 1-minute cells (JPM −0.24 VR-excess at the 30-minute horizon).

FDR survivor (excess < −0.05) other scan cells (Level 0) -0.3 -0.2 -0.1 0 0.1VR(30) excess over the MA(1)-consistent null · negative = multi-lag reversion beyond any lag-1 effect1day · 2000-2007AAPL 1day 2000-2007: VR30 1.07, excess 0.149MSFT 1day 2000-2007: VR30 0.8867, excess -0.068NVDA 1day 2000-2007: VR30 1.0224, excess -0.020AMZN 1day 2000-2007: VR30 0.7792, excess -0.240JPM 1day 2000-2007: VR30 0.9767, excess 0.046XOM 1day 2000-2007: VR30 0.5697, excess -0.319UNH 1day 2000-2007: VR30 0.6322, excess -0.393SPY 1day 2000-2007: VR30 0.7125, excess -0.211QQQ 1day 2000-2007: VR30 0.8025, excess -0.157ATRO 1day 2000-2007: VR30 0.7266, excess -0.041AXDX 1day 2000-2007: VR30 0.5734, excess -0.201BKE 1day 2000-2007: VR30 0.8367, excess -0.072CKX 1day 2000-2007: VR30 0.7351, excess -0.276HTD 1day 2000-2007: VR30 0.8951, excess -0.214ICUI 1day 2000-2007: VR30 0.7916, excess -0.149RPM 1day 2000-2007: VR30 0.6788, excess -0.141SLF 1day 2000-2007: VR30 0.5272, excess -0.4651day · 2008-2015AAPL 1day 2008-2015: VR30 1.0634, excess 0.079MSFT 1day 2008-2015: VR30 0.6912, excess -0.192NVDA 1day 2008-2015: VR30 1.0101, excess -0.003AMZN 1day 2008-2015: VR30 0.791, excess -0.169GOOGL 1day 2008-2015: VR30 0.7967, excess -0.366META 1day 2008-2015: VR30 1.1455, excess 0.132TSLA 1day 2008-2015: VR30 0.9938, excess -0.026JPM 1day 2008-2015: VR30 0.503, excess -0.297XOM 1day 2008-2015: VR30 0.4637, excess -0.251UNH 1day 2008-2015: VR30 0.81, excess -0.122SPY 1day 2008-2015: VR30 0.6838, excess -0.164QQQ 1day 2008-2015: VR30 0.8166, excess -0.084ATRO 1day 2008-2015: VR30 1.1814, excess 0.151AXDX 1day 2008-2015: VR30 0.9285, excess -0.056BKE 1day 2008-2015: VR30 0.8767, excess -0.033CECO 1day 2008-2015: VR30 0.7399, excess -0.219CKX 1day 2008-2015: VR30 0.2717, excess -0.259 (FDR survivor)CKXHTD 1day 2008-2015: VR30 1.1662, excess -0.130ICUI 1day 2008-2015: VR30 0.6223, excess -0.383PBYI 1day 2008-2015: VR30 1.0679, excess 0.112RPM 1day 2008-2015: VR30 0.8597, excess -0.062SLF 1day 2008-2015: VR30 0.8347, excess -0.092SPB 1day 2008-2015: VR30 0.884, excess -0.21830min · 2000-2007AAPL 30min 2000-2007: VR30 0.887, excess -0.086MSFT 30min 2000-2007: VR30 0.9381, excess -0.027NVDA 30min 2000-2007: VR30 0.9766, excess -0.061AMZN 30min 2000-2007: VR30 0.9883, excess -0.061JPM 30min 2000-2007: VR30 0.862, excess -0.136XOM 30min 2000-2007: VR30 0.9534, excess -0.013UNH 30min 2000-2007: VR30 0.9746, excess 0.015SPY 30min 2000-2007: VR30 1.0115, excess 0.009QQQ 30min 2000-2007: VR30 0.9476, excess -0.092ATRO 30min 2000-2007: VR30 0.5993, excess -0.029AXDX 30min 2000-2007: VR30 0.6595, excess -0.209 (FDR survivor)BKE 30min 2000-2007: VR30 1.01, excess 0.026HTD 30min 2000-2007: VR30 0.6253, excess -0.000ICUI 30min 2000-2007: VR30 0.77, excess 0.036RPM 30min 2000-2007: VR30 0.9671, excess 0.054SLF 30min 2000-2007: VR30 0.8314, excess 0.00130min · 2008-2015AAPL 30min 2008-2015: VR30 0.8804, excess -0.125MSFT 30min 2008-2015: VR30 0.8706, excess -0.166NVDA 30min 2008-2015: VR30 0.9781, excess -0.007AMZN 30min 2008-2015: VR30 0.8564, excess -0.144GOOGL 30min 2008-2015: VR30 0.9835, excess -0.020META 30min 2008-2015: VR30 0.9636, excess -0.083TSLA 30min 2008-2015: VR30 0.9828, excess 0.029JPM 30min 2008-2015: VR30 0.8974, excess -0.093XOM 30min 2008-2015: VR30 0.7897, excess -0.231UNH 30min 2008-2015: VR30 0.8428, excess -0.132SPY 30min 2008-2015: VR30 0.8984, excess -0.202QQQ 30min 2008-2015: VR30 0.8987, excess -0.166ATRO 30min 2008-2015: VR30 0.8589, excess 0.047AXDX 30min 2008-2015: VR30 0.7418, excess 0.023BKE 30min 2008-2015: VR30 0.7195, excess -0.258 (FDR survivor)BKECECO 30min 2008-2015: VR30 0.8414, excess 0.000HTD 30min 2008-2015: VR30 1.0113, excess 0.029ICUI 30min 2008-2015: VR30 0.895, excess -0.041PBYI 30min 2008-2015: VR30 0.9996, excess 0.023RPM 30min 2008-2015: VR30 0.9851, excess 0.007SLF 30min 2008-2015: VR30 1.0324, excess -0.088SPB 30min 2008-2015: VR30 0.9819, excess 0.0305min · 2000-2007AAPL 5min 2000-2007: VR30 0.8516, excess -0.052 (FDR survivor)MSFT 5min 2000-2007: VR30 0.8371, excess -0.108 (FDR survivor)NVDA 5min 2000-2007: VR30 0.919, excess -0.016AMZN 5min 2000-2007: VR30 0.9954, excess 0.013JPM 5min 2000-2007: VR30 0.968, excess 0.020XOM 5min 2000-2007: VR30 0.837, excess -0.050 (FDR survivor)UNH 5min 2000-2007: VR30 1.0124, excess -0.015SPY 5min 2000-2007: VR30 0.8906, excess 0.002QQQ 5min 2000-2007: VR30 0.9354, excess -0.035ATRO 5min 2000-2007: VR30 0.4872, excess -0.027BKE 5min 2000-2007: VR30 1.0285, excess 0.098HTD 5min 2000-2007: VR30 0.4046, excess -0.089 (FDR survivor)ICUI 5min 2000-2007: VR30 0.5688, excess -0.046RPM 5min 2000-2007: VR30 0.5994, excess -0.023SLF 5min 2000-2007: VR30 0.7249, excess -0.125 (FDR survivor)5min · 2008-2015AAPL 5min 2008-2015: VR30 0.922, excess -0.039MSFT 5min 2008-2015: VR30 0.9346, excess 0.001NVDA 5min 2008-2015: VR30 0.9139, excess -0.021AMZN 5min 2008-2015: VR30 0.9038, excess -0.045GOOGL 5min 2008-2015: VR30 0.9386, excess -0.033META 5min 2008-2015: VR30 0.9537, excess 0.010TSLA 5min 2008-2015: VR30 0.8886, excess -0.048JPM 5min 2008-2015: VR30 0.9202, excess -0.015XOM 5min 2008-2015: VR30 0.8721, excess -0.038UNH 5min 2008-2015: VR30 1.023, excess 0.001SPY 5min 2008-2015: VR30 1.0034, excess 0.048QQQ 5min 2008-2015: VR30 0.9981, excess 0.018ATRO 5min 2008-2015: VR30 0.6477, excess -0.057 (FDR survivor)AXDX 5min 2008-2015: VR30 0.6709, excess 0.006BKE 5min 2008-2015: VR30 0.8455, excess -0.089 (FDR survivor)CECO 5min 2008-2015: VR30 0.6372, excess -0.041HTD 5min 2008-2015: VR30 0.8202, excess 0.092ICUI 5min 2008-2015: VR30 0.7883, excess -0.055 (FDR survivor)PBYI 5min 2008-2015: VR30 0.9018, excess 0.125RPM 5min 2008-2015: VR30 0.8882, excess -0.029SLF 5min 2008-2015: VR30 1.1126, excess 0.100SPB 5min 2008-2015: VR30 0.7901, excess -0.0331min · 2014-2015AAPL 1min 2014-2015: VR30 0.9559, excess -0.017MSFT 1min 2014-2015: VR30 0.8708, excess -0.037NVDA 1min 2014-2015: VR30 0.8274, excess -0.061 (FDR survivor)AMZN 1min 2014-2015: VR30 0.8388, excess -0.047GOOGL 1min 2014-2015: VR30 0.8505, excess -0.014META 1min 2014-2015: VR30 0.8971, excess -0.020TSLA 1min 2014-2015: VR30 0.8903, excess 0.035JPM 1min 2014-2015: VR30 0.5772, excess -0.243 (FDR survivor)JPMXOM 1min 2014-2015: VR30 0.8777, excess -0.075 (FDR survivor)UNH 1min 2014-2015: VR30 1.0417, excess 0.027SPY 1min 2014-2015: VR30 0.9217, excess -0.041QQQ 1min 2014-2015: VR30 1.1045, excess 0.134
expC — multi-lag reversion triage on TRAIN. One dot per ticker-cell; the MA(1)-consistent null absorbs all lag-1 effects (bounce included); values beyond the axis range pile at its edge. Run 20260812T062408Z, regenerated from results.json.

Lead-lag (Experiment D). 49 pairs, two windows, with the raw-LOCF versus both-fresh comparison built in — so the non-synchronicity artifact is measured, not consumed. In 2006–2007, minute-scale market→component diffusion was real on synchronized samples (SPY led every sector ETF by +0.07..+0.15). By 2014–2015 it had collapsed to ±0.05 — the cleanest decay measurement of the project. Two structural facts survive synchronization: a splice-invariant ES↔SPY cross-serial effect (−0.032, identical across all three futures adjustments), and the SPX→SPY "lead" (+0.132) — which survives because synchronizing print times cannot fix a computed index.

raw LOCF join (artifact included) both-fresh (synchronized) 0 0.05cross-correlation at +1 min (leader → follower)ESES[contin_UNadj]->SPY ES[contin_UNadj]->SPY raw LOCF join: -0.01056 ES[contin_UNadj]->SPY both-fresh: -0.01065ES[contin_adj_ratio]->SPY ES[contin_adj_ratio]->SPY raw LOCF join: -0.01057 ES[contin_adj_ratio]->SPY both-fresh: -0.01065ES[contin_adj_absolute]->SPY ES[contin_adj_absolute]->SPY raw LOCF join: -0.01063 ES[contin_adj_absolute]->SPY both-fresh: -0.01072INDEXSPX->SPY SPX->SPY raw LOCF join: -0.01442 SPX->SPY both-fresh: -0.0145LIQUIDSPY->JPM SPY->JPM raw LOCF join: 0.04571 SPY->JPM both-fresh: 0.04577SPY->GOOGL SPY->GOOGL raw LOCF join: 0.02618 SPY->GOOGL both-fresh: 0.02542 (FDR survivor)SPY->UNH SPY->UNH raw LOCF join: 0.02274 SPY->UNH both-fresh: 0.01272SPY->NVDA SPY->NVDA raw LOCF join: 0.0101 SPY->NVDA both-fresh: 0.01006 (FDR survivor)SPY->AMZN SPY->AMZN raw LOCF join: 0.00937 SPY->AMZN both-fresh: 0.00924SPY->TSLA SPY->TSLA raw LOCF join: 0.00639 SPY->TSLA both-fresh: 0.00627SPY->AAPL SPY->AAPL raw LOCF join: -0.00193 SPY->AAPL both-fresh: -0.00195SPY->XOM SPY->XOM raw LOCF join: -0.00421 SPY->XOM both-fresh: -0.00438SPY->META SPY->META raw LOCF join: -0.01308 SPY->META both-fresh: -0.01308 (FDR survivor)SPY->MSFT SPY->MSFT raw LOCF join: -0.0149 SPY->MSFT both-fresh: -0.01506 (FDR survivor)SPY->QQQ SPY->QQQ raw LOCF join: -0.01943 SPY->QQQ both-fresh: -0.0195 (FDR survivor)RANDOMSPY->BKE SPY->BKE raw LOCF join: 0.07889 SPY->BKE both-fresh: 0.07533 (FDR survivor)SPY->ATRO SPY->ATRO raw LOCF join: 0.07129 SPY->ATRO both-fresh: 0.06777 (FDR survivor)SPY->CECO SPY->CECO raw LOCF join: 0.05331 SPY->CECO both-fresh: 0.05896 (FDR survivor)SPY->AXDX SPY->AXDX raw LOCF join: 0.05037 SPY->AXDX both-fresh: 0.06008 (FDR survivor)SPY->HTD SPY->HTD raw LOCF join: 0.0481 SPY->HTD both-fresh: 0.0728 (FDR survivor)SECTORSPY->XLF SPY->XLF raw LOCF join: 0.02823 SPY->XLF both-fresh: 0.02824 (FDR survivor)SPY->XLP SPY->XLP raw LOCF join: 0.02382 SPY->XLP both-fresh: 0.02369 (FDR survivor)SPY->XLK SPY->XLK raw LOCF join: 0.01895 SPY->XLK both-fresh: 0.01889 (FDR survivor)SPY->XLY SPY->XLY raw LOCF join: 0.01656 SPY->XLY both-fresh: 0.01646 (FDR survivor)SPY->XLB SPY->XLB raw LOCF join: 0.01336 SPY->XLB both-fresh: 0.01297SPY->XLU SPY->XLU raw LOCF join: 0.01127 SPY->XLU both-fresh: 0.01121SPY->XLI SPY->XLI raw LOCF join: 0.00997 SPY->XLI both-fresh: 0.00987SPY->XLV SPY->XLV raw LOCF join: -0.0031 SPY->XLV both-fresh: -0.00297SPY->XLE SPY->XLE raw LOCF join: -0.00842 SPY->XLE both-fresh: -0.00847
expD — lead at +1 min per pair, 2014-2015: the gap between the raw join and the both-fresh subsample is the non-synchronicity artifact (T3), measured. Run 20260812T064047Z, regenerated from results.json.
raw LOCF join (artifact included) both-fresh (synchronized) 0 0.05 0.1cross-correlation at +1 min (leader → follower)LIQUIDSPY->JPM SPY->JPM raw LOCF join: 0.03163 SPY->JPM both-fresh: 0.03153 (FDR survivor)SPY->MSFT SPY->MSFT raw LOCF join: 0.03089 SPY->MSFT both-fresh: 0.03096 (FDR survivor)SPY->AMZN SPY->AMZN raw LOCF join: 0.02902 SPY->AMZN both-fresh: 0.02918 (FDR survivor)SPY->NVDA SPY->NVDA raw LOCF join: 0.02496 SPY->NVDA both-fresh: 0.02509 (FDR survivor)SPY->QQQ SPY->QQQ raw LOCF join: 0.01916 SPY->QQQ both-fresh: 0.01913 (FDR survivor)SPY->XOM SPY->XOM raw LOCF join: 0.01823 SPY->XOM both-fresh: 0.01868 (FDR survivor)SPY->UNH SPY->UNH raw LOCF join: 0.01773 SPY->UNH both-fresh: 0.01778 (FDR survivor)SPY->AAPL SPY->AAPL raw LOCF join: 0.01129 SPY->AAPL both-fresh: 0.01135 (FDR survivor)RANDOMSPY->BKE SPY->BKE raw LOCF join: 0.1048 SPY->BKE both-fresh: 0.11961 (FDR survivor)SPY->HTD SPY->HTD raw LOCF join: 0.03871 SPY->HTD both-fresh: 0.09514 (FDR survivor)SPY->ATRO SPY->ATRO raw LOCF join: 0.00641 SPY->ATRO both-fresh: 0.00757SECTORSPY->XLK SPY->XLK raw LOCF join: 0.15039 SPY->XLK both-fresh: 0.14773 (FDR survivor)SPY->XLI SPY->XLI raw LOCF join: 0.14409 SPY->XLI both-fresh: 0.12394 (FDR survivor)SPY->XLP SPY->XLP raw LOCF join: 0.1376 SPY->XLP both-fresh: 0.13013 (FDR survivor)SPY->XLY SPY->XLY raw LOCF join: 0.13493 SPY->XLY both-fresh: 0.12449 (FDR survivor)SPY->XLV SPY->XLV raw LOCF join: 0.12624 SPY->XLV both-fresh: 0.11474 (FDR survivor)SPY->XLB SPY->XLB raw LOCF join: 0.12014 SPY->XLB both-fresh: 0.11817 (FDR survivor)SPY->XLF SPY->XLF raw LOCF join: 0.08449 SPY->XLF both-fresh: 0.08357 (FDR survivor)SPY->XLU SPY->XLU raw LOCF join: 0.07486 SPY->XLU both-fresh: 0.07391 (FDR survivor)SPY->XLE SPY->XLE raw LOCF join: 0.03372 SPY->XLE both-fresh: 0.03377 (FDR survivor)
expD — lead at +1 min per pair, 2006-2007: the gap between the raw join and the both-fresh subsample is the non-synchronicity artifact (T3), measured. Run 20260812T064047Z, regenerated from results.json.

Calendar (Experiment E). Eight pre-declared tests (day-of-week ×5, turn-of-month, pre/post-holiday) on SPY against a within-year permuted-calendar null with a family-wise max-statistic. Nothing survives (best marginal p = 0.24; family-wise p ≥ 0.93 everywhere). Turn-of-month — the last survivor in the published literature as of 2006 — fails and decays inside the train period (+7.9 bp in 2000–2007 → +1.6 bp in 2008–2015). Both pipeline controls behaved: the Monday effect stayed dead, and the volatility U-shape was strongly present.

-20 -10 0 10 20mean daily SPY return in class minus overall mean (bp) · gray bar = permuted-calendar 95% bandmon mon: observed -1.4 bp — marginal p 0.7286, family-wise p 1tue tue: observed 3.965 bp — marginal p 0.2974, family-wise p 0.99wed wed: observed -0.468 bp — marginal p 0.9075, family-wise p 1thu thu: observed 1.501 bp — marginal p 0.7031, family-wise p 1fri fri: observed -3.767 bp — marginal p 0.3368, family-wise p 0.992turn of month turn_of_month: observed 4.734 bp — marginal p 0.2444, family-wise p 0.9735pre holiday pre_holiday: observed 5.723 bp — marginal p 0.5752, family-wise p 0.9275post holiday post_holiday: observed 2.855 bp — marginal p 0.7766, family-wise p 0.9995
expE — all 8 pre-declared calendar tests on SPY (train 2000–2016): every observed effect (dot) sits inside its permuted-calendar 95% band (bar). Nothing survives; the last-survivor turn-of-month included. Run 20260812T065222Z, regenerated from results.json.
0 2 4 609:3010:3011:3012:3013:3014:3015:3009:30 — median |1min return|: 6.586 bp10:00 — median |1min return|: 4.37 bp10:30 — median |1min return|: 3.541 bp11:00 — median |1min return|: 3.162 bp11:30 — median |1min return|: 2.761 bp12:00 — median |1min return|: 2.496 bp12:30 — median |1min return|: 2.425 bp13:00 — median |1min return|: 2.387 bp13:30 — median |1min return|: 2.362 bp14:00 — median |1min return|: 2.503 bp14:30 — median |1min return|: 2.444 bp15:00 — median |1min return|: 2.602 bp15:30 — median |1min return|: 2.88 bp09:30 — median EDGE spread: 2.777 bp10:00 — median EDGE spread: 2.008 bp10:30 — median EDGE spread: 1.348 bp11:00 — median EDGE spread: 1.858 bp11:30 — median EDGE spread: 1.289 bp12:00 — median EDGE spread: 1.401 bp12:30 — median EDGE spread: 2.015 bp13:00 — median EDGE spread: 1.556 bp13:30 — median EDGE spread: 1.663 bp14:00 — median EDGE spread: 1.631 bp14:30 — median EDGE spread: 1.714 bp15:00 — median EDGE spread: 1.311 bp15:30 — median EDGE spread: 1.212 bp median |1min return| (bp) median EDGE spread (bp) half-hour bucket (RTH) · liquid 12, 1min, 2014–2015
expE / H20 — the intraday artifact profile (taxonomy input): volatility is U-shaped (6.6 bp at the open, 2.4 midday, 2.9 at the close); the spread declines monotonically (2.8 → 1.2 bp). Any "first-30-minutes" return claim faces 2–3× the midday artifact level. Run 20260812T065222Z.

5. The survival curve: statistical correction is not artifact correction#

Experiment F pushed everything the scans searched — 372 signed rules — through the correction battery: naive t-tests, Benjamini–Hochberg FDR, White's Reality Check and Hansen's SPA over stationary bootstraps, and the Deflated Sharpe Ratio.

searched universe all scanned cells/pairs/classes (×2 signs) searched universe: 372 rules (100%) 372 (100%)naive |t| > 1.96 uncorrected in-sample t-test naive |t| > 1.96: 232 rules (62%) 232 (62%)BH-FDR 5% false-discovery-rate correction BH-FDR 5%: 226 rules (61%) 226 (61%)Hansen SPA step-1 data-snooping correction Hansen SPA step-1: 68 rules (18%) 68 (18%)Survivors are gross and artifact-laden — costs (expG) are the next layer.
expF — the survival curve: what fraction of the searched rule universe survives each statistical-correction layer on TRAIN. Statistical correction fixes the search, not the mechanism. Run 20260812T065907Z, regenerated from results.json.

The result that matters is not the 18% SPA survival rate — it is what survives: rules with paper Sharpes of 10–31 annualized, physically implausible, dominated by bounce harvesting (a contrarian rule mechanically earns −autocov₁ on paper, which is precisely the spread it would pay in reality). The canary proves the point: the SPX→SPY rule — an artifact we had already measured twice — passes SPA comfortably. Statistical correction corrects for search; it is structurally blind to mechanism.

6. The cost frontier closes the loop#

Experiment G re-priced the double-filtered pool (rules that beat both the artifact nulls and the search correction; 31 rules) under a declared cost model: net = gross − κ · (EDGE half-spread) · turnover, κ swept from 0 to 2.

0.001 0.01 0.1 1 patient execution full half-spreadbreakeven cost multiplier κ* (× half-spread paid per trade, log scale) — right of a line = survives that cost levelR:CKX:1day:2008-2015 R:CKX:1day:2008-2015 — κ* = 0.2799 · gross 39.101 bp/day · turnover 1.12/day · half-spread 124.379 bpR:HTD:5min:2000-2007 R:HTD:5min:2000-2007 — κ* = 0.1923 · gross 218.132 bp/day · turnover 59.83/day · half-spread 18.956 bpR:AXDX:30min:2000-2007 R:AXDX:30min:2000-2007 — κ* = 0.1392 · gross 55.091 bp/day · turnover 3.62/day · half-spread 109.435 bpR:ICUI:5min:2008-2015 R:ICUI:5min:2008-2015 — κ* = 0.041 · gross 79.286 bp/day · turnover 70.4/day · half-spread 27.475 bpR:ATRO:5min:2008-2015 R:ATRO:5min:2008-2015 — κ* = 0.0363 · gross 133.29 bp/day · turnover 52.95/day · half-spread 69.313 bpL:SPY->ATRO:+ L:SPY->ATRO:+ — κ* = 0.0354 · gross 158.647 bp/day · turnover 154.01/day · half-spread 29.117 bpL:SPY->HTD:+ L:SPY->HTD:+ — κ* = 0.0342 · gross 24.123 bp/day · turnover 76.5/day · half-spread 9.211 bpL:SPY->BKE:+ L:SPY->BKE:+ — κ* = 0.0329 · gross 147.618 bp/day · turnover 237.6/day · half-spread 18.869 bpL:SPY->CECO:+ L:SPY->CECO:+ — κ* = 0.0317 · gross 108.571 bp/day · turnover 104.98/day · half-spread 32.629 bpL:SPY->AXDX:+ L:SPY->AXDX:+ — κ* = 0.0272 · gross 192.889 bp/day · turnover 140.98/day · half-spread 50.247 bpR:SLF:5min:2000-2007 R:SLF:5min:2000-2007 — κ* = 0.0268 · gross 37.148 bp/day · turnover 71.63/day · half-spread 19.381 bpR:MSFT:5min:2000-2007 R:MSFT:5min:2000-2007 — κ* = 0.025 · gross 39.148 bp/day · turnover 78.87/day · half-spread 19.871 bpR:BKE:5min:2008-2015 R:BKE:5min:2008-2015 — κ* = 0.0222 · gross 61.794 bp/day · turnover 79.81/day · half-spread 34.805 bpR:XOM:5min:2000-2007 R:XOM:5min:2000-2007 — κ* = 0.0182 · gross 34.032 bp/day · turnover 77.74/day · half-spread 24.048 bpR:AAPL:5min:2000-2007 R:AAPL:5min:2000-2007 — κ* = 0.0179 · gross 68.121 bp/day · turnover 79.21/day · half-spread 48.128 bpR:NVDA:1min:2014-2015 R:NVDA:1min:2014-2015 — κ* = 0.0114 · gross 117.735 bp/day · turnover 378.76/day · half-spread 27.231 bpL:SPY->XLY:+ L:SPY->XLY:+ — κ* = 0.0094 · gross 22.837 bp/day · turnover 397.98/day · half-spread 6.098 bpL:SPY->GOOGL:+ L:SPY->GOOGL:+ — κ* = 0.0067 · gross 45.636 bp/day · turnover 393.02/day · half-spread 17.413 bpL:SPY->XLP:+ L:SPY->XLP:+ — κ* = 0.0066 · gross 30.825 bp/day · turnover 395.87/day · half-spread 11.817 bpL:SPY->XLF:+ L:SPY->XLF:+ — κ* = 0.0054 · gross 52.66 bp/day · turnover 399.91/day · half-spread 24.571 bpL:SPY->NVDA:+ L:SPY->NVDA:+ — κ* = 0.0043 · gross 46.829 bp/day · turnover 398.42/day · half-spread 27.231 bpL:SPY->XLK:+ L:SPY->XLK:+ — κ* = 0.0039 · gross 37.642 bp/day · turnover 397.88/day · half-spread 24.072 bpR:JPM:1min:2014-2015 R:JPM:1min:2014-2015 — κ* = 0.0038 · gross 39.337 bp/day · turnover 393.27/day · half-spread 26.034 bpL:SPY->META:- L:SPY->META:- — κ* = 0.0037 · gross 24.615 bp/day · turnover 400.35/day · half-spread 16.839 bpL:ES[contin_adj_ratio]->SPY:- L:ES[contin_adj_ratio]->SPY:- — κ* = 0.0028 · gross 11.34 bp/day · turnover 395.96/day · half-spread 10.376 bpL:ES[contin_adj_absolute]->SPY:- L:ES[contin_adj_absolute]->SPY:- — κ* = 0.0028 · gross 11.34 bp/day · turnover 395.96/day · half-spread 10.376 bpL:ES[contin_UNadj]->SPY:- L:ES[contin_UNadj]->SPY:- — κ* = 0.0028 · gross 11.34 bp/day · turnover 395.96/day · half-spread 10.376 bpL:SPY->MSFT:- L:SPY->MSFT:- — κ* = 0.0026 · gross 18.179 bp/day · turnover 400.11/day · half-spread 17.211 bpL:SPY->QQQ:- L:SPY->QQQ:- — κ* = 0.0025 · gross 16.081 bp/day · turnover 400.33/day · half-spread 16.063 bp
expG — the cost frontier: every rule that beat the artifact nulls AND the search correction dies when it must pay a fraction of its own half-spread (median κ* = 0.0114). Run 20260812T072128Z, regenerated from results.json.

The median rule breaks even at κ* = 0.011 — it captures about 1% of one half-spread per trade. Three rules survive κ = 0.1 (all in sparse names), one survives κ = 0.25 (CKX, an ultra-sparse name with a wide, noisy spread estimate — the classic profile of an estimation artifact, forwarded to the validation split with a skeptical prior rather than discarded by hand), and zero rules — none — survive paying the full half-spread. The pre-registered falsification clause ("the costs-kill story fails if any intraday rule survives κ = 1") did not trigger.

7. What Part I establishes#

  1. A three-layer validation doctrine, demonstrated rather than asserted. Artifact nulls, search correction, and transaction costs filter different failure modes; each layer passed things the next one killed.
  2. A measured artifact taxonomy (T1–T7) for open bar data, with the magnitudes above and neutralization rules, validated on synthetic ground truth.
  3. Negative results with teeth: the calendar family is empty under an honest budget; minute-scale lead-lag decayed an order of magnitude between 2006 and 2015; and nothing in the searched universe pays for its own spread on the train split.
  4. Methodological findings: spread-based bounce nulls must be variance-consistent with the target series; synchronization on print times cannot de-artifact computed indices; LOCF distortion of lead-lag is non-monotone in staleness.

8. Limitations#

All Part-I numbers are train-split, in-sample by design — their out-of-sample fate on the untouched 2016–2021 validation split is Part II. Bar data carries no quotes: costs are estimated (EDGE), not observed. Holiday-class calendar tests are low-powered (n = 144). The universe is U.S.-equity-centric; crypto and FX hypotheses await volume-semantics verification. Capacity is out of scope.

9. Reproducibility#

Everything regenerates from the repository: commit f59e891 (data profile, client), 9b6beea (artifact baselines), c4977f2 (reversion scan), a7f66cb (lead-lag scan), b228862 (calendar scan), 7a82cc1 (survival battery), 881399d (cost frontier). Each experiment's results.json embeds the hardware manifest and the client's instrumentation; the data manifest (data_manifest/index.jsonl) indexes every API response consumed. Detector gates: benchmarks/synthetic/ (29 tests at the time of writing).

References#

Key sources (full annotated list with DOIs: bibliography): Roll (1984); Fisher (1966); Scholes & Williams (1977); Epps (1979); Lo & MacKinlay (1988, 1990); Sullivan, Timmermann & White (2001); White (2000); Hansen (2005); Bailey & López de Prado (2014); Harvey, Liu & Zhu (2016); McLean & Pontiff (2016); Novy-Marx & Velikov (2016); Chordia, Roll & Subrahmanyam (2005); Ardia, Guidotti & Kroencke (2024); Chen & Velikov (2022).