Honesty doctrine. Every candidate anomaly is an artifact until proven otherwise; in-sample results are never findings; past statistical regularity does not imply future returns. This is research on statistical properties of market data — not investment advice, not a trading system.

Research / research/state_of_the_art.md

State of the art

reviewed

State of the art · created Tue Aug 11 2026 20:00:00 GMT-0400 (heure avancée de l’Est) · updated Tue Aug 11 2026 20:00:00 GMT-0400 (heure avancée de l’Est) · Simon-Pierre Boucher

State of the art (Phase 2)#

Critical map of each anomaly family and method, in the charter §5 template. Sources: research/bibliography.md (52 verified references, accessed 2026-08-12); dataset facts: research/data_source_profile.md; measured artifact levels: research/artifact_taxonomy.md (T1–T7).

Epistemic status legend: robust / decayed / disputed / likely-artifact.


A1 — Short-horizon individual-stock reversal#

  • What it claims. Individual stock returns revert at daily–monthly horizons (Lehmann 1990 weekly; Jegadeesh 1990 monthly).
  • Granularity required. Daily suffices; our 1min adds the ability to separate close-convention effects. ✅ testable.
  • Known artifact confounds. Bid-ask bounce (T1 — Blume–Stambaugh showed it halves such effects), stale prices (T2), auction-close mismatch (T4).
  • Decay evidence. McLean–Pontiff −58 % post-publication; Chordia et al. 2014 attenuation with liquidity. Largely gone in liquid U.S. names.
  • Correct test. Cross-sectional reversal portfolios with bounce-robust prices + VR/AC1 net of the liquidity-bucket bounce null; block-bootstrap CIs (assumes stationarity within blocks).
  • Multiple-testing exposure. Moderate: horizon × universe × weighting grid. Pre-specify or FDR-correct.
  • Cost sensitivity. Extreme — highest-turnover class; Novy-Marx–Velikov prior: dies net of costs.
  • Open-source impl. Our own stats/reversion.py (gated §8.1); arch's VarianceRatio as cross-check — both build on macOS arm64. ✅
  • Main limitation. Without quote data, bounce correction is estimated, not measured.
  • Honest new test here. A 2000–2026 decay curve of daily reversal net of the measured bounce null, by liquidity bucket — a decay re-measurement, not a discovery claim. Status: decayed (gross); likely-artifact (net).

A2 — Index/portfolio variance-ratio momentum#

  • Claims. Weekly index returns positively autocorrelated, VR(q) > 1 (Lo–MacKinlay 1988).
  • Granularity. Daily/weekly from our 1day bars (2000→) and intradaily aggregation. ✅
  • Confounds. Fisher stale-constituent effect (T2/T3) inflated early index autocorrelation; largely gone in ETF prices (SPY trades fresh).
  • Decay. The classic effect faded post-1990s; on ETFs (traded prices, not stale indices) it was always weaker.
  • Correct test. Lo–MacKinlay VR with heteroskedasticity-robust CIs / block bootstrap; on BOTH the index (SPX) and the ETF (SPY) — divergence measures the Fisher artifact directly.
  • MT exposure. Low if q-grid pre-specified (q ∈ {2,5,10,30}).
  • Cost sensitivity. n/a as stated (it's a statistical property claim).
  • Impl. Ours + arch. ✅
  • Limitation. Regime breaks (2008, 2020) dominate long windows — Bai–Perron sub-periods mandatory.
  • Honest new test. SPX-vs-SPY VR divergence as a quantified Fisher artifact 2008–2026 — methodological contribution. Status: decayed; index-level residual = likely-artifact.

A3 — Lead-lag: large caps → small caps#

  • Claims. Returns of large stocks lead small stocks (Lo–MacKinlay 1990); the source of "contrarian" profits.
  • Granularity. Daily and 1min both usable. ✅
  • Confounds. Non-synchronous trading (T3) — THE canonical confound (Scholes–Williams); our expB measured SPY spuriously leading stale names +0.047 at 1min.
  • Decay. Chordia–Roll–Subrahmanyam: minute-scale predictability arbitraged within 5–60 min by 2005; expect near-zero today in fresh pairs.
  • Correct test. Lagged cross-correlation/Granger ONLY on both-fresh subsamples, against the staleness-matched null (expB machinery); Epps-aware at 1min.
  • MT exposure. High (pairs explosion) — pre-specify a small pair set.
  • Cost sensitivity. Extreme for any tradable interpretation.
  • Impl. Ours (stats/leadlag.py, gated). ✅
  • Limitation. No trade timestamps within the bar; sub-minute lead-lag invisible.
  • Honest new test. Decay curve of large→small lead-lag 2000–2026 net of the staleness null — with the artifact share reported alongside the total. Status: decayed (fresh pairs); the textbook effect is largely T3 artifact in modern data.

A4 — Futures/ETF/index lead-lag (price-discovery ordering)#

  • Claims. Futures (ES) lead cash ETFs (SPY) which lead the index print (SPX) at minute scale.
  • Granularity. 1min is coarse for this (the true lead is seconds) but the ordering may still be detectable. ⚠️ marginal.
  • Confounds. T3/T7 (session semantics, index staleness — expB measured SPX lagging SPY +0.065); futures splice choice (3 variants — testable).
  • Decay. At seconds-scale this is permanent structure; at 1min it may be fully arbitraged/invisible.
  • Correct test. Both-fresh 1min xcorr ES↔SPY with staleness null; robustness across the three futures adjustment variants.
  • MT exposure. Low (one pre-specified triple).
  • Cost sensitivity. n/a (structural claim, not a strategy).
  • Impl. Ours. ✅
  • Limitation. 1min floor; ES data starts 2008.
  • Honest new test. Is ANY ES→SPY lead detectable at 1min after the staleness null, and is SPX→anything pure artifact? Status: robust at sub-second (literature); unknown at 1min on open data — genuine gap.

A5 — Cross-asset information flow: crypto ↔ crypto-exposed equities#

  • Claims. (Thin literature.) 24/7 crypto prices embed information that equity prices can only reflect at the next open.
  • Granularity. 1min crypto (24/7) + equity opens. ✅ — this is a structural granularity advantage of our dataset.
  • Confounds. Overnight-gap conventions (T4), selection of "exposed" equities (must be pre-specified), regime dependence (crypto-equity beta varies).
  • Decay. Unknown — modern, underexplored on open data.
  • Correct test. Does BTC's Friday-close→Monday-preopen return predict the Monday opening gap of pre-specified crypto-exposed equities, vs a placebo set and a permuted-weekend null?
  • MT exposure. Low if the equity set and horizon are pre-registered.
  • Cost sensitivity. Open-auction execution is costly; report the frontier.
  • Impl. Ours. ✅
  • Limitation. Short joint history (crypto-exposed equities mostly 2018→); few independent weekends (~400).
  • Honest new test. Exactly the above — one of the few places our data can ask something not already answered. Status: unknown/genuine gap.

A6 — Weekend / Monday effect#

  • Claims. Negative Monday returns (French 1980).
  • Granularity. Daily. ✅
  • Confounds. Close conventions (T4); DST weeks (T7).
  • Decay. The cleanest corpse: gone post-publication (Schwert 2003; Marquering et al. 2006).
  • Correct test. Day-of-week means with permuted-calendar null + SPA against the full day-of-week universe (STW 2001 protocol).
  • MT exposure. High by construction — the calendar space.
  • Cost sensitivity. Any exploitation is high-turnover.
  • Impl. Ours. ✅
  • Honest new test. Re-confirmation of absence on 2000–2026 open data, published as a negative control for the calendar pipeline. Status: decayed.

A7 — Turn-of-month#

  • Claims. Returns concentrate around month boundaries (Ariel 1987; Lakonishok–Smidt 1988).
  • Granularity. Daily. ✅
  • Confounds. Month-boundary volume/flows are real mechanics (pension/401k flows) — a mechanism, not an artifact; but overlap with OpEx week and quarter-ends must be disentangled.
  • Decay. The last survivor as of Marquering et al. 2006. Post-2006 behavior on open data = open question.
  • Correct test. Pre-specified window (−1..+3 trading days), permuted- calendar null, SPA vs the full window universe, sub-period stability.
  • MT exposure. Moderate — window choice is the researcher degree of freedom; pre-register ONE window.
  • Cost sensitivity. Low-frequency (12×/year) — the rare calendar effect that could survive costs if real.
  • Impl. Ours. ✅
  • Honest new test. Did the last survivor survive 2006–2026? Status: disputed — the most interesting calendar re-test.

A8 — Intraday momentum (first → last half-hour)#

  • Claims. First half-hour market return predicts last half-hour (Gao et al. 2018); mechanism: gamma hedging (Baltussen et al. 2021).
  • Granularity. 1min/30min SPY. ✅ perfect fit.
  • Confounds. Overnight-gap inclusion choice; T4 close convention; spread seasonality (U-shape) at both ends of the day.
  • Decay. Published 2018 — post-publication window (2018–2026) is exactly what open data can measure now.
  • Correct test. Pre-registered replication (their exact spec) + OOS post-2018 sample + gamma-state split using our options chains; DSR for the spec search.
  • MT exposure. Low if the published spec is frozen.
  • Cost sensitivity. Two trades/day at the most liquid instrument's most liquid hours — survivable in principle; measure.
  • Impl. Ours. ✅
  • Honest new test. The cleanest possible decay measurement: published effect, published spec, untouched post-publication data. Status: disputed (post-2018 fate unknown).

A9 — Intraday U-shape (open/close vol & spread concentration)#

  • Claims. Volatility, volume, spreads peak at open and close (Wood et al. 1985).
  • Granularity. 1min. ✅
  • Confounds. None — this one is real microstructure, and it is itself a confounder for other intraday claims.
  • Decay. Robust across decades.
  • Correct test. Descriptive profile with bootstrap bands.
  • Cost sensitivity. n/a (input to the cost model, not a strategy).
  • Honest contribution. Measure the intraday profile of OUR bounce null (taxonomy open item) so expE can subtract it. Status: robust — use as positive control + cost-model input.

A10 — Overnight vs intraday return split#

  • Claims. Equity returns accrue disproportionately overnight.
  • Granularity. Daily open/close (+1min for convention checks). ✅
  • Confounds. T4 is central: auction close vs last bar changes overnight returns mechanically; stale opens for illiquid names.
  • Decay. Persistent in the literature but convention-sensitive — disputed as economics vs plumbing.
  • Correct test. Recompute under BOTH close conventions and both open definitions (first 1min bar vs daily open field); effect must survive all four.
  • MT exposure. Low.
  • Cost sensitivity. High (daily turnover).
  • Honest new test. Quantify how much of the overnight premium is convention-dependent on this dataset. Status: disputed.

Methods inventory (with macOS arm64 status)#

Method Use Implementation arm64
Lo–MacKinlay VR + block bootstrap Q1 scans ours (§8.1-gated) + arch.unitroot.VarianceRatio cross-check
Lo (1991) modified R/S long-memory claims to implement, gate on synthetic long-memory
Lagged xcorr / Granger Q2 ours + statsmodels grangercausalitytests
Roll / Corwin–Schultz / CHL / EDGE spreads cost model, T1 null ours (Roll); bidask package (EDGE, pure numpy) + own CS/CHL
Moving-block / stationary bootstrap all CIs ours (Künsch); Politis–Romano to add for RC/SPA
White RC / Hansen SPA / Romano–Wolf StepM expF no maintained OSS — implement ourselves, gate on synthetic ✅ (numpy)
Benjamini–Hochberg FDR scan triage trivial; statsmodels.stats.multitest
Deflated Sharpe / PBO-CSCV expF/expH implement from Bailey–López de Prado formulas
Bai–Perron breaks regime robustness statsmodels (partial) / own dynamic-programming impl

Cross-cutting conclusion. On modern liquid U.S. equities, the honest prior for every gross short-horizon effect is decay toward zero, and for every net effect, death by costs. The genuine opportunities on THIS data are: (i) decay re-measurements with artifact shares reported (A1–A3, A8), (ii) the structural questions our dataset uniquely reaches (A4 at 1min, A5 cross-asset 24/7), (iii) the last-survivor calendar re-test (A7), and (iv) methodological artifact quantifications (A2 Fisher share, A10 convention share) — all publishable regardless of sign.