Honesty doctrine. Every candidate anomaly is an artifact until proven otherwise;
in-sample results are never findings; past statistical regularity does not imply future returns.
This is research on statistical properties of market data — not investment advice, not a trading system.
P001 · version 1 · Tue Aug 11 2026 20:00:00 GMT-0400 (heure avancée de l’Est)
The Artifact Frontier: What Survives Honest Testing in Open High-Frequency Market Data?
Part I — Methods, artifact taxonomy, and train-period results (2000–2016)
We systematically scan open 1-minute-to-daily market data (hfmarketdata.io, sole source) for short-horizon mean-reversion, lead-lag, and calendar anomalies, under a pre-registered protocol: every detector must first pass a synthetic-data gate; every scan runs against measured artifact nulls; hypothesis budgets are declared before testing; and validation layers (multiple-testing correction, transaction costs) are applied in sequence. On the 2000–2016 train split, 62% of the 372 searched rules are naively 'significant' and 18% survive Hansen's SPA — yet the survivors carry physically implausible paper Sharpes, and a deliberately included known artifact (the SPX→SPY 'lead') survives statistical correction unharmed. The cost layer then eliminates essentially everything: the median surviving rule breaks even at 1.1% of one half-spread per trade, and zero intraday rules survive paying the full half-spread. The calendar family — including turn-of-month, the last published survivor — dies against a permuted-calendar null with an 8-test budget. Our principal positive contributions are a measured artifact taxonomy for this dataset and a demonstrated three-layer validation doctrine: artifact nulls, search correction, and costs are independent filters, and no one of them substitutes for another. Out-of-sample confirmation on the untouched validation split is reported in Part II.
Every number in this document regenerates from committed results.json
files, and every figure below is rendered live from them by the site. The
research log is the audit trail; the
charter is the pre-registered protocol.
Given only open high-frequency market data — 1-minute to daily bars, no
quotes — which statistical regularities are real (reproducible
out-of-sample, robust to artifacts, economically nonzero after costs), and
which are plumbing? The project's doctrine, stated before any data was
touched: every candidate anomaly is an artifact until proven otherwise;
in-sample results are never findings; negative results are first-class.
Three design rules operationalize this:
Test the tests. No detector touches real data before passing a
synthetic gate: it must find nothing in a pure random walk, recover
planted effects, and flag a pure bid-ask-bounce series as artifact.
The gate is not ceremonial — it caught two real bugs before they could
contaminate results (a variance-ratio estimator reading ~1/q on random
walks, and an end-freshness-only synchronization mask that silently
attenuated correlations via the Epps mechanism).
Pre-registration. Universes, time splits (train 2000–2016 /
validation 2016–2021 / sealed holdout 2022→), hypothesis budgets (22,
declared in research_gaps), and every
experiment's falsification criterion were frozen in writing before the
corresponding run.
A frozen cache as the reproducibility anchor. All data flows through
one client that never silently refetches; every response is indexed in a
committed data manifest. Two of the seven experiments below ran with
zero network requests, entirely from the frozen cache.
Experiment A established empirically what the source actually provides:
true 1-minute bars for ~7,700 stocks and ~5,200 ETFs from January 2000
(futures and indices from 2008, FX from 2010, crypto from 2013), plus daily
options chains with quotes and Greeks over 67 quarters. Load-bearing facts
that the documentation does not state: timestamps are US-Eastern wall-clock,
bar-start labeled; bars exist only where trades occurred (no zero-volume
placeholders — an illiquid name printed 38 bars in a full session); daily
bars carry the official auction close and consolidated volume, both absent
from the 1-minute series; and dividend-adjusted prices are re-based to the
vendor's build date, so adjusted series are not point-in-time stable.
Responses are hard-capped at 50,000 rows. Full profile:
data_source_profile.
The project's first deliverable is a catalogue of the mechanisms in this
dataset that manufacture fake anomalies — each with detection code and a
measured magnitude (artifact_taxonomy).
Experiment B measured the key ones on a pre-specified 42-ticker universe
(Q1 2024, RTH 1-minute):
expB — artifact null levels, one dot per ticker (Q1 2024, RTH 1min). Bounce pushes AC1 negative and LOCF joins make SPY spuriously lead, both in proportion to staleness. Run 20260812T055602Z, regenerated from results.json.
Highlights: bid-ask bounce alone produces AC1 of −0.23 and VR(30) of 0.55
in the stalest liquidity tercile with zero planted economics; LOCF joins
make SPY spuriously "lead" mid-staleness names (+0.047 at +1 min, Spearman
vs staleness +0.43) while diluting the lead of ultra-stale names — the
artifact is non-monotone; and the SPX index print lags SPY by one minute
(+0.065 at 0.965 contemporaneous correlation) — a Fisher (1966) effect,
measured live. The intraday profile (Experiment E) adds the time-of-day
dimension: volatility is U-shaped (6.6 bp at the open, 2.4 midday, 2.9 at
the close) while the effective spread declines monotonically (2.8 → 1.2 bp)
— any "first-30-minutes" return claim fights 2–3× the midday artifact
level.
4. The scans (train split only, Level 0 by construction)#
Mean-reversion (Experiment C). 127 cells (ticker × timeframe ×
sub-period), each tested against a bounce null and FDR-corrected. The
scan's most valuable output was about the null itself: a daily effective
spread combined with pure Roll alternation predicts impossible intraday
autocorrelations (−3 to −27), because consecutive intraday closes rarely
flip sides. The corrected, variance-consistent triage — an MA(1) null that
absorbs all lag-1 effects — leaves 14 cells of genuine multi-lag
reversion, concentrated in a daily 2008–2015 mega-cap/index family
(XOM excess AC1 −0.13, SPY −0.055, both FDR) and a few 1-minute cells
(JPM −0.24 VR-excess at the 30-minute horizon).
expC — multi-lag reversion triage on TRAIN. One dot per ticker-cell; the MA(1)-consistent null absorbs all lag-1 effects (bounce included); values beyond the axis range pile at its edge. Run 20260812T062408Z, regenerated from results.json.
Lead-lag (Experiment D). 49 pairs, two windows, with the raw-LOCF
versus both-fresh comparison built in — so the non-synchronicity artifact
is measured, not consumed. In 2006–2007, minute-scale market→component
diffusion was real on synchronized samples (SPY led every sector ETF by
+0.07..+0.15). By 2014–2015 it had collapsed to ±0.05 — the cleanest decay
measurement of the project. Two structural facts survive synchronization:
a splice-invariant ES↔SPY cross-serial effect (−0.032, identical across
all three futures adjustments), and the SPX→SPY "lead" (+0.132) — which
survives because synchronizing print times cannot fix a computed index.
expD — lead at +1 min per pair, 2014-2015: the gap between the raw join and the both-fresh subsample is the non-synchronicity artifact (T3), measured. Run 20260812T064047Z, regenerated from results.json.expD — lead at +1 min per pair, 2006-2007: the gap between the raw join and the both-fresh subsample is the non-synchronicity artifact (T3), measured. Run 20260812T064047Z, regenerated from results.json.
Calendar (Experiment E). Eight pre-declared tests (day-of-week ×5,
turn-of-month, pre/post-holiday) on SPY against a within-year
permuted-calendar null with a family-wise max-statistic. Nothing
survives (best marginal p = 0.24; family-wise p ≥ 0.93 everywhere).
Turn-of-month — the last survivor in the published literature as of 2006 —
fails and decays inside the train period (+7.9 bp in 2000–2007 → +1.6 bp
in 2008–2015). Both pipeline controls behaved: the Monday effect stayed
dead, and the volatility U-shape was strongly present.
expE — all 8 pre-declared calendar tests on SPY (train 2000–2016): every observed effect (dot) sits inside its permuted-calendar 95% band (bar). Nothing survives; the last-survivor turn-of-month included. Run 20260812T065222Z, regenerated from results.json.expE / H20 — the intraday artifact profile (taxonomy input): volatility is U-shaped (6.6 bp at the open, 2.4 midday, 2.9 at the close); the spread declines monotonically (2.8 → 1.2 bp). Any "first-30-minutes" return claim faces 2–3× the midday artifact level. Run 20260812T065222Z.
5. The survival curve: statistical correction is not artifact correction#
Experiment F pushed everything the scans searched — 372 signed rules —
through the correction battery: naive t-tests, Benjamini–Hochberg FDR,
White's Reality Check and Hansen's SPA over stationary bootstraps, and the
Deflated Sharpe Ratio.
expF — the survival curve: what fraction of the searched rule universe survives each statistical-correction layer on TRAIN. Statistical correction fixes the search, not the mechanism. Run 20260812T065907Z, regenerated from results.json.
The result that matters is not the 18% SPA survival rate — it is what
survives: rules with paper Sharpes of 10–31 annualized, physically
implausible, dominated by bounce harvesting (a contrarian rule mechanically
earns −autocov₁ on paper, which is precisely the spread it would pay in
reality). The canary proves the point: the SPX→SPY rule — an artifact we
had already measured twice — passes SPA comfortably. Statistical
correction corrects for search; it is structurally blind to mechanism.
Experiment G re-priced the double-filtered pool (rules that beat both the
artifact nulls and the search correction; 31 rules) under a declared cost
model: net = gross − κ · (EDGE half-spread) · turnover, κ swept from 0 to 2.
expG — the cost frontier: every rule that beat the artifact nulls AND the search correction dies when it must pay a fraction of its own half-spread (median κ* = 0.0114). Run 20260812T072128Z, regenerated from results.json.
The median rule breaks even at κ* = 0.011 — it captures about 1% of
one half-spread per trade. Three rules survive κ = 0.1 (all in sparse
names), one survives κ = 0.25 (CKX, an ultra-sparse name with a wide, noisy
spread estimate — the classic profile of an estimation artifact, forwarded
to the validation split with a skeptical prior rather than discarded by
hand), and zero rules — none — survive paying the full half-spread.
The pre-registered falsification clause ("the costs-kill story fails if any
intraday rule survives κ = 1") did not trigger.
A three-layer validation doctrine, demonstrated rather than asserted.
Artifact nulls, search correction, and transaction costs filter
different failure modes; each layer passed things the next one killed.
A measured artifact taxonomy (T1–T7) for open bar data, with the
magnitudes above and neutralization rules, validated on synthetic ground
truth.
Negative results with teeth: the calendar family is empty under an
honest budget; minute-scale lead-lag decayed an order of magnitude
between 2006 and 2015; and nothing in the searched universe pays for its
own spread on the train split.
Methodological findings: spread-based bounce nulls must be
variance-consistent with the target series; synchronization on print
times cannot de-artifact computed indices; LOCF distortion of lead-lag
is non-monotone in staleness.
All Part-I numbers are train-split, in-sample by design — their
out-of-sample fate on the untouched 2016–2021 validation split is Part II.
Bar data carries no quotes: costs are estimated (EDGE), not observed.
Holiday-class calendar tests are low-powered (n = 144). The universe is
U.S.-equity-centric; crypto and FX hypotheses await volume-semantics
verification. Capacity is out of scope.
Everything regenerates from the repository: commit f59e891 (data profile,
client), 9b6beea (artifact baselines), c4977f2 (reversion scan),
a7f66cb (lead-lag scan), b228862 (calendar scan), 7a82cc1 (survival
battery), 881399d (cost frontier). Each experiment's results.json
embeds the hardware manifest and the client's instrumentation; the data
manifest (data_manifest/index.jsonl) indexes every API response consumed.
Detector gates: benchmarks/synthetic/ (29 tests at the time of writing).