Honesty doctrine. Every candidate anomaly is an artifact until proven otherwise;
in-sample results are never findings; past statistical regularity does not imply future returns.
This is research on statistical properties of market data — not investment advice, not a trading system.
Independent statistics + market-microstructure research · Apple Silicon · hfmarketdata.io
Which market "anomalies" are real — and which are artifacts?
A systematic, reproducible atlas of statistical regularities in open high-frequency
market data — mean-reversion, lead-lag, calendar effects — where every claim survives (or visibly fails)
out-of-sample testing, multiple-comparison correction, artifact nulls, and realistic transaction costs.
Negative results are first-class findings.
detectable in-sample ≠ reproducible out-of-sample ≠ robust to artifacts ≠ meaningful after costs
58sources reviewed
0hypotheses registered
2/11experiments completed
0atlas findings (Level ≥ 1)
The confidence taxonomy
A finding only advances one level at a time, and only Level ≥ 1 is ever published.
A large in-sample effect with zero out-of-sample survival is a negative result —
published as one.
Level 0 — in-sample onlyscan output; never published as a finding
Level 1 — corrected & OOSsurvives multiple-testing correction and a clean out-of-sample split, artifact null subtracted
Level 2 — robust+ robust to specification choices, sub-periods, instruments
Level 3 — cost-real & held-out+ economically nonzero after realistic costs, confirmed on the once-touched holdout
Phase 1 complete: literature sweep (52 verified sources)
Question. What does the literature establish about Q1-Q3 anomalies, the
statistics of not fooling yourself, microstructure artifacts, and costs?
Method. Three parallel verification passes against OpenAlex (every
citation confirmed: title, authors, year, venue, DOI; access date
2026-08-12), plus discovery searches for 2005-2025 decay/replication work.
expB complete: artifact nulls measured (after the §8.1 gate caught a real bug)
Question. Are the bounce/staleness/non-synchronicity artifacts measurable
and material at 1min on this data?
Experiment. Built generators with known ground truth + first real stats
modules (reversion, leadlag, bootstrap, artifacts). The mandatory synthetic
gate (11 tests) CAUGHT A REAL BUG before any real data was touched: the
Phase 0.5 complete: Experiment A (data reality check)
Question. What does hfmarketdata.io actually provide (granularity,
depth, timestamps, adjustments, completeness, limits)?
Experiment. expA_data_reality — 8 probe families through the newly
implemented single client (hf_client.py: cache-first, throttled, back-off,
data-manifest index; 5 offline unit tests pass). 54 network requests,