Honesty doctrine. Every candidate anomaly is an artifact until proven otherwise; in-sample results are never findings; past statistical regularity does not imply future returns. This is research on statistical properties of market data — not investment advice, not a trading system.

Code / experiments/micro/expF_multiple_testing/analysis.md

experiments/micro/expF_multiple_testing/analysis.md 63 lines
---
project: anomaly-atlas
document: expF_multiple_testing/analysis
author: Simon-Pierre Boucher
contact: contact@spboucher.ai
data_source: hfmarketdata.io
created: 2026-08-12
modified: 2026-08-12
status: reviewed
---

# Analysis — expF_multiple_testing

Run: `results/expF_multiple_testing/20260812T072749Z/results.json`.
Battery: 372 signed rules (every cell/pair/class the C/D/E scans searched),
6 blocks, White RC + Hansen SPA (500 stationary bootstraps) + BH-FDR + DSR.
RC/SPA/DSR §8.1-gated first. **0 network requests** — the entire experiment
ran from the frozen cache (the reproducibility anchor doing its job).

## The survival curve (charter result-type E) — gross, frictionless

| layer | survivors | rate |
|---|---|---|
| universe (searched) | 372 | 100 % |
| naive \|t\| > 1.96 | 232 | 62 % |
| BH-FDR 5 % | 226 | 61 % |
| Hansen SPA step-1 | 68 | 18 % |

## The headline is methodological

The SPA survivors carry annualized Sharpes of 10–31 — **physically absurd**,
which is the charter-§12 red flag, and the diagnosis is clean:

1. **Statistical correction corrects for search, not for mechanism.** A
   contrarian rule mechanically earns −autocov₁ > 0 on paper wherever
   bid-ask bounce exists — frictionless, that is "profit"; in reality it is
   the spread you would pay. SPA has no way to know that.
2. **The canary proves it**: −L:SPX→SPY — the index-staleness artifact we
   *know* is fake (T3, measured twice) — survives SPA comfortably.
3. The expE calendar block survives nothing anywhere (naive 0, SPA p 0.87):
   the pre-registered pipeline-broken tripwire did not fire.

Honest validation therefore requires ALL THREE independent layers — artifact
nulls (the scans), search correction (this experiment), and costs (expG) —
and no one of them substitutes for another. This goes into methodology.md
as a design axiom, with this experiment as the demonstration.

## The double-filtered pool (artifact-adjusted ∩ search-corrected, gross)

12 of the 14 expC artifact-adjusted triage cells also clear SPA:
AAPL/ATRO/BKE/HTD/ICUI/MSFT/SLF/XOM 5min, AXDX 30min, CKX 1day, JPM/NVDA
1min. Lead-lag: **−L:ES→SPY** (splice-invariant basis reversion) plus the
expD 2014-2015 FDR residuals; −L:SPX→SPY is excluded from the tradable pool
(routed to candidate 03 as the artifact demonstration). This pool — and
nothing else — proceeds to expG.

## Caveats

In-sample by design (rules evaluated on the data that surfaced them;
validation split untouched until expH). DSR values are reported per block
but the bounce-driven Sharpe heterogeneity inflates `sr0`; treat them as
universe diagnostics, not verdicts.