Experiments / expF_multiple_testing
expF_multiple_testing
Multiple-testing survival: White RC, SPA, FDR, deflated Sharpe on C-E output
View benchmark implementation (benchmark.py) →
Figures
Generated server-side from the latest committed results.json — never hand-typed.
Hypothesis
Hypothesis — expF_multiple_testing#
Pre-specified 2026-08-12 before the battery ran. RC/SPA/DSR passed the §8.1 gate first (4 tests: pure noise never survives; a planted profitable rule always does).
Hypothesis
The survival curve (charter result-type E): after honest correction for
the FULL searched universe, few or none of the C/D scan leads survive.
Prior: the daily 2008-2015 reversal family and possibly ES->SPY have the
best odds; the tiny 2014-2015 lead-lag residuals and most intraday
reversion cells should die.
Falsification criterion
Not falsifiable as a directional claim — the DELIVERABLE is the measured
survival rate at each layer. The pipeline is broken (investigate, not
publish) if a rule from the expE calendar family survives SPA (expE
already showed all 8 inside the permutation band).
Artifact null(s)
The searched-universe null itself: RC/SPA bootstrap under H0 "no rule
beats zero", universe = EVERYTHING the scans looked at (not only the
FDR survivors) — 2 signed variants of every cell/pair/class.
Method (pre-declared)
Rule construction (mechanical, no tuning):
R-family (expC, 127 cells x2 signs): contrarian rule at the cell's
timeframe on RTH trade-time returns, pos_t = -sign(r_{t-1});
1day cells use the previous daily return. Daily aggregation; days
without data = 0 (idle).
L-family (expD, 49 pairs x2): follower timed by leader's previous
1min return on both-fresh minutes, daily aggregation.
C-family (expE, 8 classes x2): +/-(r_t - unconditional mean) on class
days, 0 elsewhere (drift-adjusted so "long Mondays" cannot free-ride
the equity premium).
Blocks (common day calendars): expC 2000-2007, expC 2008-2015, expC 1min
2014-2015, expD 2006-2007, expD 2014-2015, expE train.
Battery per block: (1) naive |t|>1.96 count; (2) BH-FDR on two-sided
rule p-values across ALL blocks jointly; (3) Hansen SPA (500 stationary
bootstraps, mean block 5 days, seed 42) + StepM-style step-1 survivor
count (rule t >= bootstrap max-stat 95th pct); (4) DSR of each block's
best rule, n_trials = total universe size, sr_variance across the
universe. White RC reported alongside SPA.
IMPORTANT honesty note: rules are evaluated on the SAME train data the
scans ran on — expF measures survival of the in-sample search under
correction. Out-of-sample survival is expH's job on the validation split.
Result
Run 20260812T072749Z — 0 network requests (582 cache hits; the frozen
cache carried the whole battery). Universe: 372 signed rules, 6 blocks.
Survival funnel (GROSS, frictionless): naive |t|>1.96 = 232 (62%) ->
BH-FDR = 226 (61%) -> SPA step-1 = 68 (18%). All expC/expD blocks reject
at the bootstrap floor (RC and SPA p = 0.002); the expE calendar block
survives nothing (naive 0, spa_p 0.87) — the pipeline-broken tripwire did
NOT fire. Best rules carry annualized Sharpe 10-31 — physically absurd,
the §12 red flag. Intersection with the artifact-adjusted expC triage:
12/14 cells also pass SPA. L-family step-1 includes -L:ES->SPY (the
splice-invariant basis effect) and -L:SPX->SPY (the KNOWN index-staleness
artifact, deliberately kept in the universe as a canary — it survives
statistical correction, which proves the point below).
Interpretation
(Level 0.) The survival curve's headline is METHODOLOGICAL and it is the
strongest result of the project so far: statistical correction corrects
for SEARCH, not for MECHANISM. 18% of gross rules survive Hansen SPA —
and the survivors are dominated by bounce harvesting (a contrarian rule
earns -autocov1 > 0 on paper and pays the spread in reality) plus
frictionless lead-lag timing; the known artifact (SPX->SPY) sails through
SPA unharmed. Honest validation therefore REQUIRES all three layers:
artifact nulls (scans) AND search correction (expF) AND costs (expG).
The double-filtered pool going to expG: 12 reversion cells + ES->SPY +
the expD 2014-2015 FDR set. DSR by block is reported but is mostly a
universe-heterogeneity diagnostic here (bounce-inflated Sharpe variance);
documented, not over-read.
Next experiment
expG (cost frontier) on the double-filtered pool; expH (validation split)
for whatever survives costs.Analysis
Analysis — expF_multiple_testing#
Run: results/expF_multiple_testing/20260812T072749Z/results.json.
Battery: 372 signed rules (every cell/pair/class the C/D/E scans searched),
6 blocks, White RC + Hansen SPA (500 stationary bootstraps) + BH-FDR + DSR.
RC/SPA/DSR §8.1-gated first. 0 network requests — the entire experiment
ran from the frozen cache (the reproducibility anchor doing its job).
The survival curve (charter result-type E) — gross, frictionless#
| layer | survivors | rate |
|---|---|---|
| universe (searched) | 372 | 100 % |
| naive |t| > 1.96 | 232 | 62 % |
| BH-FDR 5 % | 226 | 61 % |
| Hansen SPA step-1 | 68 | 18 % |
The headline is methodological#
The SPA survivors carry annualized Sharpes of 10–31 — physically absurd, which is the charter-§12 red flag, and the diagnosis is clean:
- Statistical correction corrects for search, not for mechanism. A contrarian rule mechanically earns −autocov₁ > 0 on paper wherever bid-ask bounce exists — frictionless, that is "profit"; in reality it is the spread you would pay. SPA has no way to know that.
- The canary proves it: −L:SPX→SPY — the index-staleness artifact we know is fake (T3, measured twice) — survives SPA comfortably.
- The expE calendar block survives nothing anywhere (naive 0, SPA p 0.87): the pre-registered pipeline-broken tripwire did not fire.
Honest validation therefore requires ALL THREE independent layers — artifact nulls (the scans), search correction (this experiment), and costs (expG) — and no one of them substitutes for another. This goes into methodology.md as a design axiom, with this experiment as the demonstration.
The double-filtered pool (artifact-adjusted ∩ search-corrected, gross)#
12 of the 14 expC artifact-adjusted triage cells also clear SPA: AAPL/ATRO/BKE/HTD/ICUI/MSFT/SLF/XOM 5min, AXDX 30min, CKX 1day, JPM/NVDA 1min. Lead-lag: −L:ES→SPY (splice-invariant basis reversion) plus the expD 2014-2015 FDR residuals; −L:SPX→SPY is excluded from the tradable pool (routed to candidate 03 as the artifact demonstration). This pool — and nothing else — proceeds to expG.
Caveats#
In-sample by design (rules evaluated on the data that surfaced them;
validation split untouched until expH). DSR values are reported per block
but the bounce-driven Sharpe heterogeneity inflates sr0; treat them as
universe diagnostics, not verdicts.
README
expF_multiple_testing#
Multiple-testing survival: White RC, SPA, FDR, deflated Sharpe on C-E output
Status: completed 2026-08-12 — survival curve measured: 372 -> 232 (naive) -> 226 (FDR) -> 68 (SPA, 18%), all gross; survivors dominated by bounce; the known SPX artifact survives SPA (statistical correction != artifact correction). Double-filtered pool -> expG.
Result runs
- 20260812T065907Z / results.json 7.8 KiB