Validation infrastructure / the governed spine

THE TOOLKIT

A sleeve's P&L is a claim; the toolkit is the cross-examination. One pipeline — luck, overfitting, specification, friction, multiplicity, allocation — that every sleeve must survive before capital. It is what turns 113 repositories into a handful of confirmed edges.

One pipeline

THE GOVERNED SPINE.

A sleeve emits a raw signal; the toolkit intercepts it and runs it down one fixed ladder. Promotion to the allocator requires surviving every layer. Each catches a different way a backtest lies, and the graveyard is full because the standard is visible rejection, not survivorship.

STRESSRegime filter / inversion wrapper — acts on the negative test, not just measures it.Regime
FRICTIONNon-linear market impact + real borrow turn gross returns into honest net.Execution
BLINDTime-shift null — the edge must beat its own randomized self.Luck
MIRAGESpecification audit — does the alpha keep its sign under every control?Spec
FDR / LEDGERCorpus-wide multiple-testing — which survivors are real after 113 shots.Multiplicity
BREAKTHROUGHRobustness-shrinkage allocation — lucky and fragile alpha is starved.Capital
The toolkit, pointed at its own corpus

113 sleeves in. Two confirmed edges out.

The flagship result is what the pipeline does to BLAQUE BAUX itself. Run on the real recorded returns of the whole corpus, the governed spine cuts the keeper list to what survives multiplicity, specification, and factor decomposition — the verdict no single sleeve's README can give.

113repositories put through the gate
+0.93expected max Sharpe by luck alone
4survive family-wise FDR (q<0.10)
2confirmed diversifying edges
1.8effective bets in the keeper book

Across the ~44 independent sleeve tests, the expected maximum Sharpe by luck alone is +0.93 — any single sleeve under roughly 0.9 is indistinguishable from the best-of-luck draw. Ranked on raw Sharpe, about a dozen "survive" — but that list ranks beta: documented nulls whose high number is a raw factor (high-beta, quality), and managed-beta sleeves at correlation ≈ 1. Ranked on edge over benchmark, only four clear Benjamini-Hochberg, and two of those are not yet specification-audited — leaving two confirmed, spec-robust, non-beta edges. The sleeve map seals it: the keeper book is 1.8 effective bets, 68% factor-R², and ~0 residual alpha — a "diversified" book that is mostly market, low-vol, and momentum beta wearing many names. The breakthrough layer is built to act on exactly this: shrink the lucky and factor-fragile sleeves toward zero and size the few real residual edges. Headline Sharpe falls; honest, out-of-sample-survivable Sharpe rises.

Read the full corpus audit →
The layers / 8

EACH LAYER CATCHES A DIFFERENT LIE.

Every layer is a runnable module. The Python prototype lives in nullbar; the canonical implementation is ported into the Julia base engine (module_14), where the friction, borrow, ADV, and point-in-time data actually live. Stdlib-only, no heavy dependencies.

L1BLINDnullbar.gate / time-shift null + Effective NLuck: Shifts entry times to build the distribution of returns a random-entry twin would earn; an edge must beat its own randomized self. Effective N counts independent trades via overlap-adjusted clustering, so a few bets dressed as many cannot inflate significance.● BUILT L2VALIDATIONPBO-CSCV / Deflated Sharpe / edge decayOverfitting: Combinatorial probability of backtest overfitting, a Sharpe deflated for the number of trials and the returns' own skew and kurtosis, and the in-sample to out-of-sample decay that exposes a curve-fit before it reaches capital.● BUILT L3MIRAGEspecification audit / Leamer extreme boundsSpecification: Enumerates every control-variable subset and reports whether the alpha keeps its sign and significance. A premium that flips under a plausible control — as the catastrophe-premium claim did — was factor beta in disguise, not a distinct edge.● BUILT L4FRICTIONAlmgren-Chriss sqrt-impact + borrow + capacity curveExecution: Flat basis-points is a lie. Square-root market impact scaled by ADV participation and volatility, plus state-dependent short borrow, re-grade thin keepers and give every sleeve a break-even AUM — a 0.6-Sharpe microstructure edge may hold $5M, not $500M.● BUILT L5LEDGER / FDRpre-registration + Benjamini-Hochberg + corpus DSRMultiplicity: Deflated Sharpe governs the trials within a sleeve; nothing governs the 113 across them. An append-only ledger records every shot taken, and corpus-wide false-discovery control finds which "keepers" are real after counting them all.● BUILT L6BREAKTHROUGHrobustness-shrinkage allocatorCapital: Standard allocators treat historical Sharpe as expected Sharpe — the flaw that funds luck. This shrinks each sleeve's expected return by its Blind p-value, Effective N, and mirage verdict before risk-parity, so lucky and spec-fragile alpha is starved and capital concentrates in the proven.● BUILT L7SLEEVEMAPcorrelation map + factor-exposureConcentration: The effective number of bets and a factor decomposition prove whether a "diversified" book is secretly one trade. On the real keeper book it returns 1.8 effective bets and 68% factor-R² — allocate on the residual, collapse correlated clusters to one risk budget.● BUILT L8STRESSregime filter / inversion wrapperRegime: Implements the negative tests — filter-only, invert, asymmetric — so a regime hypothesis is acted on, not merely measured. It intercepts the raw signal before the optimizer without rewriting a sleeve's alpha logic.● BUILT
The scorecard underneath — Jarque-Bera normality, Jensen's alpha and M² instead of Sharpe alone, skew, excess kurtosis, and crisis-correlation — is applied before any mean-variance claim. No language model sits in the order path: every allocator emits a target book, and reproducible code routes it through the gate.
Where it lives

TWO IMPLEMENTATIONS, ONE SPINE.

The toolkit is open. The prototype is a readable Python reference; the system of record is the Julia engine the live spine already runs on.

See the corpus it governs →