Nyalai
Methodology · Doctrine v0.3.3.1Contact

Methodology · one page

Five adversarial gates. One binary verdict.

Nyalai returns a CALIBRATED or NOT-CALIBRATED verdict on every probabilistic system submitted for validation. Each gate uses the statistical framework appropriate to its specific question. A strategy that clears all five is Adversarially Validated. A strategy that fails one is refused.

The problem

Institutional MRM addresses the first three quadrants. Nyalai closes the fourth.

In the taxonomy Nassim Taleb published in 2008, decisions divide along two axes: whether the payoff structure is simple or complex, and whether the underlying randomness lives in thin-tailed or fat-tailed distributions. Quantitative strategies operate in the complex-and-fat-tailed region: nonlinear payoffs, unbounded loss possibilities, regime shifts. SR 11-7 and equivalent model risk management guidance address the first three regions adequately. The fourth is where standard validation methodology fails, and where an external adversarial layer is not an addition. It is a requirement.

The five gates

Each gate uses the statistical framework appropriate to its question.

Modular validation methodology, as Joseph Simonian framed it in the Journal of Financial Data Science in 2020. Each gate below cites the peer-reviewed methodology it operationalizes. Every threshold has a stated rationale. Every deviation from the standard method is flagged in the verdict-form.

Gate 1
Statistical fragility
Test
Deflated Sharpe Ratio for return-generating strategies (Gate 1a, Bailey and López de Prado 2014), or Brier Skill Score for probabilistic prediction systems (Gate 1b, Brier 1950 as operationalized under Murphy 1973). See Doctrine §9.1.
Threshold
DSR ≥ 0.95 (Gate 1a) or BSS ≥ 0.05 (Gate 1b), per Doctrine §9.1

Detects strategies whose observed edge does not survive correction for multiple testing, non-normality of returns, and length of the observation window.

Gate 2
Cross-regime robustness
Test
Hidden Markov Model regime detection with per-regime performance decomposition
Threshold
Sharpe > 0 in every detected regime

Detects strategies whose profitability is concentrated in a single market state and would deteriorate under an unencountered regime.

Gate 3
Guided-search safety
Test
Guided-search-safe permutation test (Aronson and Masters, 2013)
Threshold
p < 0.05 after correction for the disclosed number of configurations tested

Detects strategies whose significance was inflated by directed exploration during development, the failure mode that machine-learning workflows produce by default.

Gate 4
Probabilistic calibration
Test
Brier Skill Score against naïve baseline; Expected Calibration Error
Threshold
BSS ≥ 0.05, ECE ≤ 0.05

Detects strategies whose stated confidence estimates diverge from realized outcomes. A signal that is directionally correct but miscalibrated breaks position sizing.

Gate 5
Bootstrap confidence
Test
Block bootstrap resampling (Politis and Romano, 1994)
Threshold
95% confidence interval lower bound > 0, net of costs

Detects strategies whose expected edge is positive but whose variance is large enough that the strategy may not remain profitable under realistic path variations.

Deployment vocabulary

Model Trust Levels. From MTL-1 Untested to MTL-5 Live.

A shared vocabulary for how much trust to place in a probabilistic prediction system. Numbered ascending intensity. Each level defined by empirical passage of specific gates. Strict nesting. Binary per level. Publicly contestable. Time-bounded.

MTL1
Untested
No independent validation performed. No public methodology. No track record.
Pre-production research · Not for capital allocation
MTL2
Nominally Calibrated
Passes Gate 1 (Gate 1a DSR or Gate 1b BSS, per §9.1) and Gate 4 on a documented ledger of at least 200 resolved observations.
Advisory signal · Never sole allocation signal
MTL3
Robustly Calibrated
Passes MTL-2 plus Gate 2 cross-regime robustness. Failure modes documented publicly.
Multi-signal portfolio · Measured exposure with limits, quarterly revalidation
MTL4
Adversarially Validated
Passes MTL-3 plus Gate 3 guided-search safety and Gate 5 bootstrap confidence. Sovereignty rule seven enforced: validator and strategy author structurally distinct.
Primary allocation signal · Institutional scale, biannual revalidation, annual public report
MTL5
Continuously Live-Validated
Passes MTL-4 plus rolling live verdict updates. Automatic drift detection and forced revalidation on regime change.
Mission-critical live trading · Scaled deployment under continuous supervision

Every MTL attribution is publicly contestable through the audit trail supplied with the verdict-form. Any independent third party may reproduce the computation and challenge the attribution.

What allocators receive

A verdict-form structured as a medical report.

Findings first, methodology available on request, references cited for audit. Two layers: the upper layer states the verdict in institutional vocabulary; the lower layer supplies the technical evidence. The upper layer never overstates what the lower supports.

Verdict scope statement

Regime scope, sizing envelope, portfolio-construction assumptions. Deployment outside the envelope invalidates the verdict. A verdict on a signal is not a verdict on the portfolio built from that signal.

Triangulation panel

Risk category times business context times regulatory frame. Adapted from the three-lens framework Senthil Kumar established while Chief Risk Officer at BNY Mellon. A verdict against one triangulation cell does not extend to another.

Communication layer

Upper allocator vocabulary (shortfall, drawdown, liquidity, embarrassment, per Matt Bank's Four Horsemen framework at GEM). Lower technical evidence (DSR, BSS, ECE, bootstrap CIs). Both required.

Citation guidance

The verdict-form is designed to be citable by model owners, independent validators, and internal audit across all three lines of institutional defense. Framework, not model, following Roland Stamm at Acadia.

Reproducibility appendix

Doctrine version applied, ʼCɩcɛ validator commit hash, statistical package versions, random seeds, bootstrap block sizes, HMM state counts, data provenance. Anyone can re-run.

Style drift trigger

Any material change in features, instrument universe, risk parameters, or claimed problem domain invalidates the previous verdict. Re-validation required, not optional.

Input integrity

Sequence commitment.

Nyalai never requests permission to publish a case, or to use it in any downstream decision, before a verdict is rendered. Consent for publication or downstream reference is collected after the verdict is rendered, on a document the submitter reviews. Refusal of downstream use does not alter the verdict itself.

1. Submission is frozen first. Timestamp and hash are recorded before any conversation about downstream use begins.

2. Gates run without stake exposure. The strategy author is never told, during evaluation, what downstream consequences might attach to the verdict.

3. Consent comes after the verdict. Publication, private client disclosure, aggregate research, or external reference are all conditioned on consent collected after the verdict is fixed.

This is not a mental-state promise. It is an order-of-operations commitment that Nyalai controls and any third party can audit. Per Refusal Doctrine section 12.4. A discipline of validity, not virtue.

The full methodology is public.

Every gate, every threshold, every methodological choice is documented in the Refusal Doctrine v0.3.3.1. Twelve sections plus acknowledgments plus primary-source bibliography. Published under CC0 1.0. Anyone can reimplement it, contest it, or diverge from it.

Trust is not a result of validation.
It is the product of the rigor with which we accept to contradict ourselves.

Methodology · Doctrine v0.3.3.1 · July 19, 2026 · Sebastien Assohou · Nyalai

[email protected]

This methodology page is published under CC0 1.0 Universal Public Domain Dedication.

Back