Nyalai
Contact
The calibration engine for probabilistic systems in quantitative finance

The problem is no longer producing strategies.
It is knowing which deserve to be trusted.

Know, don't believe.

Nyalai returns binary verdicts on probabilistic systems in quantitative finance. Every candidate is submitted to five adversarial validation gates. Every verdict is signed, timestamped, and publicly citable. Institutional model risk management addresses the first three quadrants of decision-making. Nyalai closes the fourth. The methodology is published in full under Creative Commons Zero. The signature is not.

LIVE CALIBRATIONJUDGE PROCESS · CLIENT DELIVERY
Pedagogy through contrast

One binds. One does not.

One system passes. One system does not. Both are examined against the same five gates, under the same Doctrine v0.3.3.1, and both receive a signed audit trail. Nyalai does not grade on a curve. The Doctrine either binds, or it does not.

iPhone displaying a Slack channel named verdicts where the CalibrationJudge bot has posted a NOT-CALIBRATED verdict for Synapse v0.3, with five of five gates failed, BSS negative 0.079, ECE 0.141, and audit trail hash cice/87c619a
Verdict-form 001 · Synapse v0.3NOT-CALIBRATEDDelivered to the institutional workflow
What CICE found
  • Five of five gates failed against a 239-prediction ledger.
  • Brier Skill Score negative 0.079: forecast quality worse than a naive baseline.
  • Expected Calibration Error 0.141 versus threshold 0.05: calibration off by a factor of 2.8.
  • Bootstrap 95 percent confidence interval on Sharpe entirely below zero.
  • Chain of custody signed and archived: audit trail hash cice/87c619a.
What a CALIBRATED verdict shows
  • Five of five gates passed under the same Doctrine v0.3.3.1.
  • Deflated Sharpe Ratio 0.97 above threshold 0.95, Brier Skill Score positive over baseline.
  • Expected Calibration Error 0.031 within tolerance of 0.05.
  • Bootstrap 95 percent confidence interval on Sharpe from 0.024 to 0.118, entirely positive.
  • Trust level MTL-4 Adversarially Validated. Revalidation cadence quarterly.
iPhone displaying a portfolio dashboard illustrating what a CALIBRATED verdict looks like when monitored, with a rising sparkline and an allocation shown within the declared envelope
Reference Predictor R-01 · Illustrative demonstrationCALIBRATEDMonitored in a portfolio dashboard

Illustrative demonstrations. Verdict-form 001 metrics on Synapse v0.3 are MEASURED against a 239-prediction ledger. The CALIBRATED example is constructed for pedagogical clarity and does not represent any specific live verdict. Full audit trail for the NOT-CALIBRATED case at nyalai.com/verdicts/001-synapse-v03.

The industry we operate in

The reproducibility crisis, in four numbers.

Quantitative finance produces published strategies at industrial scale. Very few of them survive an honest out-of-sample test. The statistical inference horizon in financial time series is short relative to the horizon over which the underlying claims need to hold, and Fama and French have documented that even ten-year windows admit substantial probability of misleading realized premia. Nyalai exists because the gap between what is published and what is real has become large enough to be its own asset class, and because internal review cannot solve the horizon problem alone.

0
Published equity risk factors
Documented in top-tier finance journals as of 2016. The majority are estimated to be statistical artifacts of trial selection, not real economic effects.
Harvey, Liu & ZhuAnd the cross-section of expected returns · Review of Financial Studies · 2016
> 0%
Probability of backtest overfitting
Estimated rate at which the best-performing in-sample strategy fails to rank among the top strategies out-of-sample. Standard result on published quant portfolios.
Bailey & López de PradoThe deflated Sharpe ratio · Journal of Portfolio Management · 2014
~$0T
Global systematic AUM
Assets under management by systematic hedge funds, quantitative CTAs, and rules-based ETFs. Every allocation decision inside this figure depends on an implicit claim of predictor calibration.
Preqin Alternative Assets DataHedge funds coverage · 2025
12
Independent binary validators
Number of institutional infrastructures that render a public CALIBRATED or NOT-CALIBRATED verdict on a probabilistic financial system with a reproducible audit trail and CC0 methodology. Until Nyalai.
NyalaiRefusal Doctrine v0.3.3.1 · 2026
Where we operate

Institutional MRM works in the first three quadrants.
Nyalai closes the fourth.

In the taxonomy Nassim Taleb published in 2008, decisions divide along two axes: whether the payoff structure is simple or complex, and whether the underlying randomness lives in thin-tailed or fat-tailed distributions. Quantitative strategies operate in the complex-and-fat-tailed region: nonlinear payoffs, unbounded loss possibilities, regime shifts. SR 11-7 and equivalent model risk management guidance address the first three regions adequately. The fourth is where standard validation methodology fails, and where an external adversarial layer is not an addition. It is a requirement.

Nyalai does not eliminate Fourth Quadrant risk. Nyalai makes the decisions rendered under that risk auditable, and refuses strategies whose evaluation would require thin-tailed assumptions the underlying data does not support. The five gates are the operational form of that discipline. Each gate uses the statistical framework appropriate to its question, following the modular validation principle Joseph Simonian formalized in the Journal of Financial Data Science.

What a Nyalai verdict looks like

Rust refuses. Green approves.

A refused system (our own meta-calibration engine, Synapse v0.3) side by side with an approved system (a reference case). The rust wax seal signals refusal. The green wax seal signals approval. A reader understands the mechanism in five seconds.

Deployment vocabulary

Model Trust Levels.
From MTL-1 Untested to MTL-5 Live.

A shared vocabulary for how much trust to place in a probabilistic prediction system. Numbered ascending. Each level defined by empirical passage of specific gates. Each level attached to a publicly declared deployment context. Strict nesting. Binary per level. Publicly contestable. Time-bounded.

MTL1
Untested
No independent validation performed. No public methodology. No track record.
Pre-production research · Not for capital allocation
MTL2
Nominally Calibrated
Passes Gates 1 and 4 (statistical fragility and probabilistic calibration) on a documented ledger of at least 200 resolved observations.
Advisory signal · Never sole allocation signal
MTL3
Robustly Calibrated
Passes MTL-2 plus cross-regime robustness under a Hidden Markov Model. Failure modes documented publicly.
Multi-signal portfolio · Measured exposure with limits
MTL4
Adversarially Validated
Passes MTL-3 plus guided-search safety (Aronson & Masters 2013) and bootstrap confidence. Sovereignty rule enforced.
Primary allocation signal · Institutional scale
MTL5
Continuously Live-Validated
Passes MTL-4 plus rolling live verdict updates and automatic drift detection. Full audit trail to LPs on request.
Mission-critical live trading · Scaled deployment
Where verdicts live

The Verdicts newsroom.

Every verdict is a long-form artifact. Signed. Versioned. Archived. Publication cadence: bi-weekly starting September 2026. Doctrine documents, method notes, and case studies live alongside.

NOT-CALIBRATED
VerdictJul 1, 2026

Synapse v0.3: NOT-CALIBRATED across all five gates

A meta-calibration engine developed by the same author is submitted to CalibrationJudge on a ledger of 239 resolved predictions. Every gate fails at published thresholds. The first Nyalai Verdict-form is a refusal of our own system. That is the intended outcome of an author who takes the sovereignty rule seriously.