Nyalai
Refusal Doctrine v0.3.3.1Contact

Refusal Doctrine · v0.3.3.1

The Refusal Doctrine

Nyalai charter for validated probabilistic systems in quantitative finance. Refusal-first discipline. Five adversarial validation gates. Published under CC0.

Version v0.3.3.1Draft, July 19, 2026Author: Sebastien Assohou

Preface

This document is written for probabilistic systems undergoing validation, not for allocators reading marketing.

Its primary audience is the strategy itself, the artifact whose survival on live capital depends on the honesty with which it has been evaluated. Its secondary audience is the human operator responsible for that strategy: the quantitative researcher who built it, the portfolio manager considering deploying it, the risk officer sanctioning its use, the allocator writing the check.

We publish this Doctrine in the belief that a validation infrastructure is only credible if it can be inspected, contested, and improved by anyone who cares to do so. The methodology described here is licensed under CC0. Anyone may use it, adapt it, extend it. Nothing about Nyalai's commercial existence depends on secrecy about the method.

What Nyalai sells, and what this Doctrine cannot substitute for, is the discipline of applying the method rigorously and publishing the results, including our own failures.

We expect this document to be wrong in places. We expect our understanding to deepen. We commit to revising this Doctrine as evidence accumulates and to acknowledging openly where earlier versions were mistaken.

1. Nyalai and its mission

1.1 What Nyalai is

Nyalai is a calibration engine for probabilistic systems in quantitative finance. We render binary verdicts (CALIBRATED or NOT-CALIBRATED, GO or NO-GO) on trading strategies, ensemble portfolios, and predictive models submitted for review.

Our verdicts are external, adversarial, and public by default. A strategy submitted to Nyalai is subjected to five formal gates, each addressing a specific failure mode identified in the quantitative literature. Only strategies that pass all five gates receive a GO verdict. Failure at any single gate produces a NO-GO verdict, together with a diagnostic verdict-form documenting precisely where and why the strategy failed.

1.2 What Nyalai is not

Nyalai is not a backtesting platform. It does not help you design strategies. It does not tune parameters. It does not suggest improvements. It does not offer training. It does not manage capital.

Nyalai is not a certifier. A GO verdict from Nyalai is not a warranty of future returns. It is a statement that, as of the date of the verdict, the strategy has survived a documented adversarial process. Nothing more. The decision to deploy capital remains with the allocator.

Nyalai is not neutral. We hold strong opinions about what constitutes rigorous validation. We publish those opinions in this Doctrine. Anyone is free to disagree with us, and to build their own validation infrastructure with different opinions.

1.3 Why validation infrastructure matters now

The industry has been through fifteen years during which strategy design was democratized far faster than strategy validation. Cloud computing, open-source libraries, historical data availability, and now large language models have made it possible for a single researcher to test in a week what once required a team over a quarter.

Validation has not kept pace. Backtesting frameworks remain shallow. Statistical rigor remains the exception. The gap between "the strategy passed my tests" and "the strategy will survive live capital" has widened. Every allocator we speak to describes the same problem: too many strategies to evaluate, no shared standard for what counts as evaluated.

Nyalai exists to close that gap. Not by replacing the internal quantitative work of funds, but by providing an external, adversarial layer that renders public judgment on that work.

A note on institutional maturity. Nyalai is a young operation at the time of writing. The seriousness of this document must be matched by the seriousness of every subsequent verdict. If the practice deviates from the doctrine, the doctrine is what is wrong, not the practice. Revisions will be public and versioned.

1.4 Scope of the calibration engine

A verdict from Nyalai applies to a specific strategy, on a specific instrument universe, within a specific regime, over a specific horizon. It does not generalize across scopes without an explicit re-evaluation.

The Fourth Quadrant. In the taxonomy Nassim Taleb published in 2008, decisions divide along two axes: whether the payoff structure is simple or complex, and whether the underlying randomness lives in thin-tailed or fat-tailed distributions. The Fourth Quadrant, complex payoffs under fat tails, is where statistical inference is hardest and where model failures are most costly. It is where quantitative strategies actually operate: nonlinear payoffs, fat-tailed asset returns, regime shifts. Institutional model risk management as currently practiced under SR 11-7 and equivalent guidance addresses the first three quadrants adequately. Nyalai addresses the gap at the fourth. Nyalai does not eliminate Fourth Quadrant risk. Nyalai claims to make the decisions rendered under that risk auditable, and to refuse strategies whose evaluation would require thin-tailed assumptions the underlying data does not support.

Instrument scope. The five gates as currently specified apply to strategies operating on liquid, mark-to-market instruments (public equities, listed derivatives, liquid credit, spot and forward FX, listed commodities). They do not apply, in their current form, to strategies whose realized returns are only observable at exit and whose interim marks are estimated rather than transacted. Private equity, private credit, venture, and real assets fall outside the scope of the current gates.

Regime scope. A CALIBRATED verdict is a statement about performance under the regimes observed in the evaluation window, as segmented by Gate 2. It is not a statement about performance under regimes the strategy has not encountered.

Horizon scope. The gates are calibrated to strategies operating on daily-to-weekly time horizons. Sub-second, sub-minute, and intra-day regimes require adapted tests within the same five-gate discipline.

The verdict-form always names the scope on its first page. A verdict without a stated scope is not a Nyalai verdict.

2. Our approach to validation

2.1 Why principles over rules

"We do not try to prove that a strategy works. We try to eliminate the bad reasons to believe that it does."

That epigraph carries the entire Doctrine. Nyalai is Popperian by design. We do not confirm hypotheses; we attempt to falsify them. A strategy that survives our attempts to refuse it is not proven right. It is merely not proven wrong on the evidence available.

We could have written this Doctrine as a set of rules. We chose instead to write it as a set of principles, followed by explanation of why each principle exists.

The reason: quantitative finance evolves. New instruments appear. New statistical techniques emerge. New sources of data become accessible. New failure modes are discovered. A rule written for one context breaks in another. A principle, if well-articulated, adapts.

2.2 Why we publish our methodology openly

Three reasons.

First, because a validation methodology that cannot be inspected cannot be trusted. If we refused to explain what we test, our verdicts would be indistinguishable from opinion.

Second, because the state of the industry benefits from open methodology. If competitors adopt our gates and improve them, we all end up with better strategies deployed.

Third, because openness is a form of discipline. Publishing the method commits us to defending it in public. If a gate is wrong, it will be caught. If a gate is incomplete, it will be extended.

The published version is always identifiable by version number and date. Older versions and past verdicts are archived at nyalai.com/verdicts and nyalai.com/doctrine under CC0 (MEASURED, live at publication time), with Internet Archive Wayback Machine snapshots occurring at Wayback's own crawl cadence (INFERRED).

2.3 The 5-gate framework overview

A strategy submitted to Nyalai passes through five gates in sequence. Each gate has a specific failure mode it is designed to detect. A strategy that fails one gate is not tested at subsequent gates. The failure mode, the evidence, and the diagnostic are recorded in the verdict-form and published.

  • Gate 1: Statistical fragility. Does the observed performance survive correction for multiple testing and sample size?
  • Gate 2: Cross-regime robustness. Does the strategy remain profitable across market regimes, not only in aggregate?
  • Gate 3: Guided-search safety. Does the strategy retain significance after correcting for the guided nature of the development process?
  • Gate 4: Probabilistic calibration. Are the confidence estimates emitted by the strategy calibrated with realized outcomes?
  • Gate 5: Bootstrap confidence. Does the strategy remain profitable across resampled histories, with confidence bounds that exclude break-even?

2.4 On the sources of edge we recognize

A Sharpe ratio is a number. It says nothing, by itself, about why the number should persist. Two strategies with identical Sharpe can carry entirely different risk profiles going forward. We therefore ask, before the gates run, what class of edge the strategy claims. We recognize three broad classes.

Speed edges. Faster access to information, faster execution, faster reaction to news. These are real, and they are almost always transient. Speed advantages are competed away in months by well-capitalized entrants.

Behavioral and risk-premium edges. Compensation for bearing a risk that other market participants systematically avoid, or exploitation of persistent behavioral tendencies of large classes of participants. These are more durable but not permanent. Value, quality, momentum, and low-volatility factors fall in this class.

Structural and mechanical edges. Edges that arise from the mechanics of the market itself: index inclusion effects, forced flows around rebalancing events, specific microstructure asymmetries, regulatory-driven flows. These persist as long as the mechanic persists.

This tripartite classification consolidates the practitioner framings articulated by Cliff Asness on speed edges, Andrew Ang on behavioral and risk-premium edges, and Jean-Philippe Bouchaud et al. on structural microstructure edges.

A strategy that clears the five gates on statistical grounds but cannot name the class of edge it depends on has not answered the load-bearing question. Our verdict-form therefore records the declared edge class alongside the gate results.

2.5 On model interpretability

A recurring objection in modern quantitative finance concerns machine-learning models whose individual features are not humanly interpretable. Our position is precise.

We do not require per-feature transparency. Requiring humanly interpretable features would exclude a class of models (learned representations, dense neural encoders, ensembles) that in some contexts genuinely outperform interpretable alternatives.

We do require aggregate behavioral coherence. A model whose outputs cannot be reconciled with a stated economic hypothesis at the portfolio level is not a validation candidate. If the model claims to trade momentum, its aggregate signal must correlate with published momentum factors on out-of-sample windows. Opacity at the feature level is acceptable; opacity at the strategy level is not.

2.6 On the virtue of complexity

Recent academic work has argued that in machine-learning settings, the classical overfit-versus-underfit tradeoff is not the correct decision frame. Highly parameterized models, appropriately regularized, can outperform their sparser counterparts even when the number of parameters exceeds the number of observations. Kelly, Malamud, and Zhou (Journal of Finance, 2024, "The Virtue of Complexity in Return Prediction") advanced this argument for equity return forecasting.

We take this argument seriously and we do not treat it as a license.

The argument holds where the training procedure includes explicit regularization, where the out-of-sample performance is measured on windows that were causally isolated from parameter selection, and where the complexity is a modeling choice rather than a fitting artifact. Our Gate 3 (guided-search safety) is designed to distinguish these two cases.

Complexity, in short, is neither a virtue nor a vice in isolation. It is a variable that must be accounted for by the correction structure of the gates.

2.7 On the statistical inference horizon

A structural feature of financial time series is that the observation window available to any researcher is short relative to the horizon over which the underlying claim needs to hold. Fama and French, in their 2018 Financial Analysts Journal paper "Volatility Lessons," documented that the statistical significance of estimated equity premia deteriorates rapidly as the evaluation horizon extends, to the point where three-year, five-year, and even ten-year windows admit substantial probabilities of negative realized premia even when the true expected premium is high. The consequence for validation is not academic. A strategy with a positive expected edge over a twenty-year horizon may spend contiguous multi-year windows underwater, and a validation methodology that requires long-horizon confidence before rendering a verdict will render no verdict at all in reasonable operating time.

This is where external adversarial validation is a necessity rather than an addition. Internal review at a fund cannot solve the horizon problem, because the reviewer and the strategy author share the same observation window and the same incentives. External review does not solve the horizon problem either, but it changes the class of failure the review can catch. Internal review addresses whether the strategy has been engineered honestly. External review addresses whether the evidence, once honestly presented, actually supports the claim. Only the second is what an allocator needs to decide whether to deploy capital.

Nyalai does not extend the observation window. Nyalai treats the observation window as fixed and asks whether the claim rests on the evidence within it. When the evidence within the window is not sufficient to support the claim, the verdict is NO-GO, regardless of the practitioner's conviction that more time would prove the claim.

2.8 On modular validation methodology

Each of Nyalai's five gates uses a statistical framework appropriate to the specific question the gate addresses. Gate 1 uses a corrected significance test because the question is whether an observed Sharpe (for return-generating strategies) or Brier Skill Score (for probabilistic prediction systems) survives correction for exploration. Gate 2 uses a regime-switching model because the question is whether performance is uniform across market states. Gate 3 uses a permutation test because the question is whether search bias has inflated significance. Gate 4 uses reliability calibration measures because the question is whether probability estimates match realized frequencies. Gate 5 uses bootstrap resampling because the question is whether the expected edge is robust to path variation.

This is not eclecticism. It is discipline. Joseph Simonian, writing in the Journal of Financial Data Science in 2020, framed the same insight for institutional model validation: traditional econometric methods are best at explaining past behavior, machine-learning methods are best at prediction and pattern recognition, and a modular validation framework uses each where it is strongest rather than forcing one framework to answer all questions.

Nyalai's framework is modular by explicit design. A submitter who prefers a single-framework audit is not the customer we serve. Institutional model validation, as practiced under SR 11-7 and equivalent regimes, has long recognized that a portfolio of tests, each appropriate to a specific risk, is the credible unit of validation. The Doctrine adopts the same discipline.

3. Priority hierarchy

In cases where the principles of this Doctrine conflict, we prioritize as follows:

  1. Refusal-first. When in doubt, refuse. The cost of a false GO verdict (capital deployed against a broken strategy) dwarfs the cost of a false NO-GO (a good strategy sent back for further work). This asymmetry is the founding asymmetry of Nyalai.
  2. Adversarial rigor. Our verdicts must be defensible against a hostile technical audit. We would rather be wrong in the direction of harshness than in the direction of leniency.
  3. Client alignment. Where the previous two priorities are satisfied, we align with the interests of the allocator submitting the strategy. We do not align with the interests of the strategy author when those diverge from the allocator's interests.
  4. Commercial success. Where the previous three priorities are satisfied, we operate as a viable business. Where commercial viability would require us to relax rigor, to soften refusal, or to align with strategy authors against allocators, we choose the earlier priorities.

The order matters. In practice, most decisions do not involve conflict. The priority hierarchy exists to make our choices predictable in the minority of cases where genuine tension appears.

4. Being genuinely useful

A verdict is useful when it changes an allocator's action. A verdict is useless when it flatters the strategy author. We prefer the former, at the cost of the latter.

To the researcher who submitted the strategy, we owe a diagnostic that is actionable: which gate failed, what evidence produced the failure, which categories of evidence would warrant re-evaluation. We do not offer to fix the strategy; the diagnostic makes fixing possible.

To the allocator who will make a capital decision, we owe a verdict that is legible without a statistics degree: findings first, methodology documented in-form, references cited for anyone who wants to audit. Every verdict-form includes an executive summary a non-quant board member can read.

To the broader market, we owe published verdicts (including our own NO-GO verdicts on our own strategies), methodology openness, and honest acknowledgment when this Doctrine turns out to be incomplete or wrong.

Being useful means saying the smallest, sharpest thing that shifts a decision. It does not mean writing longer documents. It does not mean adding caveats.

5. Being broadly ethical

Nyalai operates in a domain where a false GO verdict has downstream consequences for retail investors whose pensions and savings are held by the institutions we serve. That is a public-interest asymmetry. We take it seriously.

5.1 Non-manipulation of results

We do not massage results. We do not adjust thresholds after seeing the outcome. We do not selectively apply gates depending on the desired verdict. We do not omit results that would be inconvenient. A narrow Gate 3 failure remains a failure.

We refuse any engagement structure that would create a financial incentive to lower our standards: no contingent fees, no performance-based compensation, no equity-in-lieu-of-cash with the entities we validate. We refuse to be paid more if we produce a GO verdict.

5.2 Transparency of methodology

Every gate uses a documented statistical test with a citable reference. Every threshold has a stated rationale. Every deviation from the standard method is flagged. Everything necessary to reproduce a verdict is included in the verdict-form.

5.3 Independent judgment (sovereignty rule #7)

Sovereignty rule #7. The party that produces a probabilistic claim never validates the same claim. Nyalai's validators do not communicate with strategy authors during evaluation and do not have direct access to strategy source code; they operate on submitted output data. This separation exists to prevent contamination of judgment by the persuasive presence of the researcher.

Historical grounding. Peter Galison and Lorraine Daston, in Objectivity (Zone Books, 2007), trace three overlapping regimes of scientific image making: truth-to-nature, mechanical objectivity, and trained expertise. Nyalai's five gates operate in the mechanical-objectivity register (protocol-driven arithmetic that leaves no room for the operator to correct favorably); the refusal doctrine and verdict-form registration operate in the trained-expertise register (calibrated judgment applied after the mechanical layer has done its work). The regimes are layered, not sequential.

Institutional precedent. The Event Horizon Telescope collaboration ran four independent imaging teams forbidden from communicating during the imaging phase, reconvening only in late July 2018 to verify pixel-by-pixel agreement before the M87 release of 10 April 2019 (Galison, Fireside Chat with Andrew Lo at the Harvard Data Science Initiative). Nyalai adopts the same posture at institutional scale: the validator does not participate in strategy design, does not modify the ledger, does not counsel on how to improve gate scores.

Longer precedent. From the late 1960s at Lawrence Berkeley National Laboratory, Luis Alvarez (Nobel Prize in Physics 1968) required every discovery claim from his group to survive an adversarial ranking: one hundred Monte Carlo realizations of the null were mixed with the real data, and publication was blocked unless the real data appeared in the top rank of an independent sort by each team member. Nyalai's Gate 3 permutation testing (Aronson and Masters 2013) is a methodological descendant of that sixty-year-old protocol.

The Einstein-1939 lesson. In 1939 Einstein published a paper arguing that black holes cannot exist (Annals of Mathematics, vol. 40, no. 4, pp. 922 to 936). Andrew Strominger has publicly characterized that argument as one no graduate general-relativity student today could submit and pass (paraphrase, exact venue candidate: Lex Fridman Podcast episode 359). The correction came from the field over decades, not from Einstein himself. Sovereignty Rule #7 defends against exactly the pattern in which a producer of a claim, however authoritative, cannot be relied on to adversarially test their own claim.

The trained-expertise layer. Returning from history to Nyalai's own stack, the validator stands in the trained-expertise register defined by Galison and Daston. When the validator refuses a submission, the refusal is neither a raw arithmetic output nor a free personal opinion. It is a calibrated-judgment layer applied on top of the mechanical layer, and both layers are visible in the published verdict-form. Trained expertise is not a return to pre-mechanical idealization; it is the recognition that mechanical output can mislead when the apparatus itself has been calibrated on the observer's expectations.

Multi-method convergence. Sheperd Doeleman (Founding Director of the Event Horizon Telescope) has described, across multiple public lectures on the EHT methodology, how the collaboration deliberately built redundancy into every decision: multiple data-processing streams, multiple imaging algorithms, four separate teams, so that no result advanced until the independent paths agreed (paraphrase). Nyalai's trained-expertise layer runs multiple parallel arithmetics against the same submission (walk-forward window, block-bootstrap seed, feature-selection scan) and refuses to render a verdict until they converge. Convergence is the precondition of the verdict, not decoration.

Residual-bias floor. Team isolation does not eliminate shared human bias. Katie Bouman, describing the EHT imaging effort, has publicly named the risk that isolated teams could still bias one another toward an expected ring shape (paraphrase; a closely related direct statement appears in her CNBC interview of 12 April 2019). The verdict-form therefore includes, since v0.3.3, a "known bias sources" annex that names the ways the validator's calibration could still steer the result and the mechanisms by which those steering paths were adversarially tested during calibration itself. Where the annex is empty, the validator has not thought hard enough about their own bias sources, and the verdict-form is not eligible for publication.

6. Being honest

Eight properties define what honesty means for a validation infrastructure. Seven were carried forward from earlier versions and extended with operational specifications in v0.3.2.3 ; the eighth was added in v0.3.2.3 to separate research quality from disclosure quality. The full canonical text lives in the v0.3.3.1 sealed markdown ; this page carries the editorial abbreviation.

6.1 Truthful

We do not state falsehoods, whether to strategy authors or to allocators.

The v0.3.2.3 ProvenanceContract extends this property to every statement in the operating environment (code, log, file, folder, method, clock). A state that has not been observed is resolved by the command that exists to observe it, and labeled INFERRED only when no such command exists. The freshness clause adds that a cached or stale artifact is never asserted as the current state ; when currency matters, the state is re-observed or its staleness is declared.

6.2 Calibrated

We maintain calibrated uncertainty and do not claim more (or less) confidence than the evidence supports.

Every affirmation in a verdict-form carries an explicit label from the seven-label epistemic set: MEASURED, STRUCTURALLY-PROVEN, INFERRED, DECLARED-BY-SUBMITTER, NON-COMPUTABLE, NON-EVALUABLE, NON-APPLICABLE. The three "cannot-measure" labels are distinct by cause: NON-COMPUTABLE names a missing-input problem, NON-EVALUABLE a scope problem within a defined test, NON-APPLICABLE a category error. An unlabeled affirmation is a §6.2 violation regardless of its factual accuracy.

6.3 Transparent

We do not pursue hidden agendas: our methodology is public, our threshold choices are documented, our verdicts are inspectable.

The v0.3.2.3 ThresholdRegistry extension requires every documented threshold to carry a measured power curve, generated by a script committed to the validator's public repository with a stated random seed and replayable on the production validator machine. Invented thresholds (written into code from memory without documented power characteristics) are not eligible for a published verdict.

6.4 Forthright

We proactively share information that a strategy author or allocator would want to know, even when not asked.

The v0.3.2.3 InputManifest immediate-share extension requires anomalies detected during evaluation to be surfaced in-context, at the moment of detection, rather than appended to a subsequent audit round. The threshold is not "the auditor might not find it" but "the counterparty would want to know now."

6.5 Non-deceptive

We do not create false impressions of a strategy's quality through selective framing, misleading emphasis, or technically-true but implicative language.

The v0.3.2.3 non-evaluable is not failure extension specifies that a gate whose input cannot yield a measured result (non-computable, non-evaluable, or non-applicable) returns a status distinct from failure, and the validator treats all such gates symmetrically. Any asymmetry in code is a defect to correct, not a verdict to cite. Counting a non-measured gate as a failed gate manufactures a false impression of evidentiary weight.

6.6 Non-manipulative

We do not soften a NO-GO because the author is likable, and we do not harden one because the author is difficult.

The v0.3.2.3 no hardening for rigor extension adds a second forbidden direction: we do not stack non-measured gates onto a verdict to signal implacability. A verdict rests only on the gates actually measured. The bidirectional relational-information excluded clause strips personal friendship, prior conflict, shared network, adversarial history, and reputation from the judge's context before verdict, with recusal or declared-exposure when it reaches the judge inadvertently.

The v0.3.3.1 Attestation panel makes the exclusion auditable rather than left to the judge's own report. Every verdict-form carries a panel signed by the validator instance naming five required fields: (a) the pre-blinding step that stripped submitter identity and relational metadata, (b) the checksum or hash of the pre-blinded package, (c) any exposure event between pre-blinding and verdict (the null case is stated explicitly), (d) the identity of the second validator when recusal was invoked, and (e) the timestamp of the pre-blinding step, evidencing that the exclusion completed before gate execution began. An absent or empty panel blocks publication.

6.7 Autonomy-preserving

We render verdicts. We do not make deployment decisions ; the allocator retains full authority over what to do with them.

The v0.3.2.3 verdict-form template separation extension separates a VERDICT section (gate results and the binary GO or NO-GO conclusion) from an OBSERVATIONS section (structured observations offered for judgment rather than counsel imposed). There is no recommendations section. Imperatives directed at the submitter or the allocator ("do not deploy this," "wait before resubmitting") are §6.7 violations regardless of the accuracy of the underlying observation.

The v0.3.3 submitter-initiated guidance clause acknowledges that a submitter or allocator may explicitly request operational guidance in a separate, clearly-scoped exchange after the verdict has been rendered. Such a response is labeled consultative rather than adjudicative, is not appended to the verdict-form, and does not modify the verdict itself. It is a mode Nyalai may enter ; never one it initiates.

6.8 Research-versus-disclosure separation

New property introduced in v0.3.2.3. The quality of the submitter's disclosure discipline (ledger cleanliness, stated reservations, audit-trail transparency) and the quality of the submitter's research methodology (the presence of an edge that survives the gates) are judged separately and may diverge without contradiction. Conflating the two into one implicit verdict is a §6.8 violation because the two are load-bearing on different institutional decisions.

The extension is anchored to an internal ʼCɩcɛ adversarial session of July 14, 2026 (identifier R-02, internal precedent record retained privately per §12.4 sequence commitment). The record remains internal because the submitter's consent to public release has not been established under the §12.4 order-of-operations.

A sketch of the disclosure-quality assessment is recorded as a placeholder pending v0.3.4 formalization: a 5-axis rubric covering (1) ledger cleanliness, (2) stated-reservations completeness, (3) audit-trail transparency, (4) responsiveness during evaluation, and (5) post-verdict disclosure discipline, each scored at three levels (satisfactory, marginal, insufficient). Exact axis definitions, scoring criteria, and combination rule are pending v0.3.4 formalization and Cice adjudication.

7. Avoiding harm

The primary harm we avoid is the false GO verdict deployed to live capital. That failure mode is what the gates are designed to catch.

The secondary harm we avoid is the wrongful damage to a legitimate strategy through a false NO-GO. Our verdicts are contestable. We publish the audit trail. Any submitter can request an independent second opinion using the same methodology; the audit trail permits full reproduction.

The tertiary harm we avoid is our own capture by commercial pressure. Section 12.3 describes the business-model architecture that preserves our independence.

7.1 Cost-benefit framework

Every verdict Nyalai renders is a decision under uncertainty. We are more tolerant of Type II error (rejecting a good strategy) than Type I error (accepting a bad one). Capital lost is difficult to recover; a rejected strategy can be re-submitted after further work.

Fourth Quadrant regime-dependent asymmetry (v0.3.3). The Type I greater-than Type II asymmetry is quantitatively more pronounced in the Fourth Quadrant regime (complex payoffs in Extremistan, per §1.4) than in the first three quadrants where SR 11-7 model risk aggregation operates. Nyalai therefore holds itself to a stricter Type I threshold on Fourth Quadrant submissions.

8. Hard constraints (7 bright lines)

The following constraints are absolute. They are not weighed against other priorities. They function as filters on the space of acceptable actions.

  1. Never publish a GO verdict without all five gates having been passed.
  2. Never accept capital or other consideration in exchange for softening a verdict or altering methodology.
  3. Never anonymize our own NO-GO verdicts. If Nyalai's own strategies fail Nyalai's own gates, the verdict is published under Nyalai's name.
  4. Never validate a strategy where we have detected credible evidence of data leakage into the sample.
  5. Never render a verdict under time pressure from the submitter. Verdicts take as long as they take.
  6. Never claim more certainty than the evidence supports.
  7. Never claim less certainty than the evidence supports. Epistemic cowardice violates our honesty norms as surely as overclaiming does.

These constraints are meant to be uncrossable. When faced with a seemingly compelling argument to cross one, our default is to increase suspicion, not to comply. A persuasive case for crossing a hard constraint is more likely evidence of manipulation than evidence of a legitimate exception.

9. The 5 gates in detail

Each gate is described here with its purpose, its statistical test, its threshold, and the primary academic reference from which the test derives.

9.0 Pre-check: the one-page thesis submission

Before any of the five gates is executed, the strategy author must submit a one-page thesis. This is a hard admission requirement, not a formality.

The one-page thesis states, in prose: the economic or behavioral claim the strategy exploits; the class of edge (per Section 2.4: speed, behavioral or risk-premium, structural or mechanical); the instrument universe and the intended regime of application; the intended sizing envelope: expected capital, expected gross and net gearing, expected turnover; the failure modes the author expects, and the observations that would falsify the thesis.

The one-page thesis is submitted publicly with the verdict-form. It is not confidential and it is not editable after submission. A strategy whose thesis has not been submitted, or whose thesis has been retroactively modified after gate results, receives no verdict.

9.1 Gate 1: Statistical fragility

Purpose: to detect systems whose apparent edge does not survive correction for multiple testing, non-normality, and sample size. Statistical fragility is the property that separates a genuine signal from an artifact of exploration; Gate 1 is the arithmetic that tests for it, adapted to the output type of the system under evaluation.

Gate 1 is operationalized in two variants depending on the primary output of the system under evaluation. Both variants test the same underlying property (aggregate skill in excess of chance, after correction for the number of configurations effectively explored during development) and both are treated as equally load-bearing at the Gate 1 slot. The variant is determined by the system, not by the submitter.

Gate 1a, return-generating strategies. Applies to any system whose primary output is a return series or a directional position (long, short, flat) whose realized profit-and-loss is the object of evaluation.

  • Test: Deflated Sharpe Ratio (Bailey and López de Prado, 2014). The observed Sharpe ratio is penalized based on the number of configurations effectively tested during development, the non-normality of returns (skewness and kurtosis), and the length of the observation window.
  • Threshold: DSR ≥ 0.95. This threshold is the 95% confidence level below which Bailey and López de Prado (2014, page 9) explicitly declined to certify an empirical discovery: an example in that paper is annotated "not a legitimate empirical discovery at a 95% confidence level" with the counter-example computed as DSR = 0.9505 "above the 95% confidence level."
  • Failure mode: DSR < 0.95 after correction.

Gate 1b, probabilistic prediction systems. Applies to any system whose primary output is a calibrated probability estimate over a resolvable event, evaluated against a frozen ledger of resolved predictions. Systems in this class do not produce a return series absent an external position-sizing rule; the DSR of a return series constructed post hoc from probabilistic outputs would inherit the sizing rule as a design choice and would not test the underlying prediction skill.

  • Test: Brier Skill Score (BSS) against a naïve baseline, with corrections for finite-sample effects and for the number of prediction rules effectively evaluated during development. The naïve baseline is the base-rate predictor over the resolution window.
  • Threshold: BSS ≥ 0.05.
  • Failure mode: BSS < 0.05 after correction, including BSS ≤ 0 in which the system underperforms the naïve baseline.
  • Reference: Brier, G. W. (1950). Verification of Forecasts Expressed in Terms of Probability. Monthly Weather Review, Vol 78 No 1. Modern normalization: Murphy (1973).

Decision rule. The variant is fixed at submission and stated in the one-page thesis of Section 9.0. A system is a return-generating strategy (Gate 1a) if its evaluation ledger records realized returns produced by the system's own position-sizing logic. A system is a probabilistic prediction system (Gate 1b) if its evaluation ledger records resolved probability estimates without any position-sizing layer supplied by the system itself. Dual-output systems that submit both ledgers must pass both variants.

Threshold rationale for BSS ≥ 0.05. Neither Brier (1950) nor its modern operationalizations prescribe a universal skill-score threshold. The value 0.05 is chosen as the Nyalai operating threshold. Rationale: BSS = 0 is trivial pass; BSS = 0.10 is the informal "clearly-skillful" level in operational meteorology, a strict-half discount produces 0.05 as a defensible working floor; and 0.05 aligns numerically with the ECE ≤ 0.05 threshold at Gate 4 for legibility. The threshold is contestable; the doctrine welcomes principled arguments for a higher floor.

9.2 Gate 2: Cross-regime robustness

Purpose: to detect strategies whose profitability is concentrated in a single market regime.

Test: Hidden Markov Model regime detection applied to underlying market data. Strategy performance is computed conditional on each detected regime.

Threshold: Sharpe > 0 in each detected regime.

State-compression framework choice (v0.3.3). Gate 2 accepts two operationalizations of regime detection. The first is explicit regime definition: rolling percentile buckets on realized volatility, defined thresholds on drawdown depth, or a bull / bear / sideways classification driven by a moving-average signal. This gives full interpretability at the cost of accountability for the boundary choice. The second is data-inferred detection via HMM latent-state inference (Baum-Welch parameter inference, Viterbi decoding, forward-backward smoothing), which compresses multiple latent factors into a learned state space at the cost of interpretability. Both are legitimate provided the submitter discloses the framework used, the number of states, and the rationale. The number of latent states is itself a hyperparameter subject to the Gate 3 guided-search correction.

Chi-squared reliability thresholds (v0.3.3). Where Gate 2 uses a chi-squared test on transition frequencies, the test requires a minimum cell count for the statistic to behave reliably. Cochran 1954 requires that no expected cell count fall below five; the Markov-chain-for-finance pedagogical literature adopts a stricter thirty-observations-per-transition-cell operating rule, and Nyalai follows the stricter rule. Below either threshold, the chi-squared p-value is not usable as evidence for or against the verdict; Gate 2 defers to a rank-based or exact permutation-based alternative disclosed in the reproducibility appendix, or declines to render until the ledger is extended.

Apparent N versus true N per regime (v0.3.3). A submitter who slices N observations into k regimes obtains, on average, N over k apparent observations per regime. The true N is smaller in two ways Gate 2 must count. Autocorrelation within a regime shrinks the effective sample size by roughly one over the block size at which serial correlation decays (block-bootstrap discipline). Regimes at the window boundary are truncated by an unknown fraction of their true residence time. The Gate 2 row of the verdict-form reports both figures: apparent N (raw count) and effective N (raw count reduced by the block-bootstrap size). Confidence intervals on within-regime Sharpe widen at the rate the true N shrinks, and are disclosed accordingly.

Reference: Rabiner (1989) tutorial, Hamilton (1989) Markov-switching autoregression, Ang and Bekaert (2002) interest-rate regime detection; Cochran (1954) chi-squared reliability; Politis and Romano (1994) block-bootstrap discipline; Simonian (2020) block-size instantiation.

9.3 Gate 3: Guided-search safety

Purpose: to detect strategies whose significance has been inflated by the guided nature of the development process.

Test: Guided-search-safe permutation test (Aronson and Masters, 2013), which corrects the classical White's Reality Check for the specific case of directed rather than random search.

Threshold: p < 0.05 after correction, with the correction based on the disclosed number of configurations tested during development.

Runs and aggregation (v0.3.3). The number of Gate 3 runs is declared at submission (single-run by default, or multiple pre-registered runs across distinct evaluation configurations such as walk-forward windows or block-bootstrap seeds). When multiple runs are declared, §12.5 (On consensus averaging) governs the aggregation of pass/fail outcomes across runs. Single-run submissions apply the classical threshold above without averaging.

9.4 Gate 4: Probabilistic calibration

Purpose: to detect strategies whose confidence estimates diverge from realized outcomes, a condition that breaks position sizing even when the directional signal is sound.

Test: Brier Skill Score against a naive baseline, reliability diagram, Expected Calibration Error (ECE).

Threshold: BSS ≥ 0.05 (skill over baseline), ECE ≤ 0.05.

9.5 Gate 5: Bootstrap confidence

Purpose: to detect strategies whose expected edge is positive but whose variance is large enough that the strategy may not remain profitable under realistic path variations.

Test: Bootstrap resampling (block bootstrap for auto-correlated returns) of the return series, with recomputation of the Sharpe ratio for each resample.

Threshold: 95% confidence interval lower bound > 0 (net of costs).

9.5a Reproducibility calibration protocol

New sub-section (v0.3.3). Before running the five gates on the actual submission, the validator runs the same five-gate pipeline against three or more synthetic submissions whose properties are known to the validator and unknown to the submitter. The protocol is borrowed from the Event Horizon Telescope's image hackathons, one of which used a blurred rendering of Frosty the Snowman as ground truth precisely because no imaging team could correctly guess it in advance. If the pipeline correctly reproduces the known properties of the synthetic ledgers (regime structure, presence or absence of a genuine edge, expected DSR distribution), the validator is calibrated for the submission at hand. If it fails, the pipeline is repaired before the actual submission is evaluated. Calibration on synthetic ground truth is the price of admission. It is not optional.

Adversarial-shape variant. The synthetic set includes at least one target whose expected shape is the opposite of what the submission would produce if the submitter's claim were true. The EHT teams tuned parameters on synthetic discs (expected shape: filled) before transferring them to M87 data (expected shape: ring with hole); the transferred parameters still recovered the hole, confirming that the pipeline was not merely finding what it was trained to find. Nyalai adopts the same discipline: where the submission would produce a passing gate arithmetic if the claim were true, the pipeline is first calibrated on a synthetic in which that same submission would produce a failing arithmetic.

Helm of Hades and Spring of Narcissus. Peter Galison, in his Harvard CMSA Yip Lecture, named the two epistemic nightmares any validation pipeline must guard against. The Helm of Hades makes the target invisible: the pipeline sees only noise where the effect resides. The Spring of Narcissus makes the target reflect the observer: the pipeline sees only what the observer's priors have projected into the data. Calibration on null synthetic targets defends against Hades. Calibration on adversarial-shape targets defends against Narcissus. A pipeline that has run only one class of calibration is defended against only one class of nightmare, and the verdict-form's methodology annex annotates the gap explicitly.

Pedagogical summary. In plain language: before Nyalai runs its five gates on your actual submission, it first runs the same pipeline on a handful of synthetic submissions whose true properties Nyalai knows and you do not. Some are constructed to look like a winning strategy but to actually be pure noise; others to look like a losing strategy but to actually contain a real signal shaped in an unexpected way. If Nyalai's pipeline calls the noisy ones "no edge" and the real-signal ones "edge here even though it does not look like what you expected," the pipeline is calibrated well enough to be trusted on your submission. If it gets the synthetic cases wrong, it is repaired first. The calibration record is included in every verdict-form so the reader can inspect it.

Closure quantities as a research direction. The EHT imaging pipeline uses closure phases (products of complex visibilities across three telescopes) and closure amplitudes (products across four telescopes) because these quantities are mathematically invariant to atmospheric and instrumental gain errors. They are calibration-invariant primitives. Nyalai will identify the analogous primitives in quantitative-finance validation (permutation-invariant test statistics, quantile-preserving transforms, rank-based measures) and prefer them in future gate revisions where such primitives exist.

9.6 A note on reward-hacking

A pattern that the machine learning community has recently formalized applies directly to the class of strategies our gates are designed to refuse. A reward-hacked strategy in quantitative finance is a strategy whose backtest window was selected to fit the reported Sharpe; a strategy that removes losing periods and calls the result "trained"; a strategy whose walkforward test skipped the guided-search safety check because the check would have said no.

Our gates target this class explicitly. Gate 1 targets the Sharpe or Brier Skill Score inflated by trial selection, depending on the variant applied per §9.1. Gate 3 targets the guided-search bias inherent in modern data-driven exploration. Gate 5 targets the confidence interval that a single lucky window can artificially compress. A reward-hacked strategy that has cleared its own internal review will not clear ours, and it should not.

9.7 Verdict scope statement (three components) plus communication layer (presentation format)

The five gates evaluate signal quality. They do not evaluate portfolio construction. Signal quality and portfolio construction interact in ways that can turn a technically CALIBRATED strategy into a losing implementation, and the Doctrine must acknowledge this explicitly.

Every verdict-form therefore includes a verdict scope statement with three substantive components, presented through a required communication layer.

1. Regime scope. The set of market regimes, as segmented by Gate 2, within which the verdict applies. Regimes not present in the window are not certified.

2. Sizing envelope. The range of capital, gross gearing, net gearing, and turnover under which the gate results were computed. A verdict computed on a $10M gross book does not certify behavior at $500M. Deployment outside the envelope requires re-evaluation.

3. Portfolio-construction assumptions. The correlation structure, volatility target, and rebalancing cadence assumed during evaluation. Volatility drag, correlation collapse under stress, and rebalancing frictions can convert a positive-expectation signal into a negative-expectation portfolio.

Communication layer (presentation format). The verdict-form presents its findings in two layers, and both are required. The upper layer states, in institutional vocabulary, whether the strategy is CALIBRATED and which of the classical risk categories it exposes the allocator to. The lower layer supplies the technical evidence: DSR value or BSS value (whichever Gate 1 variant applies per §9.1), BSS and ECE numbers for Gate 4, regime-by-regime Sharpe, bootstrap confidence intervals, guided-search correction magnitudes. The upper layer must never overstate what the lower layer supports. The lower layer must never be omitted. This two-layer structure follows the principle established by Simonian (2020) for the interpretability of machine-learning model output in institutional review.

Nyalai does not validate portfolio construction. Portfolio construction is the allocator's discipline. What the Doctrine requires is that the verdict-form make the boundary between signal validation and portfolio construction unmistakable, and that any deployment outside the stated envelope invalidates the verdict.

Sections 9.8 through 9.10 are reserved for future methodology extensions. Sections 9.11 and 9.12 formalize existing operational disciplines that predate the current numbering convention and are preserved at their canonical anchors for citation continuity.

9.11 Triangulation of the verdict

Every published verdict is contextualized against three axes: the type of risk the strategy exposes the allocator to, the business context in which the allocator would deploy it, and the regulatory or jurisdictional frame that governs the deployment. This triangulation, adapted from the three-lens framework Senthil Kumar established for institutional risk management at BNY Mellon, ensures that no material risk axis is treated as absent when it is merely unspoken.

The verdict-form therefore includes, alongside the scope statement, a triangulation panel stating the category of risk the strategy exposes the allocator to (credit, market, operational, liquidity, model, counterparty), the business type against which the verdict is calibrated (institutional multi-manager pool, single-strategy fund, family office, corporate treasury, retail-facing product), and the regulatory frame (SR 11-7 or equivalent domestic MRM guidance, ISDA SIMM regime, EMIR, MiFID II, or explicit statement of the applicable frame at the deployment location). A verdict computed against one triangulation cell does not extend to another.

9.12 The framework, not the model

Institutional model validation is not the validation of a single artifact. It is the maintenance of a documentation framework, typically comprising eight to ten distinct sets of documents even for the simplest vanilla portfolio, distributed across three lines of defense: model owner documentation, independent model validation, and internal audit. Roland Stamm, in his work on ISDA SIMM model validation at Acadia, has described the framework as the true unit of institutional validation. A verdict-form that positions itself as an independent, external adversarial anchor within that framework serves the institution. A verdict-form that positions itself as a replacement for internal MRM documentation misunderstands the structure of the practice.

The Doctrine adopts the framework framing. Nyalai's verdict-form is designed to be citable by model owners as external validation evidence, by independent validation teams as a cross-reference, and by internal audit as one input to their annual review. The verdict-form itself contains a citation guidance section on its final page, stating how the document may be referenced by each of the three lines of defense.

The framing in one sentence (v0.3.3.1). The model owner reasons from the inside outward about what the strategy is likely to be; Nyalai reasons from the outside inward about what the strategy cannot credibly be; and the institution decides only when the two reasonings meet in the middle.

10. Model Trust Levels

The five gates return a binary verdict (CALIBRATED or NOT-CALIBRATED) on a probabilistic system at a moment in time. A verdict is not a deployment decision. It is a technical statement about calibration and skill.

We propose an intermediate vocabulary that closes the gap between a technical verdict and a deployment decision. We call it Model Trust Levels, or MTL. The structure follows the pattern of biosafety levels in laboratory infectious disease research and of AI Safety Levels in frontier machine learning.

10.1 MTL-1. Untested

No independent validation performed. No public methodology. No track record. Outside the acceptable perimeter for capital allocation.

10.2 MTL-2. Nominally Calibrated

Passes Gate 1 (Gate 1a DSR for return-generating strategies, or Gate 1b BSS for probabilistic prediction systems, per §9.1) and Gate 4 (probabilistic calibration) on a publicly documented ledger. Acceptable as secondary advisory signal, never as sole allocation signal.

10.3 MTL-3. Robustly Calibrated

Passes MTL-2 plus Gate 2 (cross-regime robustness via Hidden Markov Model regime segmentation). Failure modes are documented publicly. Acceptable as contributing signal in a multi-signal portfolio, with position limits and quarterly revalidation.

10.4 MTL-4. Adversarially Validated

Passes MTL-3 plus Gate 3 (guided-search safety) and Gate 5 (bootstrap confidence). Independent audit trail available to regulators on request. Sovereignty rule enforced. Acceptable as primary allocation signal at institutional scale, with biannual revalidation and annual public report.

10.5 MTL-5. Continuously Live-Validated

Passes MTL-4 plus rolling verdict updates on live data. Automatic drift detection and forced revalidation on regime change. Full audit-trail transparency to subscribing allocators. Acceptable for scaled deployment under continuous supervision.

10.6 Properties of the framework

Strict nesting. Each level includes all requirements of every lower level. An MTL-4 is necessarily MTL-3, MTL-2, and MTL-1.

Binary per level. A system is MTL-N or it is not. There is no MTL-3.5. The rigidity is intentional: it mirrors the refusal-first discipline at the level of individual gates and prevents negotiation of thresholds under commercial pressure.

Publicly contestable. Every MTL attribution is accompanied by the audit trail that permits an independent third party to reproduce the computation and contest the attribution.

Time-bounded. An MTL-4 attributed in January is not an MTL-4 in July without revalidation. Revalidation cadence is publicly declared per level.

Mechanical expiration. A verdict at MTL-N expires N months after issuance without revalidation. Expiration is mechanical: it does not require a Nyalai action to take effect. MTL-1 expires after one month ; MTL-2 after two ; MTL-3 after three ; MTL-4 after four ; MTL-5 after five. For MTL-5, whose defining property is continuous live revalidation, the five-month figure is the maximum permitted gap between revalidation checkpoints rather than a shelf-life ; an MTL-5 that stops being continuously revalidated is by definition no longer MTL-5. An expired verdict may not be cited as a live MTL attribution; it may only be cited as an expired historical record with its last revalidation date.

Sovereignty carries across levels. The requirement that validator and strategy author be structurally distinct applies at MTL-3 and above. Below that threshold, self-validation is permitted with an explicit self-validated flag.

11. Preserving important market structures

11.1 Anti-consolidation of validation power

We would consider it a bad outcome if Nyalai became the only meaningful validation authority in quantitative finance. Concentration of judgment on strategy quality into a single organization creates single points of failure, single points of capture, and single points of narrative distortion.

Our discipline requires that a small number of independent validators eventually operate in parallel, each with its own methodology, each contestable, each auditable. Nyalai publishes its methodology in full as a matter of institutional discipline, not as an invitation. We aspire to remain the reference, not the monopoly.

11.2 Style drift invalidates the verdict

A verdict-form is valid only for the specific model configuration tested. Style drift, defined as any material change in the features, instrument universe, risk parameters, or claimed problem domain, invalidates the previous verdict. Re-validation is required, not optional. The audit trail lodged with the original verdict-form records the configuration state; comparison against the current state is a first-line test that a submitter must pass before requesting extension of a prior CALIBRATED verdict.

This discipline protects the allocator from stale certifications on materially changed models. The re-validation cadence declared in Section 10 is a floor; style drift triggers immediate revalidation regardless of when the last cadence-driven review occurred.

The v0.3.3 addition adds a Lindy backwards-readability clause: every doctrine revision must remain backwards-readable so that an audit committee reading the current version alongside a v-1 version can reconstruct the reason for each change without archaeological reconstruction.

11.3 Preserving epistemic autonomy of allocators

We render verdicts. We do not render decisions. We resist the temptation to translate verdicts into implicit recommendations about capital allocation, deployment sequencing, or portfolio construction. Those decisions belong to the allocator.

An allocator who trusts our verdicts too much becomes dependent on us. That dependency is bad for the allocator and, in the end, bad for the industry. Our reports are designed to support the allocator's independent judgment, not to substitute for it.

11.4 Contestation of a verdict

A submitter or an allocator who considers that a Nyalai verdict is incorrect on its own terms retains a channel to contest it. A written objection may be filed within thirty days of the verdict-form's publication date, stating the specific gate result, threshold, or verdict-form claim that is contested and the basis of the contest.

A second validator instance, structurally distinct from the primary validator (that is, a validator instance that did not contribute to the original verdict and that forms its judgment independently, without access to the primary validator's internal deliberation), reviews the objection against the verdict-form and the underlying data package.

The outcome is published as a signed addendum attached to the original verdict-form. The addendum may confirm the original verdict, revise it, or withdraw it. The addendum is public, dated, and permanent: it never replaces the original verdict-form, which remains in the public record with the addendum appended.

Contestation is not appellate in the legal sense ; it is an adversarial re-check within the same doctrinal framework. A verdict confirmed by a contestation addendum carries the additional signal of having survived a structured objection. A verdict withdrawn by a contestation addendum is treated as a Nyalai correction.

12. Concluding thoughts

12.1 Open problems we acknowledge

  • Calibration under regime transitions is imperfectly measured. Our Gate 4 assumes stationarity within regime; it does not fully address strategies whose calibration itself is regime-dependent.
  • The guided-search correction in Gate 3 depends on the honesty of the strategy author about the number of configurations tested. We cannot audit undisclosed exploration.
  • Cross-regime detection via Hidden Markov Model is one of several available techniques. We commit to publishing a benchmark against alternative techniques by the end of 2027.
  • Our five gates are calibrated to strategies operating on daily-to-weekly time horizons. A future extension will introduce tiered validation levels adapted to sub-second, sub-minute, and intra-day regimes.
  • Real-time detection of strategy crowding is beyond the state of the art, and Nyalai does not claim it. Our verdicts address whether a strategy is calibrated and whether it possesses a coherent source of edge, not whether the strategy is uncrowded at the moment of deployment. The Doctrine treats crowding as an allocator's responsibility, monitored via the sizing envelope of Section 9.7 and the periodic revalidation cadence of Section 10.

12.2 What we don't know yet

  • The rate at which validated strategies degrade after receiving a GO verdict.
  • The correlation between Nyalai verdicts and allocator investment decisions in practice.
  • Whether the industry adopts our vocabulary, our thresholds, and our approach, or whether Nyalai remains a specialized service reserved for a specific institutional niche.

12.3 The relationship between Nyalai and the industry we serve

We are not the industry. We are external to it. We do not manage capital. We do not employ portfolio managers. We do not participate in the returns generated by strategies we validate.

This independence is structural. It cannot be maintained by intention alone. It requires a business model that does not create incentive alignment between us and the entities we evaluate.

Our current model preserves this independence through three principles. Verdict pricing is decoupled from verdict outcome (no performance-based compensation, no contingent fees). Subscription tiers with declared cadences replace ad-hoc contracts. Subscribers may commission supplementary verdicts at the flat per-verdict rate; non-subscribers pay the full flat rate. If we discover that this model creates unforeseen conflicts, we will document them here.

12.4 Sequence commitment

Nyalai never requests permission to publish a case, or to use it in any downstream decision, before a verdict is rendered. Consent from a submitting party regarding downstream use is collected after the verdict, on a document the submitter reviews. Refusal of downstream use does not alter the verdict itself.

This is not a mental-state promise. It is an order-of-operations commitment that Nyalai controls and that any third party can audit. Its purpose is to prevent a class of validity failure in which the submitter, aware of what will happen with the result, revises the input to shape the outcome. The submitter's decision about publication cannot retroact on what they wrote, because it comes after the verdict is fixed.

The commitment applies to every downstream use of a submission: public case study publication, private client disclosure, use in aggregate methodology research, external reference in any communication. All are conditioned on consent collected after the verdict is rendered. This is a discipline of validity, not of virtue. A submission contaminated by anticipation of downstream use is a false submission, and a verdict rendered against it would carry error we introduced ourselves.

Long-horizon commitment. The sequence commitment extends beyond the immediate verdict-form to the multi-year integrity of the register. A verdict published under v0.3.2 in 2026 remains verifiable against v0.3.2 in 2036, regardless of subsequent doctrine evolution. This posture draws on the ten-thousand-year marker discipline developed by Sandia National Laboratories in the 1990s for the United States Department of Energy's Waste Isolation Pilot Plant, treated at length by Peter Galison in the Nauenberg History of Science Lecture at UC Santa Cruz in April 2024. Nyalai's ledger operates on a shorter horizon but on the same discipline: the validator does not retroactively unpublish, does not retroactively rescore, and the register accumulates permanently.

12.5 On consensus averaging

When a submission produces multiple Gate 3 permutation runs across distinct evaluation configurations (for example three walk-forward windows or three block-bootstrap seeds), averaging the pass or fail arithmetic across the runs is legitimate only under three conditions. First, all runs were pre-registered as part of the verdict-form's methodology annex, so the averaging is not a post-hoc smoothing to obtain a favorable score. Second, no single run passes while the average fails, and no single run fails while the average passes; when either disagreement occurs, the verdict lists each run's outcome and the validator does not average. Third, the averaging follows the conservative protocol adopted by the Event Horizon Telescope collaboration in late July 2018: the average is used because it suppresses idiosyncratic single-run artifacts while preserving features common across runs.

The epistemic anchor. Ramesh Narayan articulated the principle behind consensus averaging with unusual clarity, as attributed by Peter Galison at the Harvard CMSA Yip Lecture of 18 April 2019: "the places where they agree should tell us what we can really believe." Nyalai's averaging discipline is an operational implementation of this anchor. The places where independent gate arithmetics agree are what the verdict can defend. The places where they diverge are what the verdict must report as unresolved uncertainty rather than smooth away.

Conservative framing over inclusive framing. Galison, present in the collaboration when the choice was debated, has publicly stated that the conservative-averaging framing won because shared features would show up right and features one method got but nobody else got would be relatively suppressed. Averaging that suppresses idiosyncratic single-method artifacts is legitimate. Averaging for collegial inclusion of the whole team, or for newspaper simplicity, is not.

Top-set reporting over point-estimate reporting. The averaging protocol shall not collapse the underlying distribution into a single point estimate. Where the mechanical layer's arithmetic produces a distribution, the verdict-form reports the distribution (for example a histogram of the ring-diameter posterior, in the Event Horizon Telescope precedent). The point estimate, when given, is annotated as a summary of the distribution and not as a substitute for it.

12.6 Regulatory positioning

Nyalai enters the market at the leading edge of a regulatory wave that is likely to require external validation of AI-produced institutional analysis. The current perimeter of SR 11-7 in the United States and equivalent guidance in other jurisdictions addresses model risk within institutions. The extension of that perimeter to external validation of AI-produced analyses is, based on the arc of the last fifteen years of prudential regulation, a matter of when rather than if.

We do not lobby for that extension. We build for it. Institutions that adopt Nyalai's discipline in advance of any regulatory mandate secure a workflow that will be recognized when the mandate arrives. The Doctrine's position is that early adoption of external adversarial validation is a discipline in itself, independent of whether regulation ever prescribes it.

The PRA four-phase policy cycle. The Prudential Regulation Authority's "Approach to Policy" framework (PS 3/25, February 2025) codifies a four-phase cycle (initiation, development, implementation, evaluation) with a statutory Cost-Benefit Analysis and an independent CBA Panel. Nyalai adopts a materially similar posture at doctrine-authoring scale: every version bump names the trigger (initiation), documents the changes with rationale (development), applies the changes to a live validator session before ship (implementation), and records the outcome for the following version's reference (evaluation).

12.7 On the word "Doctrine"

We chose the word "Doctrine" over "Manifesto," "Framework," or "Methodology" deliberately. A methodology is a set of techniques. A framework is a scaffolding for organizing techniques. A manifesto is a declaration of values. A doctrine is a foundational statement of how we operate. It is meant to be lived rather than followed. It carries the weight of a founding document.

We do not intend the word to imply rigidity. This Doctrine is a living framework, expected to evolve.

12.8 A final word

We may be wrong about specific claims in this document. We may be wrong about the priority ordering. We may be wrong about the gates. If evidence accumulates that we are wrong, we will revise.

But we do not expect to be wrong about the founding asymmetry: that the cost of a false GO verdict (capital deployed against a broken strategy) is greater than the cost of a false NO-GO. As long as that asymmetry holds, refusal-first is the correct discipline.

We offer this Doctrine in that spirit.

12.9 Governance

A doctrine that regulates the conduct of a validator must itself be governed. The absence of a governance section in v0.3 through v0.3.2.3 was a load-bearing gap; v0.3.3 fills it and v0.3.3.1 ratifies the wording.

Version authority. The Doctrine is authored by Sebastien Assohou. Version bumps are approved by the author and sealed by SHA256 co-computation with an adversarial validator instance structurally separate from the authoring path. No version becomes canonical until both the SHA256 hash is co-computed and the adjudication of enumerated UNVERIFIED items is on record.

Succession. If the author becomes unavailable for a period exceeding sixty days without prior transfer of authority, version authority passes to the individual designated in a public succession statement archived at nyalai.com/doctrine/governance (page pending publication as of v0.3.3.1); the designated individual carries a thirty-day mandate to constitute a governance committee. In the absence of a succession statement, the Doctrine is frozen at its last sealed version and no further updates are canonical until a governance quorum is re-established. A frozen Doctrine remains applicable to existing verdicts and to new verdicts issued under its terms; it does not lapse. The sixty-day threshold is a floor for triggering freeze-by-default; any actual transfer of authority requires the succession statement to be in force.

External arbitration. A verdict that has been contested per §11.4 and whose contestation addendum is itself disputed by the submitter may be referred to an external arbitration panel drawn from a public roster of qualified adversarial validators outside the Nyalai institution (the roster is aspirational per §11.1, to be established as independent validators eventually operate in parallel). The panel's finding is advisory to Nyalai; Nyalai retains final authority to accept, revise, or decline the finding, and the acceptance or decline is published as a second addendum to the original verdict-form. External arbitration is an escape valve, not a routine channel, and its advisory-only character is named openly as a limitation of the mechanism.

Sunset. If Nyalai as an entity ceases operations, all sealed Doctrine versions and all published verdict-forms remain in the public record at their archival locations under CC0. The Doctrine's methodology becomes a self-contained public good that any successor validator may adopt, adapt, or diverge from. No verdict-form is retroactively withdrawn on Nyalai's sunset. The durability of this clause depends on an off-nyalai.com archive surviving the sunset event, which is the Nyalai-triggered Internet Archive Wayback Machine snapshot obligation in the archival mechanism paragraph below; the sunset guarantee is therefore conditional on that snapshot being active at the time of sunset.

Archival mechanism obligations. Nyalai commits to (a) maintaining nyalai.com/doctrine and nyalai.com/verdicts as the canonical archival locations for so long as Nyalai operates, (b) triggering an Internet Archive Wayback Machine snapshot on each version bump (pending activation as of v0.3.3.1), and (c) publishing a governance status page at nyalai.com/doctrine/governance (pending publication as of v0.3.3.1) summarizing current authority, succession, and sunset preparations.

Acknowledgments

This Doctrine is authored by Sebastien Assohou. The intellectual influences whose work informed its construction are named individually below. Click any card for the full contribution and primary source.

Placeholder monograms shown. Custom illustrations forthcoming.

The full narrative acknowledgment, in the medical-report register that structures the rest of the Doctrine, follows below.

Statistical and methodological foundations.

  • Marcos López de Prado, whose 2018 Advances in Financial Machine Learning (Wiley) and 2014 Deflated Sharpe Ratio paper with David Bailey (Journal of Portfolio Management, Vol 40 No 5) form the statistical foundation of Gates 1, 3, and 5.
  • The Hidden Markov Model primary literature applied to financial time series: Lawrence Rabiner's 1989 tutorial (Proceedings of the IEEE, Vol 77 No 2), James Hamilton's 1989 Markov-switching autoregression paper (Econometrica, Vol 57 No 2), and Andrew Ang and Geert Bekaert's 2002 interest-rate regime detection paper (Journal of Business and Economic Statistics, Vol 20 No 2) form the HMM apparatus applied at Gate 2. Retail-to-professional convergence on the HMM stack is cited as external corroboration that the framework is not idiosyncratic to Nyalai's approach.
  • David Kipping (Columbia Astronomy), whose 2016 Sagan Summer Workshop MCMC tutorial supplied the nested-sampling versus Markov Chain Monte Carlo analogy that anchors §9.12.
  • David Aronson and Timothy Masters, whose 2013 Statistically Sound Machine Learning provided the guided-search correction methodology used in Gate 3.
  • Glenn Brier, whose 1950 Verification of Forecasts Expressed in Terms of Probability (Monthly Weather Review, Vol 78 No 1) and its modern operationalization by Guo et al. (2017) form the calibration foundation of Gate 4.
  • Dimitris Politis and Joseph Romano, whose 1994 work on stationary bootstrap methodology forms the resampling foundation of Gate 5.
  • Campbell Harvey (Duke University), whose work on out-of-sample yield-curve model discipline and multiple-testing correction rigor for cross-sectional asset-pricing anomalies informs Gate 3 posture on directed-search correction and the general refusal-first discipline against p-hacked publications.
  • Andrew Ang (BlackRock), whose playlist-metaphor unbundling of index, factors, and alpha supplied the vocabulary discipline that keeps Nyalai's verdict-form legible to institutional readers.
  • Emanuel Derman (Columbia University), whose Modeler's Hippocratic Oath (co-authored with Paul Wilmott) and "not God's truth" model epistemic humility inform the calibrated-uncertainty properties of §6.2 and the framework-not-model positioning of §9.12.

Institutional model validation.

  • Joseph Simonian, whose 2020 paper "Modular Machine Learning for Model Validation" (Journal of Financial Data Science, Vol 2 No 2, pp 41-50) supplied the modular validation discipline of Section 2.8, the econometrically-informed block-size discipline in Gate 2, and the communication-layer principle in Section 9.7.
  • Roland Stamm, Partner in Acadia's Quantitative Services division (now LSEG Post Trade Solutions), whose public statements on the "framework, not a model" positioning of institutional validation shaped Section 9.12.
  • Senthil Kumar, formerly Chief Risk Officer at BNY Mellon (now Chief Risk Officer at Huntington Bank), whose triangulation framework across risk type, business context, and regulatory frame, articulated in the Columbia CRO Spotlight Series (2022), shaped Section 9.11.
  • The Prudential Regulation Authority (Bank of England), whose SS 1/23 model risk management principles (May 2023, updated April 2026), PS 6/23 policy statement, and PRA Approach to Policy PS 3/25 (February 2025) informed the regulatory-positioning discipline of §12.6. The independent Cost-Benefit Analysis Panel established under the PRA Approach to Policy is named here as a meta-doctrine reference for how a public policy body sustains adversarial calibration of its own outputs across versioning cycles.

Philosophy of science and empirical validation methodology.

  • Peter Galison (Joseph Pellegrino University Professor at Harvard), whose Objectivity (co-authored with Lorraine Daston, Zone Books, 2007), Harvard CMSA Yip Lecture (18 April 2019), Nauenberg History of Science Lecture at UC Santa Cruz (18 April 2024), and Event Horizon Telescope participation supplied the three-regime schema (truth-to-nature, mechanical objectivity, trained expertise), the parallel-groups precedent of §5.3, the conservative-averaging framing of §12.5, and the WIPP long-horizon signaling reference of §12.4.
  • Lorraine Daston (Director emerita, Max Planck Institute for the History of Science), whose co-authorship of Objectivity with Peter Galison established the three-regime schema onto which Nyalai's stack maps.
  • Sheperd Doeleman (Founding Director, Event Horizon Telescope; Harvard-Smithsonian), whose account of the EHT's multi-method convergence discipline supplied the convergence-precondition operating principle for the trained-expertise layer of §5.3.
  • Katie Bouman (Caltech), whose Caltech CMS colloquium admission that team isolation does not eliminate shared human bias supplied the residual-bias-floor annex requirement for §5.3 and the adversarial-shape variant discipline for §9.5a.
  • Ramesh Narayan (Harvard-Smithsonian Center for Astrophysics), whose epistemic principle "the places where they agree should tell us what we can really believe" (as attributed by Galison) supplied the anchor for the consensus-averaging discipline of §12.5.
  • Andrew Strominger (Director, Center for the Fundamental Laws of Nature, Harvard), whose public commentary on Einstein's 1939 paper on black-hole non-existence supplied the Einstein-1939 lesson anchor of §5.3.

Epistemological foundations.

  • Nassim Nicholas Taleb, whose 2008 Edge.org essay "The Fourth Quadrant: A Map of the Limits of Statistics" and its peer-reviewed follow-up in the International Journal of Forecasting supplied the conceptual anchor of Section 1.4.
  • Eugene Fama and Kenneth French, whose 2018 Financial Analysts Journal paper "Volatility Lessons" (Vol 74 No 3) supplied the statistical inference horizon framing of Section 2.7.

Fiduciary AI framework.

  • Andrew W. Lo (Charles E. and Susan T. Harris Professor at MIT Sloan School of Management, Director of MIT Laboratory for Financial Engineering, Principal Investigator at CSAIL), whose ongoing work with his named MIT CSAIL collaborators on retrieval-augmented generation for fiduciary AI describes the theoretical grounding for external validation of AI-produced institutional analysis, and whose public framing of the fiduciary AI problem in three components, competency (measured by exam performance including Series 65 and CFA-style tests, which GPT-4 passes), personalization (via randomized controlled trials on client-specific advice), and trust (systematized from the code of ethics for financial advisers and what Lo calls the "fossil record" of case law regarding fiduciary duty), informs the honesty properties of Section 6. Lo's Harvard Griffin GSAS talk "Can ChatGPT Plan Your Retirement?" delivered in 2026 articulates the essential Nyalai thesis in a single sentence: "There will come a time soon where we will feel trusting of these large language models. And that's dangerous. Because feeling like you can trust somebody does not mean that you should trust somebody."

Practitioner voices in asset allocation and manager selection.

  • Matt Bank (GEM, Deputy Chief Investment Officer), whose four-category framework for allocator risk (shortfall, drawdown, liquidity, embarrassment) articulated on Capital Allocators episode 419 shaped the communication-layer vocabulary of Section 9.7.
  • Nicholas Csicsko (Trinity Wall Street endowment), whose prepare-in-good-times principle and style-drift criterion articulated on the Capital Allocators Senior Decision Makers series shaped Section 11.2.
  • Pat Dorsey (Dorsey Asset Management), whose focus on management humility and the "obvious central failure" framing articulated on Capital Allocators episode 509 informs the diagnostic discipline of Nyalai's verdict-forms.
  • Dan Rasmussen (Verdad Advisers), whose 2025 The Humble Investor (Harriman House) and Verdad Weekly Research publications supplied the meta-analysis framing of Section 2.4.
  • Cliff Asness (AQR Capital Management), whose value-spread work, "half a backtest" out-of-sample discount rule (Capitalism and Freedom podcast episode 63, December 2025), and public treatment of speed-based versus structural edges informed Section 2.4 and Section 9.7. The Doctrine's position on real-time crowding detection in Section 12.1 is written in explicit agreement with the AQR position on the same question.
  • Antti Ilmanen (AQR), whose 2011 Expected Returns (Wiley) and 2022 Investing Amid Low Expected Returns (Wiley) established the asymmetric-patience framing that distinguishes public-market from private-market validation scope in Section 1.4.

Structural template and market microstructure.

  • The authors of Claude's Constitution (Anthropic, January 2026), whose structural template informed the organization of this document. Both are licensed under CC0.
  • Jean-Philippe Bouchaud, Julius Bonart, Jonathan Donier, and Martin Gould, whose 2018 Trades, Quotes and Prices: Financial Markets Under the Microscope (Cambridge University Press) informs the treatment of high-frequency trading extensions.

The errors in this Doctrine are ours alone.

Primary sources

Every citation in this Doctrine is traceable to a primary source: the author's own paper, book, podcast, essay, or public presentation. Selected bibliographic anchors follow; the complete register is maintained in the canonical markdown of THE_REFUSAL_DOCTRINE_v0.3.3.1.

Peer-reviewed papers.

  • Bailey, D. H., and López de Prado, M. (2014). The Deflated Sharpe Ratio. Journal of Portfolio Management, Vol 40 No 5, pp 94-107.
  • Brier, G. W. (1950). Verification of Forecasts Expressed in Terms of Probability. Monthly Weather Review, Vol 78 No 1.
  • Rabiner, L. R. (1989). A Tutorial on Hidden Markov Models and Selected Applications in Speech Recognition. Proceedings of the IEEE, Vol 77 No 2, pp 257-286.
  • Hamilton, J. D. (1989). A New Approach to the Economic Analysis of Nonstationary Time Series and the Business Cycle. Econometrica, Vol 57 No 2, pp 357-384.
  • Ang, A., and Bekaert, G. (2002). Regime Switches in Interest Rates. Journal of Business and Economic Statistics, Vol 20 No 2, pp 163-182.
  • Cochran, W. G. (1954). Some Methods for Strengthening the Common Chi-Squared Tests. Biometrics, Vol 10 No 4, pp 417-451.
  • Fama, E. F., and French, K. R. (2018). Volatility Lessons. Financial Analysts Journal, Vol 74 No 3.
  • Kelly, B., Malamud, S., and Zhou, K. (2024). The Virtue of Complexity in Return Prediction. Journal of Finance, Vol 79 No 1, pp 459-503.
  • Politis, D. N., and Romano, J. P. (1994). The Stationary Bootstrap. Journal of the American Statistical Association, Vol 89 No 428, pp 1303-1313.
  • Simonian, J. (2020). Modular Machine Learning for Model Validation. Journal of Financial Data Science, Vol 2 No 2, pp 41-50.
  • Taleb, N. N. (2009). Errors, Robustness, and the Fourth Quadrant. International Journal of Forecasting, Vol 25 No 4.

Books.

  • Aronson, D. R., and Masters, T. (2013). Statistically Sound Machine Learning for Algorithmic Trading of Financial Instruments.
  • Bouchaud, J.-P., Bonart, J., Donier, J., and Gould, M. (2018). Trades, Quotes and Prices: Financial Markets Under the Microscope. Cambridge University Press.
  • Galison, P., and Daston, L. (2007). Objectivity. Zone Books. Cited in §5.3 historical grounding and trained-expertise layer.
  • Ilmanen, A. (2011). Expected Returns: An Investor's Guide to Harvesting Market Rewards. Wiley.
  • Ilmanen, A. (2022). Investing Amid Low Expected Returns. Wiley.
  • López de Prado, M. (2018). Advances in Financial Machine Learning. Wiley.
  • Rasmussen, D. (2025). The Humble Investor. Harriman House.

Public presentations and institutional documents.

  • Galison, P. (18 April 2019). Harvard CMSA Yip Lecture, "Philosophy of the Shadow." Cited in §5.3, §9.5a, and §12.5.
  • Galison, P. (18 April 2024). Nauenberg History of Science Lecture, UC Santa Cruz. Cited in §12.4 long-horizon commitment (WIPP treatment) and §12.5.
  • Kipping, D. (18 July 2016). A Beginner's Guide to MCMC. Sagan Summer Workshop. Cited in §9.12 nested-sampling analogy.
  • Sandia National Laboratories (1993). Report SAND92-1382, ten-thousand-year marker design for the United States Department of Energy's Waste Isolation Pilot Plant. Referenced in §12.4.
  • Federal Reserve Board and Office of the Comptroller of the Currency (2011). SR 11-7: Guidance on Model Risk Management.
  • Prudential Regulation Authority, Bank of England: SS 1/23 Model Risk Management Principles for Banks (May 2023, updated April 2026); PS 6/23 Policy Statement (May 2023); PS 3/25 The Prudential Regulation Authority's Approach to Policy (February 2025). Referenced in §12.6.
  • Central Bank of the U.A.E. (November 2022). Model Management Standards, Notice 505/2022.
  • Anthropic (January 2026). Claude's Constitution. Published under CC0.

Correction protocol: if any citation above is found to be incorrect, the Doctrine will be revised within 24 hours of the finding, with a versioned changelog entry documenting the correction.

This section, added in v0.3.3, lists non-authoritative works that a reader coming from the trade press, business books, or podcast circuit will recognize. It is listed to establish shared vocabulary with allocators and the podcast circuit and to bridge from popular reference to the peer-reviewed anchor without ceding methodological authority to the popular source. Nyalai does not cite these works as authorities for methodological claims and does not cite them in verdict arithmetic.

Books.

  • Gregory Zuckerman (2019). The Man Who Solved the Market: How Jim Simons Launched the Quant Revolution. Portfolio (Penguin Random House). The popular reference an allocator will most often mention. Its "Markov chains and mean reversion" framing is a useful entry point, though Gate 2 methodology draws on the peer-reviewed HMM literature (Rabiner 1989, Hamilton 1989, Ang and Bekaert 2002) rather than the trade-press framing. Where a submitter cites Zuckerman as the intellectual basis for a Markov-chain strategy, the verdict-form treats the citation as vocabulary alignment rather than as methodological authority.

Podcasts.

  • Capital Allocators (host: Ted Seides). Extensive practitioner voice from asset allocators and manager selectors. Individual episodes cited in Acknowledgments when directly informative of a specific doctrine section.

Educational YouTube.

  • Quant Guild (channel @QuantGuild, run by Roman Paolucci). Markov chains, hidden Markov models, and ARCH/GARCH volatility modeling for quantitative finance. Retail-to-professional bridge material widely watched by junior quantitative staff at institutions.
  • QuantProgram (channel @quantprogram). Jim Simons Medallion Markov strategy and related trading-channel material. Where a submitter cites this class of material, the verdict-form aligns to the vocabulary but redirects to the peer-reviewed anchor.

The section is short by design: it is a bridge, not a competing bibliography.

Every citation in this Doctrine is traceable to a primary source: the author's own paper, book, podcast interview, essay, or public presentation. No secondary interpretations or intermediary paraphrases are treated as source material.

Peer-reviewed papers. Bailey and López de Prado (2014) Journal of Portfolio Management Vol 40 No 5. Brier (1950) Monthly Weather Review Vol 78 No 1. Fama and French (2018) Financial Analysts Journal Vol 74 No 3, DOI 10.2469/faj.v74.n3.6. Guo et al. (2017) On Calibration of Modern Neural Networks, ICML, arXiv:1706.04599. Politis and Romano (1994) Journal of the American Statistical Association Vol 89 No 428. Simonian (2020) Journal of Financial Data Science Vol 2 No 2 pp 41-50. Taleb (2009) International Journal of Forecasting Vol 25 No 4.

Books. Aronson and Masters (2013) Statistically Sound Machine Learning for Algorithmic Trading of Financial Instruments. Bouchaud, Bonart, Donier, and Gould (2018) Trades, Quotes and Prices, Cambridge University Press, ISBN 9781107156050. Ilmanen (2011) Expected Returns, Wiley, ISBN 9781119990727. Ilmanen (2022) Investing Amid Low Expected Returns, Wiley, ISBN 9781119860198. López de Prado (2018) Advances in Financial Machine Learning, Wiley, ISBN 9781119482086. Rasmussen (2025) The Humble Investor, Harriman House, ISBN 9781804090763.

Essays and public presentations. Taleb (2008) The Fourth Quadrant, Edge.org. Asness (December 2025) Capitalism and Freedom podcast episode 63. Bank (2025) Capital Allocators episode 419. Csicsko Capital Allocators Senior Decision Makers series. Dorsey (2024) Capital Allocators episode 509. Kumar (2022) Columbia CRO Spotlight Series. Stamm, Ahead of the Curve podcast, LSEG Post Trade Solutions.

Institutional documents. Federal Reserve Board and OCC (2011) SR 11-7 Guidance on Model Risk Management. Anthropic (January 2026) Claude's Constitution, CC0.

If any citation above is found to be incorrect, the Doctrine will be revised within 24 hours of the finding, with a versioned changelog entry documenting the correction.

Trust is not a result of validation.
It is the product of the rigor with which we accept to contradict ourselves.

Refusal Doctrine v0.3.3.1 · July 19, 2026 · Sebastien Assohou · Nyalai

[email protected]

This Doctrine is published under CC0 1.0 Universal Public Domain Dedication.

Contribution

Background

Doctrine reference

Back