Hedge Fund

The Model Was Right, and Then the World Stopped Resembling Its Training Sample

Model risk is the possibility that a model is wrong, used wrongly, or trusted beyond what it can support. It is a recognised discipline in banking because ignoring it has repeatedly been expensive.

↩ Looking BackPart of the 2020 to 2026 retrospective, written in July 2026. The date below marks the 2025 events this piece revisits, not when it was published, so it draws on everything known through mid 2026.
Nathan Xiang·May 14, 2025

What Model Risk Means

Model risk It is the risk of loss from decisions based on the results of the model. It has three distinct sources and treating them as one thing is in itself a common failure

The model may be fundamentally flawed and based on assumptions that do not hold. It may be correct but implemented with errors. Or it may be correctly and correctly implemented but applied to a situation for which it was never designed which is the most common of the three

The third category is what this article is primarily about because it's the only one that arrives after everything has been done correctly. A model that was validated approved and worked well can become incorrect without anyone touching it simply because the world it describes has moved. Nothing failed. The environment changed beneath a description that was accurate when it was written

Why Every Model Is Wrong Somewhere

A model is a simplification. It captures the relationships that are considered important and discards the rest and that discarding is what makes it useful. The problem is that the discarded factors are only irrelevant under the conditions under which the model was developed

The classic example is any model that estimates the correlation between assets from historical data. In normal periods diversification works and correlations are moderate. In crises correlations increase sharply when investors sell everything at once. A risk model calibrated in normal periods will systematically underestimate risk precisely at the moment when risk matters because the relationship it learned no longer holds

Models fail along with what they are modeling. The conditions that make a model wrong are often the same conditions that make being wrong costly

The Shapes That Drift Takes

Degradation in a living model is not a unique phenomenon and separating the varieties is important because they are detected differently

Inputs can move. The population that arrives at the model is no longer the population on which it was built either because a company entered a new market changed its prices or acquired customers through a different channel. The relationships that the model learned may still be perfectly valid and it is asked about cases it never saw

The relationship itself can change. The inputs look very much like they always did and what they imply has changed because the behavior changed or the economic regime changed. This is the most difficult case since each input distribution check turns out clean

And the model may be right about a world it helped create. A pricing model that sets prices or a credit model that decides who gets a loan determines what observations exist to measure later. Rejected applicants do not generate payment history so the data available for the next recycling has been filtered by the current model's own decisions

The latter is the most insidious because it confirms itself. A model that wrongly rejects a category of borrower never accumulates the evidence that would show that the rejection was wrong and each subsequent review of its performance on the loans it made turns out to be excellent

The Machine Learning Version

More flexible models make this more precise than certain. A model with enough parameters can fit the historical data almost perfectly including noise which is overfitting.Then performs poorly on new data because it learned accidents from the sample rather than long-standing relationships

These models also degrade silently. A traditional model with explicit assumptions can be compared to those assumptions. A model that learned patterns without expressing them does not give any obvious signal when the patterns stop applying which is called drift and detecting it requires deliberate monitoring rather than waiting for complaints

The absence of stated assumptions is the whole difficulty. A model built on an explicit statement about say the relationship between two rates can be checked by asking whether that statement is still valid. A model that discovered its own relationships offers nothing to test itself against so the only check available is whether its results still work and that check requires waiting for the results it was predicting

How Drift Is Actually Detected

The tracking problem has a practical structure and useful signals arrive in a specific order

Input monitoring is available immediately and is the weakest evidence. Comparing the distribution of each input with the distribution the model was trained on detects the case where the population has changed and says nothing about whether the change is important

Tracking the predictions is as follows. Looking at the distribution of the model's own outputs detects a model that has started to behave differently even when no input seems unusual which is often the first visible symptom of a change in the relationship

Tracking results is the only measure that answers the real question and is delayed by the time it takes for the result to arrive. For a short-horizon forecast it is days. For a credit model it is years which means that definitive evidence about the current quality of a model will not be available until long after decisions have been made

That delay is why weaker upward signals are worth the effort. They are early and ambiguous and they are what exists in the window between a model going wrong and the losses that make it obvious

How the Discipline Responds

Banking regulators require a formal risk management model and the structure is worth knowing because it generalizes far beyond finance

controlPurpose
Independent validationReviewed by people who didn't build it.
Documented limitationsStates where the model should not be used
Continuous monitoringDetect drift before it causes losses
Model InventoryNo one can protect a model that no one knows exists.

Independent validation is load control. The people who built a model are not in a position to find its defects as they have convinced themselves. Review by a separate team with authority to reject is what makes the process real and not procedural

The word authority is where control often fails. A validation function that informs the company that it reviews or that can be overridden by the revenue it restricts produces documentation rather than questioning. The observable test is whether the function has actually ever stopped something valuable since a review process without record of rejects has not demonstrated that it can do so

Model inventory sounds like management and it isn't. Models proliferate in spreadsheets in scripts on individual machines and in internal vendor systems whose internal components no one has looked at and each of them is making decisions. An organization that can't enumerate its models can't validate them monitor them or know which ones depend on an assumption that just stopped being valid

Retraining Is Not a Solution By Itself

The obvious response to drift is to retrain on recent data which helps and introduces issues of its own worth mentioning

Retraining often means that the model in production is never the model that was validated. In fact a process that automatically updates on a schedule has replaced a revised artifact with an unreviewed one and the governance built around the original approval no longer applies to what is being executed

It also trades one failure for another. A model updated in a recent short window adapts quickly and forgets the conditions it previously saw so it becomes too attuned to the current regime and loses all the robustness it gained from having been installed in several. Retraining after a long lull produces a model that has quietly discarded its knowledge of stress

And to the previous feedback point retraining on data generated under the current model's own decisions amplifies any bias that model had rather than correcting it

The workable compromise is to treat a retrained model as a new model that requires the same checks maintain a stable baseline version to compare against and keep the training window long enough to include conditions that the current period does not contain

When Everyone Runs the Same Model

So far the debate treats a model as a private tool within an organization. Much of the risk of the model is not private and the shared version behaves differently

Widely adopted models built from public data produce similar results at similar times for many participants at once. Risk measures calculated from the same price history credit assessments based on the same ratings and the same reported financial data and valuation models calibrated with the same observable market data converge on similar responses because they are similar procedures reading the same source

What that means in practice is that the response to a change is synchronized. When shared inputs move enough to change what the models say a large number of institutions receive the same instruction in a short period of time and acting on it means transacting in the same direction simultaneously

The impact of that collective action on the market feeds back into the data read by the models constituting a naturally unbuffered loop. The model was validated on its own merits and its correctness at the level of a company has very little to do with what happens when everyone adopts it

This is the category of model risk that no internal governance addresses since each individual validation can be completely robust. The question it raises is not whether a model is correct but how many other people are running something close enough to it to act at the same time and that's not a question a validation team is in a position to answer

The Failure That Is Not Technical

The biggest flaws in models tend to be organizational rather than mathematical. A model produces a number the number appears in a report and everyone in the process treats it as a fact rather than an estimate that involves assumptions

That's why documented limitations are so important. The person who decides based on the output of a model is usually several steps away from anyone who understands what it entails and by then the uncertainty has been eliminated and only the number remains

Incentives make it worse. When a model produces favorable results such as lower capital requirements or higher valuations the pressure to question it is weak and the pressure to accept it is strong

The asymmetry worsens over time because scrutiny is applied unevenly. A model that produces inconvenient answers is frequently questioned revised and replaced so unfavorable errors are quickly corrected. A model that produces desirable answers is left alone so favorable errors persist. Therefore the surviving stock of models in an organization leaks in one direction and no one made that decision

What Reasonable Use Looks Like

The practical stance is to treat a model as an argument rather than an answer. Know what assumptions it depends on test how much production moves when they change and pay special attention to results that are surprisingly favorable since they are the least likely to be examined

Simplicity has real value here. A simpler model that is understood is usually safer than a more precise one that is not understood because someone can tell when it no longer applies

The strongest habit is to note at the moment of approval what would have to be observed for the model to be considered no longer applicable. It is the same discipline as establishing a falsification condition before performing an experiment and it works for the same reason. The judgment is made by someone who does not yet have a position to defend and turns a vague intention to monitor into something specific to pay attention to

The Bottom Line

Model risk is not about arithmetic errors. It is about applying a simplification outside the conditions under which the simplification was met and doing so with confidence because the result came as a precise number. Defense is about knowing the limitations having someone independent review them and treating a surprisingly good result as cause for scrutiny rather than celebration. A model does not announce the moment when it stops describing the world and the definitive evidence it has usually comes after the decisions that depended on it

Explore Teen Biz News →