Survivorship Bias Means You Are Only Studying the Ones That Made It
Datasets quietly delete their failures. Any conclusion drawn from what remains describes a population that was selected for success.
The Mechanism
Survivorship bias occurs when a sample includes only entities that made it to the present and those that did not are absent rather than being recorded as failures
Exclusion is rarely deliberate. It is what happens naturally when a database contains currently existing items. Funds that closed are no longer listed. Companies that went bankrupt are no longer listed in the index. Strategies that stopped working are no longer traded
Each average calculated from that sample is an average of survivors
What sets it apart from most analytical errors is that nothing in the data looks wrong. There is no outlier to detect no suspicious values no failed checks. The numbers present are precise. The problem lies entirely in the rows that were never written and no examination of the rows that exist will reveal it
The Wartime Example
The clearest example comes from a statistical problem during World War II. Analysts examined returning planes mapped where bullet holes clustered and proposed reinforcing those areas
Statistician Abraham Wald pointed out the error. The holes marked places where a plane could be hit and still return. The areas without holes in the surviving planes were precisely the areas where the damage was fatal and those were the areas to be armored
The missing observations carried the information. That's the general form of the problem and the missing observations are always the hardest to notice
The error structure is worth mentioning because it is repeated in all financial versions. Analysts selected a sample based on the outcome they were trying to study. Survival determined membership in the data set and survival was the exact variable of interest. Once this is true the sample cannot answer the question no matter how large it is or how carefully it is analyzed
Where It Distorts Finance
| Settings | what disappears | Effect |
|---|---|---|
| Fund databases | Closed and liquidated funds | Category returns are overstated |
| Index history | Delisted and bankrupt members | Exaggerated long-term returns |
| Backtesting | Companies that are no longer listed | Inflated strategy performance |
| Manager history | Discontinued strategies | Exaggerated skill |
Fund databases are the most studied case. Funds close primarily because they performed poorly so removing them eliminates the worst results. Studies estimating the size of this effect have found that it adds a significant amount to the category's reported returns enough to change conclusions about whether a category outperformed
The Arithmetic of the Inflation
The size of the distortion is not a matter of opinion. It is deduced from two quantities: how many entities disappeared and how badly they were doing when they disappeared
Take one hundred funds at the beginning of a decade purely as arithmetic. Suppose 30 of them close during the period and suppose that the 70 survivors averaged 8 percent per year while the 30 that closed averaged 2 percent per year during the period they were open
A database of currently existing funds reports the category as 8 percent. The average of everything an investor could have chosen at the beginning is 0.7 x 8 plus 0.3 x 2 which is 6.2 percent. The stated figure is 1.8 percentage points higher per year and the investor had no way of knowing in advance which group he was buying into
Aggravated that gap really works. Ten thousand dollars growing at 8 percent for ten years becomes about $21,589. At 6.2 percent it amounts to about $18,249. The object of measurement is worth about $3,340 in a ten thousand dollar position and there is none of that
Two things drive the result and both are structural rather than accidental. Attrition rates in fund categories are not small over long horizons and closing is strongly correlated with poor performance since a fund that is doing well is not the one that is liquidated. Therefore the bias is always in the same direction. You never flatter a lower category
Index History Is Rebuilt by the Winners
The index version is more subtle than the fund version because indices are not databases of survivors. Lists are maintained with documented additions and deletions and the historical series is usually calculated correctly
Distortion occurs through how the list is maintained and not through missing records. Companies are added after they have grown enough to qualify meaning their growth occurred elsewhere and off the index books. Companies are removed after they have downsized been acquired or failed meaning the decline is recorded up to the point of removal and then stops
Therefore the list is continually updated with what has performed recently. This is a reasonable design for an index intended to represent the current market. It is a poor basis for stating what an investor would have earned by holding a fixed set of companies from the beginning and the two are constantly mixed up
The error is compounded when a current voter list is applied in reverse. To study the performance of current members of a large index over the past twenty years is to study a group selected as having arrived at the present with sufficient size. Companies that were in the index twenty years ago and are not now have been silently excluded from the study by the act of choosing the sample
The Universe Is Chosen Before Any Model Runs
Every quantitative exercise begins with a decision that no one carefully documents: what values are eligible to appear in the data
That choice is made before a single calculation is performed and determines the limit of how honest the result can be. A universe assembled from a current listings file has already eliminated all bugs. A universe assembled from a point source where each date contains the values that existed and were invertible on that date does not
Proper databases include delisted securities with the reason and date which is what allows a test to distinguish a company that was acquired at a premium from one that went to zero. They both leave the data on the same day and are not the same event for whoever owned the shares
Practical failure is usually mundane. You download a file it contains the companies that exist today you proceed to the analysis and the exclusion is never a conscious decision for anyone. Nothing in the workflow raises the question so it must be deliberately asked from the beginning
Backtesting Is Where It Bites
Testing a strategy with historical data requires a universe of securities to test on. If that universe is built from companies that exist today all the companies that failed during the testing period are missing
A strategy that would have bought companies that later went bankrupt never has the opportunity to buy them because they are not in the data. The backtest reports returns that no one trading at the time could have achieved
Using a database without delisted stocks is a common and serious mistake and favors small value-oriented company strategies the most since those categories contain the highest failure rates. A rule that buys the cheapest quartile by some valuation measure is by definition buying the companies that the market has marked down and a significant proportion of them are marked down because they are in trouble. Eliminate the ones that didn't make it and the rule looks like a method of finding bargains rather than a method of containingthe difficulties
The Incubation Variant
A related practice makes it worse. A company quietly launches several small funds manages them for a period then markets the ones that performed well and closes the rest
The surviving fund has a genuine track record. It is also the result of a selection process that the investor cannot see and its history is the result of picking winners after the fact rather than skill applied in advance
The same shape appears in the individual histories without anyone pitching anything. A manager who has applied several strategies throughout his career presents the one that worked and the presentation is precise in every detail. What is missing is the denominator: how many strategies were executed during what period and what happened to the others
Beyond Finance
The pattern generalizes. Studies of successful companies that look for common traits examine only the successes so any trait shared by both successes and failures appears to be a cause of success
The same applies to advice from successful founders investors and executives. The sample was taken entirely from people whose approach worked and the same focus on people for whom it didn't work produced no books
That's why the genre is so persistently unfalsifiable. Any recommended trait can be confirmed by examining more successes and the only test that would disprove it checking how common the trait is among failures requires a data set that by definition no one collected
It Understates Risk as Well as Return
The returns problem draws attention and the risk problem is arguably worse because it is almost never mentioned
Observations removed by survival are not randomly scattered in the distribution. They are concentrated in the left tail since the entities that disappeared are largely the ones that lost the most before disappearing. If they are removed the remaining sample will have a shorter tail a smaller worst case and a smaller maximum reduction than the population it claims to describe
Measured volatility falls for the same reason. Dispersion is calculated from present observations and the most violent observations were those linked to entities that failed. Therefore the reported risk of a category is the risk experienced by members who were able to continue experiencing it
Correlations also become distorted and in the direction that matters most. The benefits of diversification are typically estimated from how a set of holdings performed together during past stresses. If the holdings that failed under that stress have been removed from the sample the estimate describes the subset that was held which is precisely the subset whose behavior is least informative about what will happen next time
The compound effect is that every risk-adjusted figure constructed from the sample inherits both errors at once. Return is overstated in the numerator and dispersion is understated in the denominator so any ratio that divides one by the other is wrong twice and in the same direction. A category may appear to have generated strong returns with modest risk when what actually happened is that episodes of severe risk removed its own evidence from the record
How to Check
Ask what happened to entities not in the data. Ask if the database includes dead stocks and closed-end funds. Ask when the sample was selected relative to the period being studied. And be wary of any long-term return figures in which the list of voters is drawn up at the end of the period rather than the beginning
Two follow-up questions do most of the remaining work. What was the attrition rate over the period since a category that loses a large proportion of its members has a large correction pending and one that loses very few does not? And demise was correlated with performance since entities that disappear for reasons unrelated to results such as a fund closing because its manager retired have no bearing
When failures truly cannot be recovered the honest response is to expand the uncertainty around the conclusion rather than proceeding with the average number of survivors and a warning in a footnote. The figure is not just uncertain. It is known to be wrong in a known direction which is more information than most warnings convey
The Bottom Line
Survivorship bias makes any sample of survivors look better than the population from which it came because failures were removed rather than recorded. It inflates fund category returns index histories backtesting and all the conclusions drawn from studying successful companies. The correction is always the same: find out what is missing and whether it was removed for performance-related reasons. The distinguishing feature of this error is that the surviving data appears perfectly clean so the checking must be done deliberately inthe moment when the universe is chosen that is before anyone has done any analysis worth defending