Sample Size Is Why Three Good Years Prove Nothing
Distinguishing skill from luck in investment returns requires far more data than anyone has. The arithmetic on how much is genuinely uncomfortable.
The Question
A manager has beaten their benchmark by 2 percent a year for three years. Is that skill?
To answer, you need to know how likely that result would be from luck alone. That depends on the volatility of the excess return, and for typical active equity strategies the annual variability of relative performance is substantially larger than the average outperformance being claimed.
When the noise is several times the signal, three observations are not close to enough.
The Uncomfortable Arithmetic
The number of years required to establish statistical confidence rises with the square of the ratio between noise and signal. That relationship is brutal.
| Excess return | Tracking error | Years for reasonable confidence |
|---|---|---|
| 2 percent | 4 percent | Roughly 16 |
| 2 percent | 6 percent | Roughly 36 |
| 1 percent | 5 percent | Roughly 100 |
These figures are approximate and the message does not depend on the precision. The horizons required exceed most careers, and they exceed the patience of almost every allocator.
By the time a track record is long enough to be statistically meaningful, the manager has often retired, the strategy has changed, or the market regime it worked in has ended.
Why Everyone Ignores It
Institutional practice evaluates managers over three to five year periods. That convention exists for governance and career reasons rather than statistical ones.
The consequence is predictable. Managers are hired after strong periods and fired after weak ones, and studies of institutional hiring and firing decisions have found that the fired managers frequently outperformed the newly hired ones over the following years.
That is what you would expect if the decisions were being driven by noise, since noise mean reverts.
The Multiple Manager Problem
The situation is worse than the single manager arithmetic implies, because allocators evaluate many managers.
With thousands of funds operating, some will produce long strong records by chance alone. Selecting the best performer from a large population and treating their record as evidence of skill ignores that the selection was made from the top of a distribution.
This is why an impressive record, considered in isolation, is weak evidence. The relevant question is how many managers were in the population from which this one was chosen.
What Helps When Data Is Short
Since waiting decades is not available, the practical approach is to gather evidence that is not return based.
Understand the process and whether it has an economic rationale. Examine the individual decisions rather than only the aggregate outcome, since a manager who was right for the stated reasons is different from one who was right accidentally. Look for consistency of approach across periods. Assess whether the same people are still making the decisions.
Increasing the number of independent observations also helps. A strategy making many decisions a year generates more evidence per year than one making a handful of large concentrated bets, which is why a high turnover systematic strategy can be assessed faster than a concentrated discretionary one.
The Same Problem in Personal Investing
An individual assessing their own ability faces exactly the same arithmetic with an even smaller sample. A few years of good returns during a rising market is not evidence of skill, and there is no way to make it into evidence.
This is one of the strongest arguments for a mechanical process. If the sample is too small to tell whether you have an edge, a rules based approach at least removes the risk of confidently acting on an edge you do not have.
The Bottom Line
Establishing statistical confidence in investment skill requires far more data than typical track records contain, because the noise in returns is several times the size of the signal. Standard three year evaluation windows measure luck. Since more data is not available, the substitute is evidence about process, decisions, and personnel, plus honest awareness of how many candidates the impressive record was selected from.