Hedge Fund

Sample Size Is Why Three Good Years Prove Nothing

Distinguishing skill from luck in investment returns requires far more data than anyone has. The arithmetic on how much is genuinely uncomfortable.

↩ Looking BackPart of the 2020 to 2026 retrospective, written in July 2026. The date below marks the 2022 events this piece revisits, not when it was published, so it draws on everything known through mid 2026.
Nathan Xiang·February 9, 2022

The Question

A manager has beaten their benchmark by 2 percent a year for three years. Is that skill?

To answer, you need to know how likely that result would be from luck alone. That depends on the volatility of the excess return, and for typical active equity strategies the annual variability of relative performance is substantially larger than the average outperformance being claimed.

When the noise is several times the signal, three observations are not close to enough.

The Uncomfortable Arithmetic

The number of years required to establish statistical confidence rises with the square of the ratio between noise and signal. That relationship is brutal.

Excess returnTracking errorYears for reasonable confidence
2 percent4 percentRoughly 16
2 percent6 percentRoughly 36
1 percent5 percentRoughly 100

These figures are approximate and the message does not depend on the precision. The horizons required exceed most careers, and they exceed the patience of almost every allocator.

By the time a track record is long enough to be statistically meaningful, the manager has often retired, the strategy has changed, or the market regime it worked in has ended.

Why Everyone Ignores It

Institutional practice evaluates managers over three to five year periods. That convention exists for governance and career reasons rather than statistical ones.

The consequence is predictable. Managers are hired after strong periods and fired after weak ones, and studies of institutional hiring and firing decisions have found that the fired managers frequently outperformed the newly hired ones over the following years.

That is what you would expect if the decisions were being driven by noise, since noise mean reverts.

The Multiple Manager Problem

The situation is worse than the single manager arithmetic implies, because allocators evaluate many managers.

With thousands of funds operating, some will produce long strong records by chance alone. Selecting the best performer from a large population and treating their record as evidence of skill ignores that the selection was made from the top of a distribution.

This is why an impressive record, considered in isolation, is weak evidence. The relevant question is how many managers were in the population from which this one was chosen.

What Helps When Data Is Short

Since waiting decades is not available, the practical approach is to gather evidence that is not return based.

Understand the process and whether it has an economic rationale. Examine the individual decisions rather than only the aggregate outcome, since a manager who was right for the stated reasons is different from one who was right accidentally. Look for consistency of approach across periods. Assess whether the same people are still making the decisions.

Increasing the number of independent observations also helps. A strategy making many decisions a year generates more evidence per year than one making a handful of large concentrated bets, which is why a high turnover systematic strategy can be assessed faster than a concentrated discretionary one.

The Same Problem in Personal Investing

An individual assessing their own ability faces exactly the same arithmetic with an even smaller sample. A few years of good returns during a rising market is not evidence of skill, and there is no way to make it into evidence.

This is one of the strongest arguments for a mechanical process. If the sample is too small to tell whether you have an edge, a rules based approach at least removes the risk of confidently acting on an edge you do not have.

The Bottom Line

Establishing statistical confidence in investment skill requires far more data than typical track records contain, because the noise in returns is several times the size of the signal. Standard three year evaluation windows measure luck. Since more data is not available, the substitute is evidence about process, decisions, and personnel, plus honest awareness of how many candidates the impressive record was selected from.

Explore Teen Biz News →