Statistical Significance Does Not Mean the Effect Is Large
A finding can be highly significant and completely trivial. The threshold answers a question about certainty and no question at all about magnitude.
The Two Separate Questions
There are two things you might want to know about an effect. Whether it exists, and whether it is big enough to care about.
Statistical significance addresses the first. Effect size addresses the second. They are independent, and conflating them is one of the most common analytical errors in finance.
Significance is a claim about confidence. Magnitude is a claim about importance. A finding can have plenty of one and none of the other.
Why Sample Size Breaks the Intuition
The test statistic grows with the square root of the number of observations. With a large enough sample, an arbitrarily small effect becomes statistically significant.
Given millions of observations, a strategy with an average edge of one hundredth of a percent per trade can be significant at extreme confidence levels. The effect is real and it does not survive a single transaction cost.
The reverse also happens. A genuinely large effect measured on a small sample may fail to reach significance, and reporting it as no effect confuses absence of evidence with evidence of absence.
The Four Combinations
| Significant | Large effect | Interpretation |
|---|---|---|
| Yes | Yes | Worth acting on |
| Yes | No | Real and useless |
| No | Yes | Promising, needs more data |
| No | No | Nothing here |
The second row is where most of the damage happens in finance, because it is the row that gets published and marketed. A statistically robust finding sounds authoritative regardless of whether the magnitude justifies any action.
What Significance Actually States
A p value is the probability of observing data at least this extreme if there were no real effect. That is a conditional statement running in a specific direction.
It is not the probability that the hypothesis is true. It is not the probability the result was a fluke. It is not one minus the probability of replication. Each of these misreadings is common enough to appear in professional commentary.
The distinction matters because the probability that a finding is true depends on how plausible it was beforehand. An implausible hypothesis with a significant p value is more likely to be a false positive than a plausible one with the same p value.
The Costs That Decide It
In finance the magnitude question has a specific form: is the effect larger than the cost of exploiting it.
Transaction costs, market impact, taxes, financing, and the fees charged to access a strategy all subtract from a gross edge. A statistically robust gross effect smaller than that stack is not an opportunity.
Capacity is the related constraint. An effect that exists reliably in very small illiquid securities may be untradeable at any meaningful size, which makes it real, significant, and commercially irrelevant.
Better Practice
Report confidence intervals rather than p values alone. An interval shows both whether zero is excluded and how large the effect plausibly is, which is the information actually needed.
State effect sizes in units that mean something: basis points after costs, dollars of profit, percentage of variance explained.
And decide in advance what magnitude would be large enough to change a decision. An effect below that threshold does not become interesting by being measured precisely.
The Bottom Line
Statistical significance tests whether an effect is distinguishable from zero and says nothing about whether it matters. Large samples make trivial effects significant, which is exactly the situation in finance, where datasets are enormous and real edges are small. Always ask for the magnitude after costs, and set the threshold for what would be worth acting on before looking at the result.