P Hacking Is Torturing Data Until It Confesses
Run enough tests on the same data and something will clear the significance threshold. The threshold was never designed to survive that treatment.
What the Threshold Means
A finding is conventionally called statistically significant if there is less than a 5 percent probability of observing a result that extreme if no real effect existed.
Read carefully, that means one in twenty tests conducted on pure noise will produce a significant result. The threshold was designed for a world where a researcher formed one hypothesis and ran one test.
That is not the world any longer.
The Arithmetic of Searching
Run 20 tests on random data and you expect roughly one significant result. Run 100 and expect about five. Run 10,000, which a computer does in seconds, and expect around 500 apparently significant findings in data containing nothing at all.
The threshold is a statement about a single test. Applied after a search, it is a statement about nothing.
The Forms It Takes
| Practice | Effect |
|---|---|
| Testing many variables | Some pass by chance |
| Trying many time periods | Find the one where it works |
| Adjusting the sample | Drop inconvenient observations |
| Changing the measure | Try definitions until one is significant |
| Stopping when significant | Data collection ends at the right moment |
| Hypothesising after the result | Present a discovery as a prediction |
Most of these do not feel like misconduct while they are happening. Trying a different definition of a variable because the first one seemed poorly constructed is a reasonable analytical decision. Doing it repeatedly until significance appears is a search, and the significance reported at the end does not account for the search.
The Finance Version
Academic finance has a well documented version of this problem. A very large number of factors have been published as explaining returns, and independent replication has found that a substantial proportion do not survive out of sample or in later periods.
The incentive structure explains it. Journals publish positive findings, so tests producing nothing are rarely written up. That means the published literature is a filtered sample of everything tested, selected on significance, which is survivorship bias applied to research.
Some researchers have proposed raising the required threshold substantially for finance factors, precisely because the effective number of tests conducted across the field is enormous.
The Commercial Version
The same mechanism operates whenever a product is built from a search. A fund launched on a strategy discovered by testing many strategies has a backtest that cleared a threshold designed for a single test.
Index providers have faced the same criticism: an index constructed with rules selected because they performed well historically will show excellent historical performance by construction.
The Corrections
Pre registration is the strongest. Stating the hypothesis, the test, and the sample before looking at the data removes the ability to search, and it is standard practice in clinical research for exactly that reason.
Multiple testing adjustments raise the required threshold according to how many tests were run. They require honestly counting the tests, including the ones abandoned.
Out of sample confirmation on data not touched during the search is the most practical check available to an outsider.
And an economic rationale formed before the test is the cheapest filter of all. A relationship with no plausible mechanism, discovered by searching, is almost certainly noise regardless of its statistics.
What to Ask
When shown a significant finding, the useful questions are how many other things were tested, whether the hypothesis existed before the analysis, whether it holds in different periods and markets, and whether there is a reason it should be true.
A finding that cannot answer these is a coincidence with a p value attached.
The Bottom Line
P hacking exploits the gap between a threshold designed for one test and a process that runs thousands. It is usually not fraud, it is the accumulation of reasonable choices in a search that nobody counted. The defences are pre registration, adjusting for the number of tests, out of sample confirmation, and requiring an economic reason before the statistics are consulted.