P Hacking Is Torturing Data Until It Confesses
Run enough tests on the same data and something will clear the significance threshold. The threshold was never designed to survive that treatment.
What the Threshold Means
A finding is conventionally considered statistically significant if there is less than a 5 percent chance of observing such an extreme result if there were no real effect
Read carefully that means that one in twenty tests performed with pure noise will produce a significant result. The threshold was designed for a world where a researcher formulated a hypothesis and ran a test
That's not the world anymore
It's worth reflecting on how arbitrary the number is. Five percent is a convention adopted for convenience in an age of manual computing and printed tables and is not intended to be the correct place to draw a line. Nothing changes in the world as a result of crossing from just above to just below. What changes is whether the finding is considered real written up and funded which is a big consequence of holding on to a boundary that no one chose for a reason
What a P Value Is Not
Almost every misuse of the threshold begins with a misunderstanding of the quantity itself so it pays to be precise about the direction of the statement
A p value is the probability of seeing data at least at this extreme assuming there is no real effect.It is not the probability that there is no real effect and it is not the probability that the finding is wrong. Those are different questions and converting between them requires knowing how likely the hypothesis was before the data arrived about which the p-value contains no information
The threshold is a statement about a single test. Applied after a search it is a statement about nothing
Three other readings are common misstatements. A p-value is not a measure of effect size so a highly significant result may describe an effect that is too small to be worth acting on and with a large enough sample almost any effect becomes significant. It is not a measure of the likelihood that the result will replicate. And a result just above the threshold is not proof that there is no effect. Not detecting something is not the same as establishing its absence even if it is constantly reported
The Arithmetic of Searching
Run 20 tests on random data and you'll expect about one significant result. Run 100 and expect about five. Run 10,000 which a computer does in seconds and expect about 500 seemingly significant findings on data that contains nothing at all
The reader of a result is never told this number. An article reports the test that worked a presentation reports the strategy that was launched and in both cases the denominator resides with the person who performed the analysis and is not part of what is published. The threshold can only be interpreted in conjunction with a count to which the audience structurally does not have access
Peeking Changes the Threshold
One entry in that table deserves to be separated because it corrupts the arithmetic even when only a single hypothesis is considered
Suppose the analysis is run the result does not reach the threshold and more data is collected. The analysis is run again. There is still more left so collect more run again and stop when everything becomes clear. Nothing here involves testing multiple hypotheses. Everywhere the same question is asked
However the false positive rate is well above 5 percent because each look is another chance for chance to produce the result and the stopping rule is defined by the result rather than fixed in advance. The estimate drifts as the data accumulates and a procedure that stops the moment the drift crosses a line will eventually cross it only with noise if given enough attention
This has a particular flavor in finance because the data comes in continuously. Each week a signal that is being monitored in production is examined and the natural time to declare it validated is the week in which the statistics look convincing. It is optional to stop with an accompanying market feed. The correction is to fix the sample size or evaluation date before starting or use a sequential test created for repeated looks which sets a higher bar for each for exactly this reason
The Forms It Takes
| Practice | Effect |
|---|---|
| Testing many variables | Some happen by chance |
| Testing many time periods | Find the one where it works |
| Adjusting the sample | Leave inconvenient comments |
| Changing the measurement | Test definitions until one is meaningful |
| Stop when it matters | Data collection ends at the right time |
| Formulate hypotheses after the result. | Present a discovery as a prediction. |
Most of these don't feel like bad behavior while they're happening. Trying a different definition of a variable because the first one seemed poorly constructed is a reasonable analytical decision. Doing it repeatedly until a meaning appears is a search and the meaning reported at the end doesn't count toward the search
The Choices Nobody Counted as Tests
The version of this that surprises careful people involves no repetition. A single analysis still contains dozens of decisions each of which could justifiably have taken another path
How are outliers handled? Are returns measured daily weekly or monthly? Does the sample start at the beginning of the data or at the point where the data becomes reliable? Are financial firms excluded or retained as is often the case? Is the variable used in levels or in changes. Is the result reported gross or net of any adjustments?
Each choice is made once in good faith with a reason available for it. But the reason is frequently provided after taking a look at how the analysis is going and the analyst who would have made a different decision if the first one had produced nothing has effectively run several tests while reporting one. The path through the decisions was guided by the results and nothing in the final report records that he could have gone elsewhere
That's why the honest question is never whether someone cheated. It's about how many analyzes were possible and to what extent the report was shaped by seeing the data
The Finance Version
Academic finance has a well-documented version of this problem. A large number of factors explaining returns have been published and independent replication has found that a substantial proportion do not survive out of the sample or into later periods
The incentive structure explains this. Journals publish positive results so tests that do not produce results are rarely written down. That means that the published literature is a filtered sample of everything tested selected based on importance which is a survivorship bias applied to research
Some researchers have proposed substantially increasing the required threshold for financial factors precisely because the effective number of tests performed across the field is enormous
Finance has an aggravating characteristic that laboratory sciences do not have. Everyone works with the same data. There is a price story and thousands of researchers have examined it which means that the effective number of tests performed against that story is much greater than the number reported by any single paper. A new finding drawn from the same series is arriving after an unknown but enormous number of previous attempts and its threshold must be set accordingly
Underpowered Tests Make It Worse
The other half of the problem receives much less attention than the threshold. Statistical power is the probability that a test detects an effect that actually exists and in finance it is usually low because the real effects are small and the usable history is short
Low power does not simply mean missing real findings. It changes the composition of the findings that are reported. Take a thousand hypotheses of which one hundred are actually true by pure arithmetic. At 50 percent power the true ones produce about 50 significant results. The nine hundred false ones yield about 45 with a 5 percent threshold. Thus about 95 significant findings appear and about 47 percent of them are false even though each individual test performsexactly as designed
If the power is reduced to 20 percent which is not unusual the true hypotheses will produce only 20 detections versus the same 45 false ones. Now about 69 percent of the significant results are incorrect. Nothing inappropriate happened in either scenario. The threshold did its job in each test
There is a second effect that compounds this. When power is low an effect only exceeds the threshold in samples in which chance worked in its favor so the surviving estimates systematically exaggerate the size of the effect. The finding is less likely to be real and when it is real it is smaller than reported
The Commercial Version
The same mechanism operates whenever a product is built from a search. A fund launched with a strategy discovered by testing many strategies has a backtest that exceeded a threshold designed for a single test
Index providers have faced the same criticism: an index constructed with rules selected because they performed well historically will show excellent historical performance by construction
The commercial environment also eliminates the weak corrective that academia has. A journal at least sees one article that other researchers can try to replicate. A company that tested forty strategies and launched one has no obligation to mention the other thirty-nine and marketing materials are not the place where that number will appear
The Corrections
Preregistration is the strongest. Stating the hypothesis testing and sampling before looking at the data eliminates the ability to search and is standard practice in clinical research for exactly that reason
Multiple test settings increase the required threshold based on how many tests were run. They require honest counting of tests including abandoned ones
Out-of-sample confirmation of data that was not touched during the search is the most practical verification available to an outsider
And an economic justification formed before testing is the cheapest filter of all. A relationship without any plausible mechanism discovered by a search is almost certainly noise regardless of its statistics
One more helps and costs nothing: report the size of the effect and its uncertainty rather than the verdict. A coefficient with a confidence interval tells the reader how large the effect might be and how strongly the data define it. The word significant compresses all of that into a binary at an arbitrary cutoff and the compression is where most of the damage occurs
What to Ask
When showing a significant finding useful questions are how many other things were tested whether the hypothesis existed before the analysis whether it holds up across different periods and markets and whether there is a reason it should be true
One finding that cannot answer these questions is a match with an accompanying p-value
Anything that is sold rather than published is worth adding two more. How large the effect is expressed in one unit is what matters since significance says nothing about magnitude. And what would have been reported if this test failed since one operation that only shows results that worked is to present a filtered sample no matter how rigorous each individual test was
The question of magnitude is what resolves most arguments in practice. A factor can be statistically significant over decades of data and offer an advantage measured in a few basis points per year that is real unambiguous and less than the cost of trading it. Significance answers whether the effect is distinguishable from zero. It says nothing about whether it can be distinguished from worthless and those two questions are combined every time a t-statistic is presented as if it were a reason to allocate capital
The Bottom Line
The P hack exploits the gap between a threshold designed for a test and a process that runs thousands. It's not fraud it's the accumulation of reasonable decisions in a search that no one counted. The defenses are pre-registration adjusting for number of tests off-sample confirmation and requiring an economic reason before consulting statistics. And the threshold itself deserves less deference than it receives since a p-value answers a limited question about a test in a field where everyone has already searched for the same price history