Hedge Fund

P Hacking Is Torturing Data Until It Confesses

Run enough tests on the same data and something will clear the significance threshold. The threshold was never designed to survive that treatment.

↩ Looking BackPart of the 2020 to 2026 retrospective, written in July 2026. The date below marks the 2024 events this piece revisits, not when it was published, so it draws on everything known through mid 2026.
Nathan Xiang·April 12, 2024

What the Threshold Means

A finding is conventionally considered statistically significant if there is less than a 5 percent chance of observing such an extreme result if there were no real effect

Read carefully that means that one in twenty tests performed with pure noise will produce a significant result. The threshold was designed for a world where a researcher formulated a hypothesis and ran a test

That's not the world anymore

It's worth reflecting on how arbitrary the number is. Five percent is a convention adopted for convenience in an age of manual computing and printed tables and is not intended to be the correct place to draw a line. Nothing changes in the world as a result of crossing from just above to just below. What changes is whether the finding is considered real written up and funded which is a big consequence of holding on to a boundary that no one chose for a reason

What a P Value Is Not

Almost every misuse of the threshold begins with a misunderstanding of the quantity itself so it pays to be precise about the direction of the statement

A p value is the probability of seeing data at least at this extreme assuming there is no real effect.It is not the probability that there is no real effect and it is not the probability that the finding is wrong. Those are different questions and converting between them requires knowing how likely the hypothesis was before the data arrived about which the p-value contains no information

The threshold is a statement about a single test. Applied after a search it is a statement about nothing

Three other readings are common misstatements. A p-value is not a measure of effect size so a highly significant result may describe an effect that is too small to be worth acting on and with a large enough sample almost any effect becomes significant. It is not a measure of the likelihood that the result will replicate. And a result just above the threshold is not proof that there is no effect. Not detecting something is not the same as establishing its absence even if it is constantly reported

The Arithmetic of Searching

Run 20 tests on random data and you'll expect about one significant result. Run 100 and expect about five. Run 10,000 which a computer does in seconds and expect about 500 seemingly significant findings on data that contains nothing at all

The reader of a result is never told this number. An article reports the test that worked a presentation reports the strategy that was launched and in both cases the denominator resides with the person who performed the analysis and is not part of what is published. The threshold can only be interpreted in conjunction with a count to which the audience structurally does not have access

Peeking Changes the Threshold

One entry in that table deserves to be separated because it corrupts the arithmetic even when only a single hypothesis is considered

Suppose the analysis is run the result does not reach the threshold and more data is collected. The analysis is run again. There is still more left so collect more run again and stop when everything becomes clear. Nothing here involves testing multiple hypotheses. Everywhere the same question is asked

However the false positive rate is well above 5 percent because each look is another chance for chance to produce the result and the stopping rule is defined by the result rather than fixed in advance. The estimate drifts as the data accumulates and a procedure that stops the moment the drift crosses a line will eventually cross it only with noise if given enough attention

This has a particular flavor in finance because the data comes in continuously. Each week a signal that is being monitored in production is examined and the natural time to declare it validated is the week in which the statistics look convincing. It is optional to stop with an accompanying market feed. The correction is to fix the sample size or evaluation date before starting or use a sequential test created for repeated looks which sets a higher bar for each for exactly this reason

The Forms It Takes

PracticeEffect
Testing many variablesSome happen by chance
Testing many time periodsFind the one where it works
Adjusting the sampleLeave inconvenient comments
Changing the measurementTest definitions until one is meaningful
Stop when it mattersData collection ends at the right time
Formulate hypotheses after the result.Present a discovery as a prediction.

Most of these don't feel like bad behavior while they're happening. Trying a different definition of a variable because the first one seemed poorly constructed is a reasonable analytical decision. Doing it repeatedly until a meaning appears is a search and the meaning reported at the end doesn't count toward the search

The Choices Nobody Counted as Tests

The version of this that surprises careful people involves no repetition. A single analysis still contains dozens of decisions each of which could justifiably have taken another path

How are outliers handled? Are returns measured daily weekly or monthly? Does the sample start at the beginning of the data or at the point where the data becomes reliable? Are financial firms excluded or retained as is often the case? Is the variable used in levels or in changes. Is the result reported gross or net of any adjustments?

Each choice is made once in good faith with a reason available for it. But the reason is frequently provided after taking a look at how the analysis is going and the analyst who would have made a different decision if the first one had produced nothing has effectively run several tests while reporting one. The path through the decisions was guided by the results and nothing in the final report records that he could have gone elsewhere

That's why the honest question is never whether someone cheated. It's about how many analyzes were possible and to what extent the report was shaped by seeing the data

The Finance Version

Academic finance has a well-documented version of this problem. A large number of factors explaining returns have been published and independent replication has found that a substantial proportion do not survive out of the sample or into later periods

The incentive structure explains this. Journals publish positive results so tests that do not produce results are rarely written down. That means that the published literature is a filtered sample of everything tested selected based on importance which is a survivorship bias applied to research

Some researchers have proposed substantially increasing the required threshold for financial factors precisely because the effective number of tests performed across the field is enormous

Finance has an aggravating characteristic that laboratory sciences do not have. Everyone works with the same data. There is a price story and thousands of researchers have examined it which means that the effective number of tests performed against that story is much greater than the number reported by any single paper. A new finding drawn from the same series is arriving after an unknown but enormous number of previous attempts and its threshold must be set accordingly

Underpowered Tests Make It Worse

The other half of the problem receives much less attention than the threshold. Statistical power is the probability that a test detects an effect that actually exists and in finance it is usually low because the real effects are small and the usable history is short

Low power does not simply mean missing real findings. It changes the composition of the findings that are reported. Take a thousand hypotheses of which one hundred are actually true by pure arithmetic. At 50 percent power the true ones produce about 50 significant results. The nine hundred false ones yield about 45 with a 5 percent threshold. Thus about 95 significant findings appear and about 47 percent of them are false even though each individual test performsexactly as designed

If the power is reduced to 20 percent which is not unusual the true hypotheses will produce only 20 detections versus the same 45 false ones. Now about 69 percent of the significant results are incorrect. Nothing inappropriate happened in either scenario. The threshold did its job in each test

There is a second effect that compounds this. When power is low an effect only exceeds the threshold in samples in which chance worked in its favor so the surviving estimates systematically exaggerate the size of the effect. The finding is less likely to be real and when it is real it is smaller than reported

The Commercial Version

The same mechanism operates whenever a product is built from a search. A fund launched with a strategy discovered by testing many strategies has a backtest that exceeded a threshold designed for a single test

Index providers have faced the same criticism: an index constructed with rules selected because they performed well historically will show excellent historical performance by construction

The commercial environment also eliminates the weak corrective that academia has. A journal at least sees one article that other researchers can try to replicate. A company that tested forty strategies and launched one has no obligation to mention the other thirty-nine and marketing materials are not the place where that number will appear

The Corrections

Preregistration is the strongest. Stating the hypothesis testing and sampling before looking at the data eliminates the ability to search and is standard practice in clinical research for exactly that reason

Multiple test settings increase the required threshold based on how many tests were run. They require honest counting of tests including abandoned ones

Out-of-sample confirmation of data that was not touched during the search is the most practical verification available to an outsider

And an economic justification formed before testing is the cheapest filter of all. A relationship without any plausible mechanism discovered by a search is almost certainly noise regardless of its statistics

One more helps and costs nothing: report the size of the effect and its uncertainty rather than the verdict. A coefficient with a confidence interval tells the reader how large the effect might be and how strongly the data define it. The word significant compresses all of that into a binary at an arbitrary cutoff and the compression is where most of the damage occurs

What to Ask

When showing a significant finding useful questions are how many other things were tested whether the hypothesis existed before the analysis whether it holds up across different periods and markets and whether there is a reason it should be true

One finding that cannot answer these questions is a match with an accompanying p-value

Anything that is sold rather than published is worth adding two more. How large the effect is expressed in one unit is what matters since significance says nothing about magnitude. And what would have been reported if this test failed since one operation that only shows results that worked is to present a filtered sample no matter how rigorous each individual test was

The question of magnitude is what resolves most arguments in practice. A factor can be statistically significant over decades of data and offer an advantage measured in a few basis points per year that is real unambiguous and less than the cost of trading it. Significance answers whether the effect is distinguishable from zero. It says nothing about whether it can be distinguished from worthless and those two questions are combined every time a t-statistic is presented as if it were a reason to allocate capital

The Bottom Line

The P hack exploits the gap between a threshold designed for a test and a process that runs thousands. It's not fraud it's the accumulation of reasonable decisions in a search that no one counted. The defenses are pre-registration adjusting for number of tests off-sample confirmation and requiring an economic reason before consulting statistics. And the threshold itself deserves less deference than it receives since a p-value answers a limited question about a test in a field where everyone has already searched for the same price history

Explore Teen Biz News →