Here's a number worth sitting with: in a typical psychology experiment, the odds of getting a statistically significant result by chance — if you test just one hypothesis, once, cleanly — are 5%. That's what a p-value threshold of 0.05 means. But if you test twenty different outcomes, or slice your data by age and then by gender and then by age-and-gender combined, or try three different statistical models until one cooperates, your actual false-positive rate climbs toward certainty. You haven't cheated, exactly. You've just been flexible. And flexibility, in statistics, is a slow-motion disaster.
This is the problem of multiple comparisons, and it sits at the heart of why so much published research fails to replicate. I've written before about statistical significance as a threshold rather than a verdict — but the multiple comparisons problem is what happens upstream of that threshold, in the choices researchers make before they ever run the final test.
The Garden of Forking Paths
The statistician Andrew Gelman has a useful metaphor: the "garden of forking paths." At every stage of an analysis — which outliers to exclude, which covariates to include, how to operationalize the main variable — researchers face choices. Each choice is defensible. None of them, individually, looks like cheating. But the cumulative effect is that the analysis has quietly explored many possible versions of itself and reported only the one that worked.
This is distinct from outright fraud, and that distinction matters. A researcher who runs twenty tests and publishes only the one that reached p < 0.05 is doing something genuinely problematic. But a researcher who makes a series of reasonable-seeming analytical decisions, each of which nudges the result toward significance, may not even realize they're doing it. The incentive structure does the work. Journals publish significant results; careers depend on publications; so researchers — consciously or not — find their way to significant results.
The technical fix is straightforward to describe and genuinely difficult to implement: if you're going to test multiple hypotheses, you need to correct for that. The Bonferroni correction, the most conservative approach, simply divides your significance threshold by the number of tests. Testing twenty outcomes? Your threshold drops from 0.05 to 0.0025. Other corrections — Benjamini-Hochberg, for instance — are less aggressive but still force the researcher to account for the full scope of what they tested.
The problem is that these corrections only work if you know, and honestly report, how many tests you actually ran. And that's where the system breaks down.
What the Antibody Scandal Reveals About Invisible Flexibility
The multiple comparisons problem usually gets discussed in the context of academic papers, but a recent investigation suggests the same logic applies wherever data gets selectively presented. Ars Technica reported that a post-doc at Northwestern University, Reese Richardson, found manipulated images in the marketing materials of commercial antibody suppliers — the companies that sell research tools to labs worldwide. Working with collaborator Sholto David, Richardson partially automated the detection process and identified problematic image manipulations across roughly 17,500 antibodies from 16 different companies, with nearly 7 percent of examined images showing signs of manipulation.
The connection to p-hacking is structural, not incidental. In both cases, the problem is selective presentation: showing the version of the data that supports the claim you want to make, while the versions that don't support it disappear. In academic research, those disappearing versions are the failed analyses that never got reported. In commercial antibody marketing, they're the experiments that showed the antibody didn't work as advertised. The mechanism is different; the epistemic damage is the same. Downstream researchers build on results — published findings, or supplier-validated reagents — that were curated to look better than they are.
The Preregistration Partial Solution
The most widely adopted structural response to p-hacking is preregistration: publicly committing to your hypotheses, sample size, and analysis plan before you collect data. If you've declared in advance that you're testing exactly one outcome with exactly one statistical model, the garden of forking paths closes considerably.
But preregistration has limits that are worth being honest about. I covered the "adaptive preregistration" problem earlier this year — the practice of amending preregistered plans in ways that functionally restore analytical flexibility while maintaining the appearance of rigor. And preregistration does nothing for the existing literature, which was built without it.
The Nature Neuroscience paper on OCD and tic disorder genetics published recently offers a useful contrast. The researchers explicitly applied false discovery rate corrections throughout their analysis — reporting FDR thresholds, adjusting for multiple comparisons across thousands of genetic variants, and distinguishing between high-confidence findings and more speculative ones. That's what the methodology section of a paper is supposed to do: show you the correction math, not hide it.
Most papers don't make it that easy to check. Which means the reader's job — and the science journalist's job — is to ask the question that the methods section often buries: how many things did they actually test before they found the one they reported?
That number is almost never in the abstract.
