Hero image for "Flexible Analysis Isn't a Bug in Neuroscience Research. It's Often the Feature."

Flexible Analysis Isn't a Bug in Neuroscience Research. It's Often the Feature.


A researcher runs a brain imaging study on 28 participants. The primary analysis doesn't reach significance. So they try a different region of interest. Still nothing. They adjust the motion-correction threshold. Getting warmer. They exclude two participants as outliers. There it is: p = 0.043. The paper reports a significant finding. The press release announces a breakthrough.

This is p-hacking — or, in the more charitable framing researchers sometimes prefer, "researcher degrees of freedom." And a meta-epidemiological study published in Trials last month offers a useful window into why neuroscience and psychiatry are particularly exposed to it. Analyzing 18,609 neurology and psychiatry drug trials registered on ClinicalTrials.gov between 2000 and 2023, the authors found that most were early phase, single-center, and small — typically enrolling between 11 and 50 participants — and that 38% were open label, meaning neither participants nor researchers were blinded to treatment assignment.

Small samples. No blinding. These are the conditions under which flexible analysis does the most damage.

Why Small Samples Amplify the Problem

Statistical significance is not a fixed property of a true effect. It's a threshold that a noisy estimate has to clear. In a small sample, estimates are noisy by definition — the confidence intervals are wide, and the difference between p = 0.06 and p = 0.04 can hinge on a single data point, a single preprocessing choice, a single analytical decision made after looking at the data.

This is the core mechanism behind p-hacking: when you have many legitimate-seeming analytical choices available — which covariates to include, how to handle missing data, which time window to analyze, which participants to exclude — and you make those choices after seeing the results, you are effectively running multiple tests while reporting only one. The nominal false positive rate of 5% no longer applies. You've inflated it, often dramatically, without any record of having done so.

The mathematics here are unforgiving. A recent paper from the University of Chicago and Cambridge examining false discovery rate control shows that even under relatively favorable conditions — what statisticians call "compound p-values" that are valid only on average across true nulls — the Benjamini-Hochberg procedure can produce a false discovery rate of up to 1.93 times the nominal level. Under positive dependence between tests, the inflation can scale as O(log m), where m is the number of hypotheses. The point isn't that this paper is directly about p-hacking; it's that it illustrates how fragile false discovery rate guarantees become when the validity conditions on p-values are weakened even slightly. Researcher degrees of freedom weaken those conditions in exactly this way — silently, and without any formal record.

The Neuroscience Exposure

The Trials study wasn't designed to measure p-hacking directly. What it measured was the structural conditions that make p-hacking easy: small samples, open-label designs, and — critically — poor results reporting. Among trials completed after 2007, only 50% had reported results, with a mean delay of 12 months. That's a lot of dark matter. Studies that don't report results don't just disappear; they shape the literature by their absence, leaving only the positive findings visible.

The most studied conditions in this dataset — pain, schizophrenia, depression, Alzheimer's disease — are also among the fields with the most notorious replication problems. That's not a coincidence. These are areas where the underlying biology is complex, effect sizes are often modest, and the pressure to find something publishable is intense. When your sample is 30 people and your analysis pipeline has a dozen decision points, the probability that at least one combination clears p < 0.05 by chance is not small.

Preregistration was supposed to fix this. If you commit to your analysis plan before collecting data, you can't selectively report the version that worked. But as I wrote in June, "adaptive preregistration" has emerged as a workaround — a way to modify the registered plan mid-study under the cover of legitimate-sounding flexibility. The structural problems the Trials study documents — small samples, open labels, missing results — suggest that even where preregistration exists, the underlying research conditions may not have changed enough to matter.

What Would Actually Help

The Trials authors note that industry funding of neurology and psychiatry trials has been declining, with university sponsorship rising to fill the gap. That sounds like good news until you consider that academic incentives — publish or perish, novelty over replication, grants tied to positive results — are not obviously better than industry incentives at producing reliable science. Different pressures, similar distortions.

The structural fix is unglamorous: larger samples, blinded designs, mandatory results reporting, and pre-specified analysis plans that don't have escape hatches. The Trials dataset found that 10% of trials weren't even randomized. That's not a p-hacking problem; that's a design problem that precedes any statistical analysis.

Watch for whether the results-reporting gap documented here — half of completed trials going dark — becomes a target for registry enforcement. ClinicalTrials.gov has had mandatory reporting requirements for years. The compliance rate suggests those requirements need teeth, not just paperwork.