Here's a scenario that should unsettle you. A research team runs a study on 40 participants, finds a striking effect, publishes in a respectable journal, and watches the press release get picked up by a dozen outlets. Six months later, a replication attempt with 400 participants finds nothing. The original authors defend their work. The replication team defends theirs. Readers are left wondering who to believe.
The answer, often, is neither — because the original study was never equipped to answer the question it claimed to answer. This is the statistical power problem, and it sits at the root of a lot of what's broken in published science.
What "Power" Actually Means — and Why Journals Ignore It
Statistical power is the probability that a study will detect a real effect if one exists. A study with 80% power has an 80% chance of finding a true signal; it will miss that signal 20% of the time. The conventional target, established by statistician Jacob Cohen decades ago, is 0.80. Most published studies don't come close.
The reason is structural, not accidental. Running more participants costs money and time. Journals reward novelty and positive results, not methodological rigor. Researchers working under grant pressure and publication timelines make the rational choice: run a smaller study, hope the effect is large enough to clear the significance threshold, and publish. If the effect is real but modest, an underpowered study will miss it. If the effect is noise, an underpowered study will occasionally — by chance — find something that looks significant anyway.
That second scenario is the one that should keep you up at night. When a study is underpowered and still finds a "significant" result, that result is more likely to be a false positive than the same result from a well-powered study. This is sometimes called the winner's curse: the studies that make it through the publication filter in small-sample research tend to be the ones that got lucky, not the ones that found something real.
I've written before about how p-values mislead — how a threshold of p < 0.05 tells you almost nothing about whether a drug works or an effect is real. The power problem is the upstream version of that same failure. A p-value from an underpowered study is a p-value from a study that was structurally prone to producing misleading results in the first place.
The Incentive Architecture That Keeps This Going
The persistence of underpowered studies isn't a mystery. It's a predictable output of how science gets funded, evaluated, and published.
Consider what happens when a researcher runs a properly powered study. If the effect is real, they find it — good outcome. If the effect doesn't exist, they find nothing — and a null result is genuinely hard to publish. The asymmetry is brutal: a well-powered null result represents months of work and often ends up in a file drawer. An underpowered positive result, meanwhile, clears peer review and generates press coverage.
Nature and PNAS, the journals that set the tone for what "important science" looks like, explicitly prioritize significance and novelty over methodological conservatism. Nature's acceptance rate runs around 6%; PNAS sits closer to 10–15%. The desk rejection filter at these journals is tuned for impact, not for sample size adequacy. A technically flawless study confirming a known effect with 2,000 participants is a harder sell than a surprising finding from 35.
That 35-participant number isn't hypothetical. A recent Nature Neuroscience paper on noninvasive brain-computer interfaces — a genuinely interesting piece of work — demonstrated its core decoding method in exactly 35 healthy volunteers. The authors are careful about this; they report character error rates with appropriate nuance, and they frame the work as a proof of concept rather than a clinical solution. But 35 participants is 35 participants. The best-case performance figures (18% character error rate for top performers) come from a subset of an already small cohort. That's not a criticism of the science — it's a description of where the field is. Early-stage research runs small. The problem is when small-sample early-stage findings get treated as established results.
What Honest Reporting Would Look Like
A study's power calculation — if it has one — should be in the first paragraph of any science news coverage, not buried in the methods section or omitted entirely. Readers deserve to know: was this study designed to reliably detect the effect it claims to have found?
The honest version of most science headlines would read something like: "Researchers found a promising signal in a small sample; a larger study is needed to know if it's real." That's less exciting than "Scientists discover X." It's also more accurate.
The replication crisis didn't emerge from bad intentions. It emerged from a system where the incentives for running small, optimistic studies consistently outweigh the incentives for running large, rigorous ones. Fixing that requires changes to how journals evaluate submissions, how funders reward null results, and how science journalists read methods sections — not just abstracts.
Until then, the gap between what a study claims and what it demonstrates will keep generating headlines that don't survive contact with replication. The studies that don't make news — the ones that tried to confirm a finding and couldn't — are often doing more honest work than the ones that did.
