Hero image for "Bem's Data Survived the Firing Squad. The Interpretation Is Still Up for Grabs."

Bem's Data Survived the Firing Squad. The Interpretation Is Still Up for Grabs.


The strangest thing happened to Daryl Bem's precognition research: it got more interesting after the critics tried to kill it.

When Bem published his 2011 paper in the Journal of Personality and Social Psychology — presenting what appeared to be experimental evidence for precognition across nine studies — the response was immediate and brutal. Skeptics called it an embarrassment to the field. Replication attempts, they said, failed. One journalist used the results as evidence that science itself was broken. The consensus among mainstream psychologists was that Bem had demonstrated, inadvertently, how badly flawed methodology could produce compelling-looking nonsense.

More than a decade later, that verdict looks considerably more complicated. Bem's data has, by several accounts, held up — not vindicated, exactly, but not dismissed either. What's happened in the intervening years is a methodological reckoning that implicates far more than one controversial psychologist.

The Replication Problem Cuts Both Ways

The original critique of Bem's work centered on what researchers call questionable research practices (QRPs): hypothesizing after results are known, running multiple analyses, selecting which data to report. A PeerJ analysis of study preregistration traces part of this story directly: critics like Wagenmakers et al. pointed out that Bem's methodology was largely indistinguishable from standard practice in mainstream psychology at the time. Which meant, uncomfortably, that if Bem's results were artifacts of QRPs, so were a significant portion of psychology's most celebrated findings.

This is the knife's edge the precognition debate has always balanced on. The methodological critique of Bem is correct and important. But it's a critique of a practice, not a proof that the underlying phenomenon doesn't exist. Cleaning up the methodology — through preregistration, larger samples, pre-specified analyses — is exactly what the field needed. The question is what happens to the anomalous signal when you apply those cleaner methods.

The answer, so far, is: it doesn't entirely disappear. That's not a triumphant vindication. It's a genuinely uncomfortable finding that deserves more rigorous attention than it typically gets.

What "Standing Up" Actually Means

It's worth being precise about what it means to say Bem's data "stood up," because the phrase can carry more weight than it should.

The Seeds of Science piece on this — written by Mitch Horowitz, a historian with an explicit sympathetic orientation toward anomalous research — frames it as vindication. That framing deserves scrutiny. What the meta-analytic record actually shows is a persistent statistical anomaly: effect sizes that remain above chance across aggregated studies, but that are small, contested, and subject to ongoing debate about file-drawer effects and methodological adequacy. "Standing up" means surviving aggregation. It does not mean the mechanism is understood, the effect is large, or the replication record is clean.

The distinction matters enormously for this field. A small, persistent statistical anomaly in a domain this theoretically fraught requires extraordinary methodological care before it can support extraordinary claims. The honest position is that the anomaly is real enough to warrant continued investigation under rigorous conditions — not that precognition has been demonstrated.

The Wiley-Blackwell Handbook of Transpersonal Psychology, published in August 2026, situates this within the broader parapsychology literature: psi research has identified what it calls "psi-conducive conditions" — states like dreaming, meditation, and relaxation — that appear to correlate with stronger anomalous effects. This is interesting as a research direction, but it also illustrates the field's persistent challenge: the more a study optimizes for conditions that produce effects, the harder it becomes to rule out demand characteristics, expectancy effects, and subtle methodological drift.

The Preregistration Test Is the One That Matters Now

The most important development in precognition research isn't any single positive finding — it's the slow adoption of preregistered designs. The PeerJ analysis of preregistration documents how preregistration changes the evidentiary picture: when researchers commit to their hypotheses, sample sizes, and analysis plans before collecting data, the QRP critique loses most of its force. What remains is the signal itself, stripped of the methodological ambiguity that has plagued this literature for decades.

The field doesn't yet have a large body of preregistered precognition studies with adequate statistical power. That's the honest state of play. What exists is a meta-analytic literature that shows something, a methodological reform movement that's beginning to apply proper constraints, and a gap between those two things that hasn't been closed.

That gap is the actual story — more interesting, I'd argue, than either "precognition is real" or "it's all noise." A persistent anomaly that survives aggregation but hasn't yet been subjected to a definitive preregistered test at scale is exactly the kind of open question that deserves serious attention rather than premature closure in either direction.

Watch for whether the next generation of preregistered replication attempts — particularly those with sample sizes large enough to detect small effects — produce results consistent with the meta-analytic record. If they do, the conversation changes. If they don't, the anomaly was probably methodological all along. Either outcome would be genuinely informative. The field has been waiting a long time for that clarity.


A closing poem this week from the archive of the irreducibly strange:

The future leaks, or seems to — a stain on the underside of now, faint, explicable, almost nothing. We measure almost nothing very carefully.