Hero image for "The Landmark Study That Didn't Hold Up — and the System That Made It Hard to Know"

The Landmark Study That Didn't Hold Up — and the System That Made It Hard to Know


A 2003 paper by Dan Ariely and Klaus Wertenbroch became one of behavioral economics' most-cited findings: that people perform better when they can precommit to evenly spaced deadlines rather than cramming at the last minute. It was the kind of result that felt intuitively right, generated hundreds of follow-on studies, and made its way into productivity advice, management consulting, and university syllabi. The only problem: a replication published this year in Psychological Science found that evenly spaced deadlines had a "negligible effect" on performance.

Not a smaller effect. Not a weaker effect. Negligible.

That's a significant word to find in a peer-reviewed replication of a landmark study. And it raises two questions worth sitting with: What does it mean when a finding that shaped a field doesn't hold up? And why is it so hard to find out when that happens?

What the Replication Actually Found

The new study, by Kyle Hyndman and Alberto Bisin, replicated Study 2 from the original Ariely-Wertenbroch paper using adult participants at a large public university in the United States. Their results showed that changes in deadline structure had negligible effects across three performance metrics and several survey measures. Externally imposed evenly spaced deadlines — the specific intervention the original paper championed — didn't stand out for reducing procrastination.

The replication also found patterns of participant behavior consistent with what you'd expect regardless of deadline structure, which is methodologically interesting: it suggests participants were responding to factors the original study may not have adequately controlled for. The original finding may have been capturing something real but much narrower than its conclusions implied — or something that doesn't generalize beyond the original sample.

This is exactly the kind of gap that the author's perspective here is built around: what researchers claim they found versus what they actually demonstrated. The original study demonstrated that a particular group of participants, under particular conditions, showed a particular pattern. The leap to "deadlines help people overcome procrastination" was always bigger than the data warranted. The replication just made that visible.

The Visibility Problem Is Structural

Here's what makes this more than a story about one failed replication: the Hyndman-Bisin paper exists, but most researchers citing the original Ariely-Wertenbroch study won't encounter it. Replication studies are systematically harder to find than original studies, and that asymmetry compounds over time.

A project described in Nature is trying to address this directly. A team of researchers has begun posting replication studies to the post-publication peer-review platform PubPeer, linking them to the original papers they attempted to reproduce. The source for those replication studies is the FORRT Library of Reproduction and Replication Attempts (FLoRA), which currently catalogs roughly 2,400 studies. Of those, nearly 1,000 have been successfully replicated independently, and 865 replication attempts failed. The remainder produced mixed results.

That's a meaningful failure rate — but the more important number is how often those failures are visible to the researchers who need to know about them. "Replications are generally less cited and less visible than original studies," says Josefine Weinerova, a psychologist at Birkbeck, University of London, who is co-leading the PubPeer project. The FORRT database launched in 2024, but its creators acknowledge it's still not easy to find replication studies because they're typically not linked to original papers in major databases.

The result is a citation ecosystem where a failed replication can sit in the literature for years while the original paper accumulates citations from researchers who have no way of knowing it's been challenged.

Why the System Produces This Outcome

A recent chapter on replicability and questionable research practices in the behavioral sciences frames the structural problem clearly: the replication crisis isn't just about individual researchers cutting corners. It's about a set of incentives — publication bias toward positive results, limited rewards for replication work, and editorial practices that favor novelty — that systematically underweight the evidence that would correct the record.

The proposed solutions are familiar at this point: preregistration, open data sharing, registered replication reports. What's less discussed is the discovery problem. Even when replication studies exist, the infrastructure for connecting them to original papers is fragile. PubPeer is a reasonable intervention, but it's a workaround for a problem that lives upstream — in how journals index, link, and surface post-publication evidence.

The Ariely-Wertenbroch deadline study will continue to be cited. The replication will continue to be hard to find. And somewhere, a management consultant is building a workshop around evenly spaced deadlines.

The fix isn't complicated to describe: replication studies need to be indexed alongside the papers they replicate, with the same discoverability. What's complicated is that this requires journals, databases, and funding bodies to treat replication as a first-class scientific activity rather than a footnote. The PubPeer project is a step. Whether it scales into something that actually changes citation behavior is the question worth watching over the next year or two.