Hero image for "The Incentive System Was Already Broken. LLMs Are Just Accelerating It."

The Incentive System Was Already Broken. LLMs Are Just Accelerating It.


When I wrote about underpowered studies last June, the argument was structural: small samples dominate published research not because scientists are careless, but because the system rewards publication over rigor. A new preprint posted to arXiv on July 19 makes that argument more uncomfortable — because it suggests the problem is about to get significantly worse.

More Papers, Not Better Ones

The preprint, posted by a team including Carl Bergstrom, a biologist at the University of Washington, uses optimal-foraging theory to model how scientists will reallocate their effort once LLMs become standard research tools. The framing is clever: optimal-foraging theory was developed to analyze how animals maximize energy gain while conserving resources. Apply it to scientists, and you get a model of how researchers distribute effort across phases of a project — initial hypothesis generation, required work like drafting and figure-making, and discretionary development like follow-up experiments and polishing.

The model's prediction is blunt. LLMs speed up every phase of the research process. But that acceleration doesn't translate into better papers. It translates into more papers. Faster discovery and drafting phases mean researchers can hit their publication quotas more quickly — and then move on to the next project rather than spending extra time on the one they just finished. "LLMs are rarely the problem themselves," Bergstrom told Nature. "LLMs hold up a mirror to problems that we already have."

That phrase deserves to sit with you for a moment. The problem isn't the tool. The problem is what the tool reveals about the incentive structure it's being dropped into.

What "Discretionary Development" Actually Means

The preprint's breakdown of research phases is worth unpacking, because "discretionary development" is doing a lot of work in that model. This is the category that includes follow-up experiments, additional controls, and the kind of methodological tightening that separates a suggestive finding from a robust one. It's also, by definition, the category that gets cut first when researchers are under pressure to publish.

This is precisely where statistical power lives. Running a study with adequate sample size to reliably detect an effect of the expected magnitude isn't required work in the model's framework — it's discretionary. You can publish without it. Journals will accept the paper. The press release will go out. The finding will enter the literature. The fact that the study was underpowered to detect anything reliably is a detail that rarely makes the abstract.

The LLM acceleration model predicts that this discretionary category will shrink further, not grow. Scientists will use the time saved to start new projects, not to strengthen existing ones. The incentive to publish more papers outweighs the incentive to publish better ones — and LLMs, by making the minimum viable paper faster to produce, tighten that logic rather than loosening it.

One Data Point That Cuts the Other Way

There's a counterweight worth noting. A separate preprint, examined in a Nature Index report, analyzed 116,359 papers published in PLOS journals since May 2019 and found that papers with publicly available peer-review reports were retracted at a meaningfully lower rate than those without. Among papers with open referee reports, 105 were later retracted; among those without, 321 were retracted — out of a pool where open-review papers represented roughly 40% of the total.

The study doesn't claim to explain the mechanism, and its authors are careful not to. But the lead hypothesis is that researchers who opt into open peer review tend to be the same researchers who embrace open-science practices more broadly — preregistration, data sharing, methodological transparency. The decision to make your review process public may function as a proxy for the kind of researcher who was already doing the discretionary development work.

If that interpretation holds, it suggests the problem isn't uniformly distributed. There's a subset of researchers for whom rigor is intrinsic rather than incentivized. The question is whether that subset grows or shrinks as LLM-assisted speed becomes the norm.

The Mirror Problem

Bergstrom's "mirror" framing is the most useful thing either of these studies offers. The replication crisis, the underpowered-study problem, the publication bias toward positive results — none of these are new diagnoses. What's new is that a technology is arriving that will stress-test the system's existing pathologies at scale.

The researchers whose incentives were already misaligned will publish faster and cut more corners. The researchers who were already doing careful work will probably continue doing careful work, possibly more efficiently. LLMs don't change the underlying incentive structure; they amplify whatever was already there.

That's not a reason for fatalism. It's a reason to focus reform efforts on the incentive structure itself — on what journals reward, what hiring committees count, and what "productivity" means in a field where a single well-powered, carefully controlled study is worth more than a dozen underpowered ones that will never replicate. The mirror is showing us something. The question is whether anyone with institutional power is looking at it.