Psychology's replication crisis has been well-documented in these pages — the landmark Ariely findings that didn't hold up, the publication bias that buries null results, the structural incentives that reward novelty over rigor. But a new preprint adds a genuinely unsettling layer to that story, and it has nothing to do with p-hacking or underpowered samples.
The question worth sitting with right now: if we're struggling to replicate findings from studies written by humans exercising their own judgment, what happens when the original studies themselves were substantially written by AI?
The Numbers Are Harder to Dismiss Than They Look
A preprint posted to arXiv on August 12 — and covered by Nature — estimates that nearly 90% of biomedical papers published in December 2025 showed signs of AI-assisted writing. The figure for all of 2025 sits at roughly 77%, up from an estimated 52% for 2024. Dmitry Kobak, a computer scientist at Ghent University and co-author of the study, told Nature he was initially convinced his team had made an error: "I was sure that we did something wrong." Further checks convinced him the data held.
The methodology matters here, as it always does. Previous estimates of LLM use in scientific literature relied on analyzing paper abstracts and found figures around 13.5% for 2024. This new study analyzed full text and used a more sensitive detection method — one that yields direct estimates rather than lower bounds. That methodological shift accounts for most of the gap between 13.5% and the new figures. Whether that sensitivity is a feature or a bug depends on how you read it: the method may be catching genuine AI use that abstract-only analysis missed, or it may be overcounting. The authors acknowledge more analysis is needed.
What makes the numbers harder to wave away is the corroboration from survey data. A 2025 survey found that 71% of researchers reported using AI for writing assistance — and the preprint's authors note that self-reported figures typically undercount actual behavior. The gap between what researchers admit and what they do is probably not closing in the direction of less AI use.
What This Means for Replication, Specifically
Here's where the replication crisis angle sharpens. The classic failure mode in psychology research involves a researcher making dozens of small decisions — how to code ambiguous responses, which outliers to exclude, when to stop collecting data — and those decisions, consciously or not, nudging results toward significance. The problem isn't fraud; it's that human judgment is porous and motivated.
AI writing assistance introduces a different kind of porosity. The concern isn't that LLMs are fabricating data. It's subtler: AI tools are trained on existing literature, which means they've absorbed the genre conventions of scientific writing — including the conventions of how findings get framed, how limitations get minimized, and how conclusions get stated with more confidence than the data warrants. A researcher who uses an LLM to draft their discussion section may end up with prose that sounds more authoritative than their results justify, not because they intended to mislead, but because that's what the model learned to produce.
The Nature coverage notes that paper introductions and discussion sections show more signs of AI use than results sections. That's the exact distribution you'd expect if researchers are using AI to do the interpretive work — the parts of a paper where the gap between "what we found" and "what we claim" is most likely to widen.
A New Variable in an Already Complicated Equation
There's a preprint on arXiv from this month examining replicability of statistically significant findings more broadly — though its full content wasn't available for review. The directional question it raises is one the field has been circling for years: how much of what gets published as significant actually holds up?
That question now has an additional complication. Replication studies assume the original paper reflects the researchers' actual reasoning and judgment. If a substantial portion of the interpretive scaffolding was generated by a model optimized to produce confident-sounding scientific prose, then replication isn't just checking whether the effect is real — it's checking whether the original paper accurately represented what the researchers even thought they found.
This is genuinely new territory. The replication crisis was, at its core, a story about human psychology: motivated reasoning, career incentives, the seductive pull of a clean result. AI assistance doesn't eliminate those pressures. It adds a layer of automated amplification on top of them.
The field's methodological reformers — preregistration advocates, open-data proponents, replication consortia — built their tools to catch human bias. Whether those tools are calibrated for this new variable is a question worth watching. The more immediate one: as journals begin grappling with AI disclosure policies, the papers most worth scrutinizing may be the ones where the discussion section reads most fluently.
