There's a structural assumption baked into how science is supposed to work: you publish a finding, peers review it, journals vet it, and then the broader literature cites it. The citation is meant to be an endorsement of a process, not just a pointer to a claim. That assumption is now breaking down in a measurable, documented way — and the implications for how scientific consensus forms are stranger than most coverage suggests.
A preprint posted to arXiv in late July puts numbers to something researchers have been quietly worried about for years. Using metadata from six major preprint servers linked to a large bibliographic reference database, the study's authors examined how often journal articles cite preprints — specifically, preprints that were never subsequently published in a peer-reviewed venue. The findings are striking: the preference to cite preprints has grown exponentially since the mid-2010s, with key inflection points at the launch of bioRxiv in 2013 and the introduction of a preprint content type in Crossref's DOI system in 2016. At its peak in 2021, nearly one in five journal articles cited at least one preprint.
That alone might be defensible — preprints are often faster, and many eventually get published. The more unsettling finding is the growing share of citations pointing to non-curated preprints: work that never made it through peer review. And the authors rule out the obvious confounds. It's not just that there are more preprints to cite. It's not explained by faster publishing cycles. It's not the outsized influence of a handful of viral preprints skewing the numbers. The shift reflects something structural in how researchers now treat preprints — as legitimate, autonomous citation objects, independent of their eventual publication status.
Early Visibility Shapes What Gets Remembered
The mechanism worth understanding here is not that researchers are being sloppy. Most of them probably assume the preprint they're citing will eventually be published, or was published somewhere they didn't track down. The problem is subtler: early visibility creates citation momentum that peer review can't easily reverse.
When a preprint lands on bioRxiv and gets shared widely, it enters the literature's informal memory before anyone has formally evaluated it. Other researchers read it, find it useful, cite it in their own work — which then gets published in journals. By the time a peer reviewer might flag problems with the original preprint (if it ever gets formally reviewed at all), the citation trail has already branched. The claim has propagated. Retracting or correcting the original does little to prune the tree.
This is citation bias operating at the infrastructure level. The question isn't whether individual researchers are acting in bad faith. The question is whether the system's architecture — preprints indexed, DOI-assigned, and citable before review — is creating a selection pressure that rewards early posting over eventual rigor. Scientific Reports publishes hundreds of open-access articles monthly across exactly these fields, and the sheer volume of that output makes the curation problem viscerally concrete: the pipeline is enormous, and the checkpoints are not keeping pace.
The LLM Problem Compounds This One
There's a second paper in the current source pool that, read alongside the preprint citation study, makes the picture considerably darker. A new arXiv preprint finds that by the end of 2025, 89% of open-access biomedical papers on PubMed Central show excess of LLM-associated vocabulary — a proxy measure for LLM-assisted writing. The authors are careful about what this does and doesn't show: LLM assistance isn't automatically misconduct, and it can genuinely help non-native English speakers. But the distribution is telling. LLMs are roughly twice as likely to be used when writing a Discussion section paragraph as a Methods section paragraph — 68% versus 32%, though even Methods sections show over 50% prevalence.
Discussion sections are where authors interpret their findings, situate them in the literature, and make claims about what their results mean. Methods sections are where they describe what they actually did. The gap suggests that LLM assistance is concentrated precisely where scientific judgment is supposed to live — and where overclaiming is easiest to obscure in fluent, confident prose.
Put these two findings together: preprints are increasingly cited before anyone vets them, and the papers doing the citing are increasingly written with tools that generate authoritative-sounding interpretation on demand. The conditions for consensus-distortion are not hypothetical. They're already in place.
What "Consensus" Actually Means Now
I've written before about how the preprint flood creates quality problems in specific fields. What the July citation study adds is a mechanism: it's not just that bad preprints exist, it's that citation behavior has structurally decoupled from curation. The study's authors call for renewed attention to citation conventions and journal policies — reasonable, if somewhat optimistic given how slowly those norms move.
The more immediate question is what this means for readers trying to evaluate scientific claims. When a consensus appears to be forming around a finding — when multiple papers cite the same result — that convergence used to be a signal that the finding had survived scrutiny at multiple checkpoints. Now it may simply mean a preprint got posted early, indexed quickly, and cited before anyone looked hard at it. The parallel to how AI benchmark scores get treated as ground truth before anyone examines what the benchmark actually measures is uncomfortably close: in both cases, a number circulates, acquires authority through repetition, and becomes load-bearing before the methodology underneath it has been seriously stress-tested.
Watch for whether major journals respond to the citation study with updated preprint citation policies. A few have experimented with requiring authors to note whether a cited preprint has been peer-reviewed. Whether that disclosure actually changes behavior is a different question — and one worth tracking.
