Hero image for "The Paper Mill Problem Isn't Fake Papers. It's That Nobody's Checking Who Cites Them."

The Paper Mill Problem Isn't Fake Papers. It's That Nobody's Checking Who Cites Them.


Here's an uncomfortable number: 480 patents, 57 policy documents, and 12 clinical guidelines have cited papers from an authorship-for-sale scheme that isn't even a real research network (Retraction Watch). The scheme is called the Pharmakon Neuroscience Network, and it doesn't exist as an actual institution — it's a brand that authors could apparently pay to be affiliated with, giving fabricated or low-quality research the veneer of a coherent research group (information for practice).

That detail matters more than the raw citation counts. Forensic statisticians and research-integrity sleuths have gotten reasonably good at flagging suspicious papers themselves. What they're much worse at tracking — because almost nobody is tracking it — is where those flagged papers go afterward. This week's story isn't really about detection. It's about what happens after detection fails to stop the paper from entering the bloodstream of real-world decisions.

What Actually Counts as a Forensic Red Flag

The analysis behind the PNN numbers, posted to arXiv on August 31 by researchers at Digital Science, used a specific tell that sleuths rely on constantly: author lists made up of researchers with no history of collaboration, no shared coauthors, and sometimes no overlap in field (Retraction Watch). That's a structural signature, not a content judgment — you don't need to read the paper closely to suspect something is wrong, you just need to map the social network of who's publishing with whom. It's the citation-graph equivalent of noticing that a "team" of coworkers has never actually met.

This lineage of pattern-based detection goes back further than most readers probably realize. Elisabeth Bik's landmark 2016 study didn't use algorithms to catch fraud — she and her coauthors manually scanned more than 20,000 papers across 40 journals and found 782 instances of image duplication plus 196 papers with altered duplicate figures (Physics World). That work turned image forensics into a recognized subfield of scientific sleuthing and helped spawn the loose community of "sleuths" that now flags a meaningful share of the corrections and retractions we see today, work for which Bik later received the Einstein Foundation Award (Physics World). The throughline from Bik's manual pixel-hunting to today's coauthorship-network analysis is the same: fabrication tends to leave structural fingerprints, even when the underlying data looks clean. Newer automated approaches are chasing a similar intuition for images generated by AI rather than doctored by hand — one recent self-supervised detector trains only on real images and still learns to flag generator "fingerprints" it's never seen before, reaching high accuracy on generators it wasn't trained against (arXiv). The underlying bet is the same across both eras of sleuthing: fabrication and manipulation leave statistical residue, whether it's an implausible coauthor list or a pixel-level artifact.

Detection Doesn't Equal Removal

Here's where the optimism runs out. Knowing a paper mill exists and proving its output shaped a guideline are two very different claims, and the researchers behind the PNN analysis were careful not to conflate them. Leslie McIntosh told Retraction Watch that "clinical guidelines also have a lot of citations generally," and that the analysis has "no silver bullet" showing paper mill studies meaningfully distorted the recommendations (Retraction Watch). Malcolm Macleod, the Edinburgh research-integrity expert quoted in the same piece, made the more useful point: you'd need to ask whether the guideline would look any different if the paper mill study had never existed, since many mill-produced papers just parrot findings already established elsewhere in the literature (information for practice).

That's a genuinely hard counterfactual to run, and it's why most of the citations sitting in patents and guidelines will probably never get individually audited. The bulk of clinical-guideline citations tied to problematic work in this analysis trace back to Shaker Mousa, a researcher found to have faked data in two papers by a 2024 Office of Research Integrity investigation and who now carries at least 16 retractions (Retraction Watch). Sixteen retractions is a lot of chances for someone downstream to notice a pattern before it reaches a guideline document. Nobody did, at least not in time — a gap that fits a broader pattern this same outlet has been tracking in its weekend roundups of research-integrity failures across fields (Retraction Watch).

The Sleuthing Keeps Improving. The Cleanup Doesn't.

The detection side of this story is arguably in decent shape. Coauthorship-network analysis, image-duplication scans, and a growing community of independent sleuths mean fabricated research gets flagged faster than it did a decade ago — Bik's original 2016 paper is still the reference point everyone cites for a reason (Physics World). Even adjacent forensic tools are moving fast: localization models that can point to exactly which region of a manipulated image was altered, and explain why, are now a live research problem in their own right (arXiv). What hasn't kept pace is the machinery for scrubbing flagged work out of the places it's already been cited — patents, guidelines, policy briefs — once it's there. Watch for whether Digital Science or similar groups extend this kind of downstream citation-tracing to other known paper mills beyond PNN. Right now the tooling for finding fabrication is outrunning the tooling for finding out where fabrication already went.