Last spring, eighteen teams of neuroscientists sat down with identical brain recordings and tried to answer the same question: which region had the highest density of hippocampal sharp-wave ripples? Twelve of the seventeen reporting teams reached the same conclusion — no significant difference between regions. Sounds like consensus. Except when you look at what each team actually measured, the agreement evaporates. Some teams detected virtually no ripples at all. Others found up to ten ripples per minute in each region. The shared conclusion of "no difference" was built on observations so divergent that they might as well have been analyzing different brains.
This is the finding from the Brainhack hackathon organized at the Champalimade Foundation in March, written up by Gaëlle Chapuis and Mattia Chini in The Transmitter. It's a small study, and the authors are careful not to overclaim. But the result is genuinely unsettling, because it targets a corner of neuroscience that was supposed to be immune to this kind of problem.
Electrophysiology Was the Safe Harbor
Brain imaging's replication troubles are well-documented at this point. When I covered this last May, the core problem with fMRI was clear: small samples, flexible analysis pipelines, and the sheer number of researcher decisions baked into any given workflow. The field has been grappling with this for years, and a landmark 2020 Nature study — in which 70 independent teams analyzed the same fMRI dataset and reached substantially different conclusions — became a standard reference for why. That study is cited directly in The Transmitter's Brainhack writeup as the methodological precedent for the hackathon design.
Electrophysiology was supposed to be different. You're recording actual electrical signals from neurons, not inferring blood-flow proxies. The data is harder, more direct, less dependent on modeling assumptions. Ripple detection in particular has a long methodological history and even a recent consensus paper meant to standardize how researchers identify these events.
The Brainhack results suggest the consensus paper didn't solve the problem. Teams used three broad approaches: the consensus-paper method, deep-learning detectors, and classical bandpass-and-threshold techniques. All are defensible. All produced radically different absolute estimates. The apparent agreement on the final answer — "no difference" — was a statistical coincidence, not a sign of methodological convergence.
The Deeper Issue: We Don't Know What We're Measuring
What makes this finding stick is the specific nature of the disagreement. The teams weren't arguing about statistical thresholds or p-values. They were disagreeing about what a ripple is — how to define it, how to detect it, what counts as a real event versus noise. This is a conceptual problem, not a computational one.
That kind of foundational ambiguity has structural causes. One is publication pressure: a 2017 survey of 1,151 psychology journals found that just 3 percent explicitly stated they welcome replication studies, which means the incentive to nail down definitions through repeated testing is weak. Journals reward novelty; they don't reward the unglamorous work of establishing that two labs are even measuring the same thing.
The infrastructure compounds the problem. A recent arXiv preprint describing EEG-Dash, an open-source platform cataloguing nearly 800 publicly archived neurophysiological recordings, found that in a May 2026 audit, only about one in three OpenNeuro EEG datasets passed the official Brain Imaging Data Structure (BIDS) validator — and even validator-compliant datasets sometimes failed to load correctly. When researchers can't reliably share or reuse each other's datasets, the field loses its ability to catch conceptual disagreements before they harden into published literature.
The authors of the Brainhack study have launched a follow-up project, CON²PHYS (CONceptual CONsistency in electroPHYSiology), aimed at systematically quantifying how much disagreement exists when neuroscientists interpret fundamental concepts in their own field. The preliminary results suggest the answer is: more than anyone suspected.
What Consensus Actually Requires
The Brainhack study is worth sitting with because it reframes what "replication" means in practice. We tend to think of replication as running the same experiment again and checking whether the result holds. But if eighteen teams analyzing identical data can't agree on what they found, then replication in the traditional sense may be insufficient. The problem isn't that experiments fail to replicate — it's that the original measurements themselves are underdetermined.
Some methodological responses are already in motion. The normative modelling framework updated in a bioRxiv preprint from Marquand's group at Radboud represents one serious attempt to address variability at the level of computational psychiatry — building models that explicitly account for individual variation rather than assuming group averages are meaningful. Separately, new sensor technologies like the red-shifted acetylcholine sensors described in Nature Neuroscience are pushing toward more direct, multiplexed measurements of neural activity, which may eventually reduce the interpretive ambiguity that plagues indirect proxies.
None of that fixes the upstream conceptual problem. And the political environment isn't helping: a proposed OMB rule described in The Atlantic would shift control over public science funding toward political appointees, threatening the kind of long-horizon, methodologically unglamorous work — replications, standards development, null results — that a field in conceptual crisis most needs.
The CON²PHYS project is asking the right question. Before neuroscience can solve its replication problem, it needs to know exactly how deep the disagreement runs — not just across labs, but within the basic vocabulary researchers use to describe what they see. Watch for their first systematic results, which should clarify whether Brainhack was a worst case or a representative sample of how the field actually operates.
