
Latest issue
The Benchmark Is the Product: How AI Safety Scores Get Built to Impress
7/28/2026
Pick any major AI model release in the past year. Somewhere in the announcement, you'll find a table: rows of benchmark names, columns of model versions, and a satisfying diagonal of bold numbers showing the new model winning. What that table almost never shows is what the benchm…
Recent posts
A 2003 paper by Dan Ariely and Klaus Wertenbroch became one of behavioral economics' most-cited findings: that people perform better when they can precommit to evenly spaced deadli…
Here's a number worth sitting with: in a survey of principal investigators running randomized controlled trials in sub-Saharan Africa, roughly 14% of those who failed to publish ci…
There's a version of p-hacking that everyone in research methods knows about: run your analysis twenty different ways, report the one that clears p < 0.05, and call it a day. It's…
Here's a scenario that should unsettle you. A research team runs a study on 40 participants, finds a striking effect, publishes in a respectable journal, and watches the press rele…
Last spring, eighteen teams of neuroscientists sat down with identical brain recordings and tried to answer the same question: which region had the highest density of hippocampal s…
A paper published in the latest issue of Methods in Ecology and Evolution proposes something called "adaptive preregistration" — a framework that, its authors argue, handles the me…
Imagine a drug that gets tested in twelve independent clinical trials. Three find a statistically significant benefit. Nine find nothing. If all twelve get published, the picture i…
A clinical trial enrolls a thousand patients. The new antidepressant beats placebo. The p-value is 0.03 — statistically significant, publishable, press-releasable. The drug reduces…
The problem with most fMRI research isn't that the findings are wrong. It's that they're underpowered to be right — and the field has known this for years without doing much about…
There's a particular irony in AI-generated hallucinations contaminating the research literature on AI safety. The field dedicated to making AI systems more reliable is partly built…
There's a version of open science that exists entirely on paper. Journals adopt data availability policies. Funders mandate sharing plans. Researchers check the box. And then, some…
A trial enrolls 800 patients. The new drug reduces symptom scores by 1.2 points on a 100-point scale. The p-value is 0.03. The press release says: "statistically significant result…
Two papers landed in Nature on April 1st — not a joke, though the timing has a certain poetry — and together they represent the most systematic accounting of social science reprodu…
There's a counterintuitive finding buried in a recent arXiv preprint that deserves more attention than it's likely to get: the papers that receive the harshest peer review tend to…
The average paper accepted by Nature spends somewhere between 8 and 14 months in the pipeline from submission to publication. That's not a bug. It's closer to the intended design.…
The standard critique of placebo controls in psychiatry goes like this: it's unethical to withhold treatment from people who are suffering. Give someone with severe depression a su…
Every major climate report comes with a range. Not a point estimate — a range. The IPCC's equilibrium climate sensitivity figure, to take the canonical example, has carried an unce…
The word does a lot of work it hasn't earned. When researchers publish an AI bias study and report that their model achieves "fairness," they almost always mean one specific, narro…
Most AI bias research doesn't actually study bias. It studies disparity — and those are not the same thing. Here's the distinction that gets quietly buried in methods sections: a d…


















