The manuscript that disappears before anyone votes on it

You ran the experiment. You followed the original protocol with the kind of fidelity that borders on devotion, recruited a larger sample than the study you were testing, pre-registered the design on OSF before collecting a single data point, and got nothing. Not a weak effect. Nothing. So you write it up, submit it to the journal that published the original finding, and wait. Six weeks later, a desk rejection lands in your inbox. No reviewers saw it. The editor thanks you for your interest and wishes you luck placing the work elsewhere.

That sequence is not random bad luck.

It is the internal escalation structure of academic publishing doing exactly what it was built to do, just not in the way most scientists will say out loud.

The ladder no one draws on the masthead

Every journal with a serious submission volume runs on triage. At the bottom sits the handling editor, often an early-career academic or a rotating associate who reads incoming manuscripts and makes a first cut. Above them, a senior or section editor. Above that, the editor-in-chief, who in practice touches only a small fraction of submissions. At the very top, for the prestige titles, an editorial board whose members rarely read individual papers at all but whose stated preferences filter down through every decision below them, the way a dress code shapes a room without anyone announcing it.

A replication failure almost never climbs this ladder. It gets stopped at the first rung.

Handling editors filter, consciously or not, for what the journal's brand rewards. High-impact journals built their reputations on surprising positive findings, and a paper reporting that a celebrated result does not hold carries a different prior into the editorial office than a paper announcing a new phenomenon. The handling editor knows, from years of reading the journal's output, what the editor-in-chief finds interesting. Null results rarely fit that picture. The desk rejection goes out. The decision never reaches anyone senior enough to override it.

This structural mechanism matters more than any individual editor's bias, and it is worth being direct about that. The hierarchy diffuses responsibility so completely that no single person feels they are suppressing anything. The handling editor thinks they are doing triage. The senior editor never sees the paper. The editor-in-chief would, in many cases, genuinely not know the submission existed.

What "not a priority for our readership" actually means

Desk rejection letters for null results cluster around a few polite formulations: the work is technically sound but not of sufficient general interest, or it does not represent a sufficient advance, or it falls outside current editorial priorities. These phrases are doing real work, and they deserve to be read as policy statements rather than courtesies.

"Sufficient advance" encodes a directional assumption: that science moves forward by adding confirmed effects to a growing pile. A replication failure, by this logic, merely removes something from the pile. Backward motion. The journal's impact factor, which depends on citations, is not well served by papers that mostly get cited as a footnote in methods sections.

Consider what this looks like at ground level. Two researchers, call them Priya and Marcus, both attempt to replicate a widely-cited social priming study. Priya's lab is well-resourced; she runs 400 participants, pre-registers, and finds an effect size near zero. Marcus, working with a smaller budget, runs 120 participants and finds a marginal positive result. Marcus's paper gets sent to peer review. Priya's gets a desk rejection in five weeks. The editorial system never formally compared the two submissions. It just filtered by pattern-matching against what looks publishable, and the larger, more careful study lost.

This is not invented cynicism. The Reproducibility Project, which coordinated replication attempts across dozens of psychology studies, found that many null replications struggled to find publication homes even after the project's overall findings made the replication crisis a public conversation. The mechanism hadn't changed. Only the external pressure had, and only temporarily.

How prestige cascades downward

There is a second-order effect that makes the desk-rejection problem worse, and it runs through the hierarchy in a less obvious direction.

When a high-prestige journal turns away a replication failure, the authors resubmit to a lower-tier journal. That journal's handling editor recognises the original paper being challenged as something published in a prestigious outlet. Paradoxically, the halo of that prestige works against the replication. The handling editor reasons that if the original passed rigorous review at that level, accepting a null result against it feels presumptuous. So the mid-tier journal also rejects, or requests revisions that soften the claims until the paper reads more like a report of a somewhat different effect than a straightforward failure to reproduce the original finding.

Meanwhile, the authors of the original finding are often asked to peer-review the replication attempt when it finally reaches a willing journal. Many journals send replications to the original authors as a matter of courtesy, framing it as a right of reply. Even where those authors act in good faith, the structural conflict is baked in, and good faith doesn't neutralise it.

Then there is what happens to replications that do reach peer review. Senior editors at several journals have written publicly about the pressure to solicit a third reviewer when the first two disagree on a null-result paper, and to weight the more skeptical review more heavily. This isn't always cynical. Editors genuinely worry about publishing a flawed replication that unfairly damages a researcher's career. But the asymmetry is real and difficult to defend as a consistent standard. A paper claiming a new effect needs two reviewers who are broadly convinced. A paper claiming that effect doesn't replicate often needs to survive four reviewers, a statistical consultant, and a demand for additional data collection. By the time that gauntlet ends, the authors have frequently moved on, and the null result sits in a drawer.

The drawer, incidentally, is not a metaphor. The file-drawer problem, named formally by Robert Rosenthal in 1979 and documented continuously since, describes exactly this accumulation of unpublished null results. What the escalation-structure analysis adds is an explanation of why the drawer fills: not primarily because researchers hide their failures, but because the editorial hierarchy routes null results away from publication before a meaningful decision is ever made.

The journals that tried to fix this, and what happened

Some journals created dedicated replication sections or adopted registered reports, a format where peer review happens before data collection, so that a pre-approved protocol guarantees publication regardless of outcome. PLOS ONE committed early to judging papers on methodological soundness rather than novelty of result. Several psychology journals introduced explicit replication policies once the crisis became impossible to ignore.

The results have been mixed, and that is worth sitting with rather than glossing over.

Registered reports do get null results published. But they represent a small fraction of total output at even the journals that championed them. The broader escalation structure remains intact. The editor-in-chief still sets the journal's brand. Handling editors still filter accordingly. And registered reports require authors to commit to a study before running it, which works reasonably well in psychology but fits awkwardly with how biology, chemistry, and clinical research actually operate in practice.

The journals that created separate replication sections also discovered an unintended consequence. Sequestering replications in a dedicated section turns out to be a polite form of burial. If replication failures live in a corner that regular readers skip, they accumulate without actually correcting the scientific record. A finding cited 4,000 times doesn't get uncited because a null result appears in a supplementary section of a niche outlet. The correction exists on paper. It travels nowhere.

The correction that never circulates

So here is the honest version: scientific publishing was designed to announce discoveries, not to maintain an evolving, self-correcting record. The escalation hierarchy reflects that original purpose faithfully, perhaps too faithfully. Editors at prestigious journals are rewarded, in prestige and in metrics, for publishing work that generates attention and citation. A replication failure, even a methodologically impeccable one, closes a question rather than opening one. It generates less downstream activity almost by definition.

If the system were working as its defenders claim, the threshold of methodological scrutiny applied to a replication failure would not routinely exceed the threshold applied to the original positive finding that created the problem in the first place. That the reverse is so consistently true points to something structural rather than incidental.

The deeper issue is that the escalation structure makes the failure invisible in a way that feels, from the inside, like ordinary professional judgment. Nobody formally decided to suppress the finding. The handling editor made a reasonable local decision. The senior editor was never consulted. The editor-in-chief would point, with complete sincerity, to their journal's stated commitment to rigor. And the null result sits unread, while the original paper continues to be cited, taught in undergraduate courses, and built upon in grant applications across a dozen countries.

The hierarchy didn't suppress the truth. It just made sure the question never reached anyone with the authority, or the information, to answer it.