WARRANT · Summit Cognitive

← All episodes

Season 2 · Episode 09 In development

Replication

A result no one can reproduce is a rumor with a p-value.

Science calls something knowledge once it survives being done again. The reproducibility crisis is what happens when that step gets skipped — when a finding is published, cited, and built upon before anyone has confirmed it holds. This conversation reads that crisis not as a statistics problem but as an admissibility problem: a result that cannot be reproduced is not yet entitled to count, no matter how convincing the figure looks.

Listen

Episode 09 — Replication

Warrant · Season 2

Not yet recorded

Audio publishes when Season 2 launches. Subscribe below and it will appear here and in your app the day it drops.

Show notes

Science earns the right to call something knowledge only after the result survives being produced again — by other hands, in another lab, from the same setup. That is what makes the reproducibility crisis so disquieting: a large share of published findings, across fields, turn out not to reproduce when someone actually tries. The usual diagnosis points at statistics — underpowered studies, flexible analysis, the quiet hunt for a significant p-value. But the deeper issue is not that the numbers are wrong. It is that a finding which cannot be reproduced has never cleared the bar that would make it admissible in the first place. It is a claim, often an accurate-looking one, that no one has yet been allowed to test. A result no one can reproduce is a rumor with a p-value.

The reforms that the field has converged on are, read closely, a discipline of provenance. Preregistration fixes the question and the analysis before the data arrive, so the hypothesis cannot be quietly rewritten to fit the result — it is a commitment recorded in advance, the way a chain of custody records who touched the evidence and when. Sharing the data and the code does the rest: it turns a printed conclusion back into the inputs and the steps that produced it, so that the path from raw observation to reported number is open to inspection. Each of these is the scientific analog of a reconstructable record — not a louder assertion that the finding is real, but the materials a skeptic would need to find out for themselves.

And the strongest check of all is not a better explanation; it is a replay — taking the same data and the same code and re-running the analysis to see whether the same number comes out. A study whose author can hand you the inputs and the procedure, and whose result reappears when you execute it yourself, has done something a beautifully argued discussion section never can: it has made its conclusion contestable on equal terms. This is the same move the rest of the show keeps returning to. Legitimacy does not come from the prestige of the journal or the authority of the lab. It comes from the fact that the route to the result can be rebuilt and checked — replays, not explanations. The book this series accompanies, Admissible Reality, is built on exactly that standard: a conclusion counts when its warrant survives reconstruction.

What makes the unreproducible result genuinely dangerous is that it does not look broken. It is accurate-looking — a clean effect, a tidy figure, a confident abstract — and that surface plausibility is precisely what lets it travel. It gets cited, it shapes the next grant, it becomes the premise other work is stacked upon, all before anyone has confirmed it holds. The conversation puts this to a metascientist: if the cure for the crisis is provenance and replay rather than rhetoric, then reproducibility is not a hygiene step at the end of science but the moment a claim actually becomes evidence — and everything that cannot pass it, however polished, is still only a rumor with a p-value.

Chapters

  1. 00:00Cold open — a famous result that would not come back
  2. 04:10What "reproduce" actually means, and why fields disagree
  3. 11:00A finding is not yet knowledge — it is a claim awaiting a hearing
  4. 18:25Preregistration as a commitment recorded in advance
  5. 26:00Shared data and code: turning a conclusion back into its inputs
  6. 33:40Replay — re-running the analysis as the strongest check
  7. 41:15The danger of the accurate-looking, unreproducible result
  8. 47:30Close — reproducibility as the moment a claim becomes evidence

Chapter titles and timestamps are illustrative; the final episode is forthcoming.

"A finding you cannot run again is not a result yet. It is a rumor that happened to clear a threshold." — Warrant, Episode 09

What this episode asks

  1. If a finding that cannot be reproduced is not yet knowledge, what is it — and at what exact point in the scientific process should a claim be allowed to start counting as evidence?
  2. Preregistration, shared data, and shared code are all forms of provenance. Which one does the most work, and what failure mode does each of them actually close?
  3. Re-running an analysis on the same inputs — a replay — is a weaker test than an independent new study, yet many published results fail even that. Why is the cheapest check so often the one nobody runs?
  4. An unreproducible result is dangerous precisely because it looks accurate. How do you tell a finding that is true from one that is merely accurate-looking before anyone has reproduced it?
  5. If legitimacy comes from reconstructability rather than authority, what would it take for a field to treat "we could replay it and the number returned" as the threshold a claim must pass to be cited at all?

The episode is forthcoming; these are the questions the conversation will put to the guest, not answers given in advance.

Subscribe

Season 2 is in development. Subscribe and this episode lands in your app the day it drops.