Season 1 · Episode 02 Forthcoming
Not all true things are allowed to decide.
Courts spent four centuries building rules for which evidence is allowed to settle a question — relevance, hearsay, authentication, chain of custody. This conversation puts that tradition next to the way machine-learning evaluation treats a test set, and asks what the older discipline knows that the newer one keeps forgetting: that a fact can be perfectly true and still inadmissible.
Audio publishes when Season 1 launches. Subscribe below and it will appear here and in your app the day it drops.
The law of evidence is one of the longest-running experiments in human reasoning, and its central discovery is unsettling: not everything true is allowed to count. Over roughly four centuries, common-law courts assembled a discipline of admissibility — relevance, hearsay, authentication, chain of custody, the rules against character and unfair prejudice — whose entire purpose is to keep certain true and even persuasive facts out of the decision. A statement might be accurate and still excluded because no one can be cross-examined on it. A document might be genuine and still excluded because its origin cannot be established. The tradition treats the question is this true? as separate from, and downstream of, the question is this allowed to decide? That second question is the one the law spent generations refining, and it is the one this episode is about.
Set that beside how machine-learning evaluation actually works. The held-out test set is supposed to be the courtroom — the protected place where a claim of competence is proven. But in practice the discipline treats nearly any data that moves a metric as fair to use. There is no doctrine of hearsay for a benchmark contaminated by training leakage, no authentication requirement for a label of unknown provenance, no chain of custody for a dataset that has been quietly scraped, relabeled, and re-split a dozen times. Evaluation asks did the number go up? and rarely asks was this evidence admissible to raise it? A model can post a state-of-the-art result that is, in the legal sense, built entirely on inadmissible proof — true-looking, metric-improving, and disqualified the moment you ask where it came from and whether the other side could have tested it.
The through-line of this show is that a conclusion is legitimate not because it was reached but because the path to it can be reconstructed and contested. Evidence law is, in that light, four centuries of machinery for keeping the path contestable — every exclusionary rule is really a rule about whether the losing party could have challenged the step. That is precisely the safeguard that vanishes when a system reports an answer without reporting what was allowed to produce it. The book this series accompanies, Admissible Reality, takes the word seriously on purpose: the goal is not merely accurate machines but machines whose conclusions could survive a hearing, where the evidence behind a claim is the kind a careful adversary would have been permitted to attack.
So the conversation is not a metaphor dressed up as one. It asks an evidence-law scholar to take the actual doctrine — its tests, its exceptions, its centuries of hard-won exclusions — and hold it against the loose evidentiary culture of model evaluation, to see what transfers and what does not. Where the law refuses a true fact, is it protecting something a benchmark also needs? Where it admits with conditions, what would the analogous condition be for a dataset or a score? The wager of the episode is that the older discipline has already done much of the thinking the newer one is about to need, and that admissible is the more demanding standard machine reasoning has been avoiding.
Chapter titles and timestamps are illustrative; the final episode is forthcoming.
"A benchmark records whether the number went up. Evidence law asks whether the number was allowed to." — Warrant, Episode 02
The episode is forthcoming; these are the questions the conversation will put to the guest, not answers given in advance.
Season 1 is in production. Subscribe now and this episode lands in your app the day it drops.