WARRANT · Summit Cognitive

← All episodes

Season 1 · Episode 02 Forthcoming

Admissible

Not all true things are allowed to decide.

Courts spent four centuries building rules for which evidence is allowed to settle a question — relevance, hearsay, authentication, chain of custody. This conversation puts that tradition next to the way machine-learning evaluation treats a test set, and asks what the older discipline knows that the newer one keeps forgetting: that a fact can be perfectly true and still inadmissible.

Listen

Episode 02 — Admissible

Warrant · Season 1

Not yet recorded

Audio publishes when Season 1 launches. Subscribe below and it will appear here and in your app the day it drops.

Show notes

The law of evidence is one of the longest-running experiments in human reasoning, and its central discovery is unsettling: not everything true is allowed to count. Over roughly four centuries, common-law courts assembled a discipline of admissibility — relevance, hearsay, authentication, chain of custody, the rules against character and unfair prejudice — whose entire purpose is to keep certain true and even persuasive facts out of the decision. A statement might be accurate and still excluded because no one can be cross-examined on it. A document might be genuine and still excluded because its origin cannot be established. The tradition treats the question is this true? as separate from, and downstream of, the question is this allowed to decide? That second question is the one the law spent generations refining, and it is the one this episode is about.

Set that beside how machine-learning evaluation actually works. The held-out test set is supposed to be the courtroom — the protected place where a claim of competence is proven. But in practice the discipline treats nearly any data that moves a metric as fair to use. There is no doctrine of hearsay for a benchmark contaminated by training leakage, no authentication requirement for a label of unknown provenance, no chain of custody for a dataset that has been quietly scraped, relabeled, and re-split a dozen times. Evaluation asks did the number go up? and rarely asks was this evidence admissible to raise it? A model can post a state-of-the-art result that is, in the legal sense, built entirely on inadmissible proof — true-looking, metric-improving, and disqualified the moment you ask where it came from and whether the other side could have tested it.

The through-line of this show is that a conclusion is legitimate not because it was reached but because the path to it can be reconstructed and contested. Evidence law is, in that light, four centuries of machinery for keeping the path contestable — every exclusionary rule is really a rule about whether the losing party could have challenged the step. That is precisely the safeguard that vanishes when a system reports an answer without reporting what was allowed to produce it. The book this series accompanies, Admissible Reality, takes the word seriously on purpose: the goal is not merely accurate machines but machines whose conclusions could survive a hearing, where the evidence behind a claim is the kind a careful adversary would have been permitted to attack.

So the conversation is not a metaphor dressed up as one. It asks an evidence-law scholar to take the actual doctrine — its tests, its exceptions, its centuries of hard-won exclusions — and hold it against the loose evidentiary culture of model evaluation, to see what transfers and what does not. Where the law refuses a true fact, is it protecting something a benchmark also needs? Where it admits with conditions, what would the analogous condition be for a dataset or a score? The wager of the episode is that the older discipline has already done much of the thinking the newer one is about to need, and that admissible is the more demanding standard machine reasoning has been avoiding.

Chapters

  1. 00:00Cold open — a true fact the court threw out
  2. 03:30Admissible is not the same as true
  3. 10:15Relevance, hearsay, and the right to cross-examine
  4. 17:40Authentication and chain of custody: where did this come from?
  5. 25:20The test set as a courtroom — and how it leaks
  6. 33:05What a benchmark has no doctrine for
  7. 40:30Could this score survive a hearing?
  8. 46:50Close — porting four centuries forward

Chapter titles and timestamps are illustrative; the final episode is forthcoming.

"A benchmark records whether the number went up. Evidence law asks whether the number was allowed to." — Warrant, Episode 02

What this episode asks

  1. If a fact can be perfectly true and still inadmissible, what is the corresponding category in machine-learning evaluation — and why does the field act as if it does not exist?
  2. Hearsay exists because no one can cross-examine an absent speaker. What is the test set's equivalent of cross-examination, and where exactly does it break down?
  3. Authentication and chain of custody ask where did this come from and who has touched it? What would it mean to demand the same lineage from a dataset or a label before a result is allowed to stand?
  4. The law excludes some true evidence because it is unfairly prejudicial. Is there a metric so persuasive that it should be excluded for the same reason?
  5. If we held a model's reported result to the standard of a hearing — could the losing side have challenged each step — how many state-of-the-art claims would actually be admissible?

The episode is forthcoming; these are the questions the conversation will put to the guest, not answers given in advance.

Subscribe

Season 1 is in production. Subscribe now and this episode lands in your app the day it drops.