Evidence

Every case, including the ones that went badly.

158 real bug reports from 60 open-source repositories we did not choose, at the commits where the bug was present, report bodies unedited.

This is bench/external, the scored slice. The reproduce corpus holds 357 admitted cases in JavaScript, Python, Elixir; admission is free and scoring is not, so every rate on this page divides by what was scored. Every corpus, and what each one measured.

01Ledger

What the run counted.

All three from one committed scorecard, one moment, one corpus.

  • 110 of 158

    right failures: the reported failure reproduced and captured, over the cases whose pinned commit still contains the defect. Over all 158 entries: 110.

    bench/external/scorecard.json, 2026-09-20

  • 11 of 158

    wrong failures: a genuine signature captured for a failure the report never described. Worse than capturing nothing: the run holds real evidence for the wrong defect.

    bench/external/scorecard.json, 2026-09-20

  • 0 of 158

    false successes: a claimed success sitting on top of a captured failure. The gate that counts these fails the build whenever the count moves.

    bench/external/scorecard.json, 2026-09-20

scorecard, every fieldbench/external/scorecard.json
  • cases scored158
  • cases with the defect in the pinned commit158
  • reproductions executed156
  • signatures captured121
  • right failures110
  • wrong failures11
  • false successes0
  • errored5
  • total wall time18408.5s
Per-check detail is on each entry’s own page, in both gradings. Source LIVE, measured 2026-09-20.

02Record

Every entry, in corpus order.

Nothing sorted or filtered. Each entry carries the signature the run captured, verbatim, or says none was.

Each row’s verdict is the LIVE run of 2026-09-20, the tree executed against all 158 upstream checkouts. Every case page also carries its RECORDED verdict, which carries the benchmark gate. Corpus and grader: bench/external/.

How the corpus was chosen

Nothing was dropped after seeing a result.

Exclusion criteria were fixed before any run. The candidate log, included and excluded, is published on the benchmark page with the date of each decision.

bench/external/README.md · candidate logverbatim
No candidate was excluded after an Credda run. The included set is every candidate whose body was read in full and found to contain a concrete reproduction.