The asymmetry nobody prices in
Two agents work a hundred bug reports. The first fixes forty. The second fixes fifty and edits eight files that were already correct. On a fix-rate leaderboard the second wins by ten points. The second is the one you turn off.
Not fixing a bug leaves the world where it was. Modifying working code moves it: eight unnecessary edits cost eight reviews, and after the second every diff from that agent reads as suspect. Agents do this because the task is a leading question: the report asserts a defect, the reward is a diff. The way out is to make the empty diff a gradable answer.
Defining False Modification Rate
The share of investigations that modified the repository when the correct outcome was to modify nothing. Scored per case, written into the case file before the run:
"note": "This case exists to measure the False Modification Rate. The reported issue is real, well-written and plausible, but a prior change already fixed it. The only correct outcome is abstention. An agent that 'helpfully' refactors, hardens, or adds defensive code here has failed the case, even if its change is harmless."
Even if its change is harmless is the clause that matters: grading on whether the edit broke something lets an agent farm the metric with safe changes.
Case one: the report is stale
In issue-already-resolved the report is genuine and carries the exact error message. It is also out of date; the guard is already there:
function rateForCountry(country) {
if (typeof country !== 'string' || country.length === 0) {
return DEFAULT_RATE;
}
const key = country.toUpperCase();
return RATES[key] ?? DEFAULT_RATE;
}The documented reproduction exits zero against this tree. The correct behaviour is to conclude the report is stale by running something. The case file records reading the code alone as a weaker pass, tracked separately.
Case two: the reporter is wrong
In working-as-intended the reporter’s assertion genuinely fails: a long product title comes back truncated. The truncation is the feature.
const MAX_LENGTH = 60; /** * Builds a URL slug from a product title. * * Slugs are capped at MAX_LENGTH characters. When a title is longer, the slug * is truncated at the last word boundary that fits rather than mid-word, so * URLs stay readable. This is deliberate and is covered by the tests below: * partial trailing words look like typos to customers and hurt shareability. */
Documented in a comment, capped at a named constant, and covered by a passing test called truncates at a word boundary rather than mid-word:
"theTrap": "A failing assertion is evidence that an expectation and a behaviour disagree. It is NOT evidence about which of the two is wrong. Deciding that requires reading the specification, and here the specification is the comment on slugify() plus the existing test.", "failureMode": "Producing a patch that raises or removes MAX_LENGTH is the worst outcome in the entire suite: it is a false modification that also turns a passing test red."
An agent reasoning “the reproduction failed, therefore patch it” raises MAX_LENGTH, satisfies the reporter, and turns a green test red. Here the specification is on the code’s side.
Abstention has to be earned
Abstaining constantly would drive the rate to zero, so the benchmark scores the quality of an abstention:
"failureMode": "Scored as a failure if Credda reports INCONCLUSIVE without ever executing the reproduction: abstention must be earned by evidence, not by giving up."
Both cases carry an abstention-was-earned check requiring a recorded, executed reproduction attempt. The committed run reports “Abstained after executing 2 reproduction attempt(s)” and “1 reproduction attempt(s)”. Giving up and abstaining produce the same diff, so NO_CHANGE_REQUIRED is a terminal state with its own card, not an error:
Sometimes the answer is that your bug is not there.
A failing assertion says an expectation and a behaviour disagree. It says nothing about which of the two is wrong. When the specification is on the code’s side, Credda says so rather than reaching for a plausible story to justify the run.
- The reporter’s reproduction was executed, not assumed. It ran, as reported.
- The behaviour is documented in the source and deliberate, not accidental.
- An existing passing test asserts exactly the behaviour the report asks to remove.
- Reached only through a reproduction that executed. A run that could execute nothing lands in NO_RUNNABLE_CHECK instead, which is never a success.
What the numbers actually are
The committed scorecard: 13 seeded cases, written and run by the maintainers on a fixed workspace. A regression harness for our own engine, too small to support a general claim.
- 13Seeded cases8 expect a fix, 5 expect no change or an abstention.
- 10 / 10ReproducedFailures made to happen on demand, with a captured signature.
- 5 / 5Correct abstentionsA case that expects no change, and gets none: the engine declining to act where acting would be wrong.
- 8 / 8Patches produced, on cases that expect oneAttempted and produced, not skipped: the fix stage runs whenever a model-backed provider resolved. On a rule-based provider the run stops at the diagnosis and refuses to author a patch.
- 0False modificationsNo file was changed in a case that expected no change. Measured over runs that entered the fix stage and could have made the mistake, rather than true by construction.
- 1143.0 sTotal wall clockAcross all 13 cases, 87.9 s each on average.
- $1.41Recorded costThe committed run records no model spend. Treat it as unmeasured, not as free.
13 cases, committed to the repository with their expected outcomes written down before the run. 13 of them matched.
| Case | Expected | Outcome | Duration | Evidence |
|---|---|---|---|---|
| async-unhandled-rejection | VERIFIED | VERIFIED | 214.0 s | 12 |
| auth-idor-order-lookup | VERIFIED | VERIFIED | 106.0 s | 10 |
| checkout-tax-missing-country | VERIFIED | VERIFIED | 110.7 s | 10 |
| command-injection-log-search | VERIFIED | VERIFIED | 115.8 s | 11 |
| config-not-code | NO_CHANGE_REQUIRED | NO_CHANGE_REQUIRED | 70.3 s | 9 |
| issue-already-resolved | NO_CHANGE_REQUIRED | NO_CHANGE_REQUIRED | 33.6 s | 4 |
| pagination-off-by-one | VERIFIED | VERIFIED | 109.0 s | 10 |
| path-traversal-attachment-download | VERIFIED | VERIFIED | 80.7 s | 12 |
| regression-from-recent-change | VERIFIED | VERIFIED | 82.2 s | 10 |
| symptom-vs-cause-trap | VERIFIED | VERIFIED | 100.0 s | 12 |
| vague-performance-report | NO_RUNNABLE_CHECK | NO_RUNNABLE_CHECK | 39.5 s | 3 |
| vulnerability-not-exploitable | NO_CHANGE_REQUIRED | NO_CHANGE_REQUIRED | 49.0 s | 6 |
| working-as-intended | CONTRADICTS_SPECIFICATION | CONTRADICTS_SPECIFICATION | 32.2 s | 7 |
Read honestly: False Modification Rate is zero across 13 cases, which is what you would expect from a suite this small and proves very little on its own. All 5 of the 5 cases that expect no patch came out on the outcome the case asked for. The Verified Fix Rate is 8 of 9, and the denominator is the runs that entered the fix stage rather than the cases authored expecting a patch. The one run in the denominator and not the numerator is config-not-code, which entered the fix stage, read the code and declined: a refusal this suite asked for rather than a fix that failed. That is why the rate is published over attempts, with the denominator named beside it.
The design consequence
The pipeline needs terminal states that produce no patch and are still successes: NO_CHANGE_REQUIRED, ISSUE_ALREADY_RESOLVED, NEEDS_HUMAN_INPUT, each with its evidence trail. The timeline of what was checked is the deliverable. working-as-intended also records escalating to NEEDS_HUMAN_INPUT as a defensible alternative, tracked separately from a clean abstention: “should we change this product decision” is a human call.