The test that always passes
Ask a coding agent to fix a bug and it hands back a patch and a test. The test was written after the patch, by the same process, with the patch already in the tree. Its author had every opportunity to write an assertion the old code satisfied equally well. A passing test tells you the test and the patch agree, and nothing about the defect.
One cheap experiment turns it into evidence, and almost nothing in this category runs it: execute the test against the code before the patch, and require it to fail.
Ordering is the whole mechanism
/** * Independent verification of a patch. * * The ordering here is the point. The regression test is written and executed * against UNPATCHED code first, and must FAIL. Only then is the patch applied * and the same test run again. A test authored after a patch that passes proves * only that the test agrees with the patch; a test that fails before and passes * after proves the defect existed and is gone. * * Everything below is mechanical. No model opinion enters the verdict. */
runRegressionTestBeforePatch() treats a non-zero exit code as the desired result, and a passing test as a problem to report:
const failed = outcome.result.exitCode !== 0;
const note = failed
? null
: 'The regression test passed against unpatched code, so it does not demonstrate the reported defect.';
return { status: failed ? 'FAIL' : 'PASS', evidenceIds: [outcome.evidenceId], note };A test that passes against unpatched code means it does not exercise the defect, or the defect is not there. A reviewer needs to see either one.
A verdict is a function, not an opinion
No model decides what happened. Verification records observations, each the exit status of a command that ran:
export interface VerificationSignals {
readonly reproductionBefore: 'FAIL' | 'PASS' | 'NOT_RUN';
readonly reproductionAfter: 'FAIL' | 'PASS' | 'NOT_RUN';
readonly regressionTestBefore: 'FAIL' | 'PASS' | 'NOT_RUN';
readonly regressionTestAfter: 'FAIL' | 'PASS' | 'NOT_RUN';
readonly existingTests: { passed: number; failed: number; total: number } | null;
readonly build: 'PASS' | 'FAIL' | 'NOT_RUN' | 'NOT_APPLICABLE';
readonly typecheck: 'PASS' | 'FAIL' | 'NOT_RUN' | 'NOT_APPLICABLE';
readonly lint: 'PASS' | 'FAIL' | 'NOT_RUN' | 'NOT_APPLICABLE';
}No confidence score, no probability, no “likely fixed”. NOT_RUN means a check was applicable and did not happen; NOT_APPLICABLE means the repository has no such command. Those are different facts about a patch.
const reproFixed = s.reproductionBefore === 'FAIL' && s.reproductionAfter === 'PASS';
const regressionProves =
s.regressionTestBefore === 'FAIL' && s.regressionTestAfter === 'PASS';
// Without a demonstrated before-failure there is nothing to have fixed.
if (s.reproductionBefore !== 'FAIL' && s.regressionTestBefore !== 'FAIL') {
return 'INCONCLUSIVE';
}
// Any hard gate failing rejects the patch outright.
if (s.build === 'FAIL' || s.typecheck === 'FAIL') return 'REJECTED';
if (s.existingTests !== null && s.existingTests.failed > 0) return 'REJECTED';
if (s.reproductionAfter === 'FAIL' || s.regressionTestAfter === 'FAIL') return 'REJECTED';
if (!reproFixed && !regressionProves) return 'UNVERIFIED';
const fullySupported =
reproFixed && regressionProves && s.existingTests !== null && s.existingTests.failed === 0;
return fullySupported ? 'VERIFIED' : 'PARTIALLY_VERIFIED';Reading the gates
No before-failure, no verdict. If neither the reproduction nor the regression test failed beforehand, the result is INCONCLUSIVE. No demonstrated defect, nothing to have fixed.
Collateral damage is rejection. A failing build, a failing typecheck, or a single newly failing existing test produces REJECTED outright.
Partial is a real answer. VERIFIED requires all three of: the original failure gone, a regression test that failed before and passes after, and a parsed, clean existing-test run. Miss one and it is PARTIALLY_VERIFIED. Nothing rounds up.
Unparseable output degrades honestly. When the runner’s output cannot be parsed into pass/fail counts, Credda does not guess at numbers:
if (!summary.parsed) {
// Reporting invented counts would be worse than reporting none. The exit
// code still carries a usable signal.
notes.push(
`Test output from ${profile.testRunner} could not be parsed; falling back to the exit code (${outcome.result.exitCode}).`,
);
return outcome.result.exitCode === 0 ? null : { passed: 0, failed: 1, total: 1 };
}Returning null stops deriveVerdict() reaching VERIFIED, which requires existingTests !== null. Not knowing costs the run its top verdict.
What the signal table looks like
Because the verdict is a function of the signals, the signals can be shown verbatim. The checkout-tax-missing-country case, whose repository ships ten existing tests:
Signals are recorded observations. The verdict is derived from them mechanically.
| Signal | Before patch | After patch |
|---|---|---|
| Reproduction | FAIL | PASS |
| Regression test | FAIL | PASS |
| Existing tests | 10 / 10 | 10 / 10 |
| Build | not recorded | PASS |
| Typecheck | not recorded | PASS |
A reviewer can disagree with the verdict by disagreeing with a row. A confidence percentage can only be believed.
Where this is honest about its limits
None of this establishes that a patch is good, only that a demonstrated failure is gone, a test catches its return, and nothing the repository already checked has broken. A patch can clear every gate and still be the wrong design.
A project with no test command gets existingTests: null and cannot reach VERIFIED: Credda inherits the state of your CI rather than substituting for it.