Documentation
Everything Credda does today
Pre-release. Every command, flag and environment variable below is transcribed from the table credda --help renders.
What Credda is
Credda reads your repository on a push and reports the defects nobody filed. Label a bug report instead and it runs in your own CI, makes the failure happen, finds what caused it, writes the patch, and proves it with a test that fails before it and passes after. Credda proposes; it never merges.
Where the fix stage is today. The Fixer and Verifier have been on the investigation path since 2026-08-27. The provider that resolved decides a run’s shape.
- SignalA report, a failing command, a stack trace.
- InvestigateRead the repository; find where the claim lands.
- ReproduceMake the failure happen on demand; capture its signature.
- DiagnoseNarrow hypotheses against evidence until one is confirmed, or record that none was.
- ReportWhat executed, what it established, and what it did not.
- Fixnot enteredWrite the patch, inside Credda’s disposable checkout, never in your tree.
- Verifynot enteredRun a test that must fail before the patch and pass after, or discard the patch.
5 of 7 stages
The run stops at REPRODUCED_AND_DIAGNOSED and publishes the report. The scorecard records the fix rate as NOT_ATTEMPTED with that run’s attempt count, not a zero.
The invariant
Credda makes no claim it did not execute. A report it cannot turn into a runnable check stops there. A check that captures no failure says so.
No change required is a success: a reproduction ran and the reported thing did not happen.
NO_RUNNABLE_CHECK, where nothing ran, never carries a success tone.
Install
One workflow file and one label. No App, no webhook, no credential beyond the token the job already has.
Step one
# .github/workflows/credda.yml
name: Credda on labeled issues
on:
issues:
types: [labeled]
permissions:
contents: read # checkout of the repository under test
issues: write # the one report comment
id-token: write # mint the OIDC token that fetches the engine
jobs:
investigate:
if: github.event.label.name == 'credda'
runs-on: ubuntu-latest
timeout-minutes: 30
steps:
- uses: actions/checkout@v4
- uses: Credda-io/action@v1
with:
# @v1 defaults this to `credda,codereef`, so the run works without
# the line. Pin it anyway: the second name is there to carry old
# installs through the rename and is documented as temporary.
label: credda
# REQUIRED on a private repository: uncomment it. POST /v1/engine
# answers one with no licence 402 Payment Required, and the fetch
# fails before `Run Credda`. Public repos are never asked for a key.
# license: ${{ secrets.CREDDA_LICENSE }}
# Optional. Without it the heuristic provider reproduces, diagnoses and
# reports but writes no patch: the fix stage needs a model-backed one.
# with:
# anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}
At .github/workflows/credda.yml. It runs on a labelled issue, nothing else.
Step two
gh label create credda --description 'Credda reproduces this bug in a sandbox and comments what it established.'
Once, from a clone, or anywhere with --repo owner/name. The workflow runs because a label was applied, so it cannot create it.
Pin the tag, not the branch. A branch is whatever was last pushed to it. Pinning one means agreeing in advance to run code nobody has shown you, in your own repository, with your own token.
@v1 is the only published tag.
The three permissions
- contents: read
- Check out the repository under test. Credda cannot run code it cannot see.
- issues: write
- Post one comment on the issue that triggered the run. How a run reports back, and the only write Credda performs against a repository today.
- id-token: write
- Mint a short-lived OIDC token that identifies the workflow to Credda. How the engine is fetched: the repository is private, so the download is exchanged for a token GitHub signs about your job. It grants nothing over yours.
Not granted, so not spendable: contents: write, pull-requests, actions, packages, deployments.
What the first run does
- You label an issueThe credda label on any open issue is the whole trigger. Removing and re-adding it runs again after a push.
- The runner builds a sandboxThe first run builds the workspace image: minutes of package installation. Later runs on the same runner image reuse it, rebuilding only when the sandbox Dockerfile changes.
- Your repository executes with no networkRepository code runs in a Docker container with every network interface removed. A runner that cannot provide that plane gets a refusal, not a fallback to the host.
- One comment arrives, or the job summary doesA run that established something comments on the issue. One that established nothing writes to the job summary instead, and that document is about Credda, not about your bug.
Optional: a note on every issue opened
# MERGE this into the workflow above; do not append it.
# Replace that file's `on:` block with this one (`opened` arms triage) and add
# this job beside `investigate`. Appending gives the file two `on:` keys and two
# `jobs:` keys, which GitHub rejects as invalid YAML.
on:
issues:
types: [labeled, opened]
jobs:
triage:
if: github.event.action == 'opened'
runs-on: ubuntu-latest
timeout-minutes: 10
concurrency:
group: credda-issue-${{ github.event.issue.number }}
cancel-in-progress: false
steps:
# Optional. With it a refusal can say "this repository holds no
# `repro.js`" instead of staying quiet.
- uses: actions/checkout@v4
- uses: Credda-io/action@v1
with:
mode: triage
# The same name, OPPOSITE reason: in triage mode this is what Credda
# stays QUIET for, because an issue opened carrying it is about to get
# a full investigation. @v1 defaults to `credda,codereef`, so this
# name is already covered; pin it so the quiet set matches the label
# the workflow above triggers on and does not drift from it.
label: credda
# Uncomment on a private repository. Same 402 as above.
# license: ${{ secrets.CREDDA_LICENSE }}
A second job on the same trigger, reading issues as they open. No container, no install, no model call, no key. It asks for what a report is missing, or says nothing.
Reproductions are free on every public repository. A private one needs a licence key, as the license: line above says.
Keys are minted by hand and sent by email within a day. Pricing.
What a run reports about itself. One receipt per job at most: salted hashes of the owner, repository and triggering login, an outcome token, a duration, Credda’s version, and a public or private flag.
No repository name, issue text, source code, path, command output, branch or commit SHA, and no field in the protocol that could hold one. metering-url: '' makes no request at all. A metering failure cannot fail your job.
Outcome states
One terminal state per investigation, four of them successes. NO_RUNNABLE_CHECK is not among them: nothing ran.
- Ready for reviewREADY_FOR_REVIEW
A reviewable change backed by executable evidence. Credda proposes it; a human merges it.
- Patch rejectedPATCH_REJECTED
Credda discarded its own patch rather than hand it over unverified.
- Reproduced and diagnosedREPRODUCED_AND_DIAGNOSED
Reproduced, and a cause established from cited evidence. No code was changed.
- Reproduced, cause not establishedREPRODUCED_NOT_DIAGNOSED
Reproduced and its signature captured. No cause is offered, because none was confirmed against evidence. No code was changed.
- Reproduced; the tests assert itCONTRADICTS_SPECIFICATION
The reported behavior happened, and this repository owns passing tests asserting exactly that. Proved by running the one named test. Naming a cause would make this a defect claim, so none is named. No code was changed.
- No change requiredNO_CHANGE_REQUIRED
A reproduction ran and the reported failure was not there. A success, reachable only through a reproduction that actually executed.
- No runnable checkNO_RUNNABLE_CHECK
Nothing ran: no runnable check could be derived from the report, so nothing was established about this repository. Not an all-clear, and not a defect alleged.
- Insufficient evidenceINSUFFICIENT_EVIDENCE
Not enough was observed to justify a conclusion.
- CanceledCANCELLED
Stopped before completion.
- Run blockedFAILED
Credda's own run errored before reaching a verdict. An environment failure, not a defect in your code. Nothing was observed about the repository.
The evidence model
Every material claim cites a recorded artifact: a type, a phase, an ordinal strength.
BEFORE_PATCH, AFTER_PATCH or INDEPENDENT: which side of the comparison an artifact belongs to, or that the patching process did not write it.
STRONG, MODERATE or WEAK, never a percentage: there is no calibrated probability model.
The command, exit code, error class, normalised message, and origin file and line. Verification compares signatures, not raw log text.
ESTABLISHED, PARTIALLY_ESTABLISHED or NOT_ESTABLISHED, with the list of what this record does not establish.
Every answer the before/after comparison can give
Six, and two decline to answer.
- Failure resolvedFAILURE_RESOLVED
The same command, re-run against the patched workspace, exited cleanly.
- Failure unchangedFAILURE_UNCHANGED
The failure signature is identical after the patch. Nothing was fixed.
- Failure changed shapeFAILURE_MUTATED
The command still fails, but differently. This is the dangerous case: the original defect may be masked, or a new one introduced.
- New failureNEW_FAILURE
No failure signature was recorded before the patch, and one was recorded after it.
- Not re-runAFTER_NOT_RUN
A failure was demonstrated, but the failing command has not been re-run since. No claim is made about whether it still fails.
- Not comparableINCOMPARABLE
There is no before/after pair of observations of the same command, so no comparison is claimed.
Running the CLI from a clone
Published as `credda` on npm. `npm install -g credda` installs the same engine the Action runs.
For running the engine against a checkout on your own machine. The Action needs no local toolchain.
This clone fails outside our organisation.
The engine repository is private and stays private, so an unauthenticated git clone of it returns 404. The commands under it are unaffected once you have a checkout.
git clone https://github.com/Credda-io/core.git cd core corepack enable pnpm install # Every command below is 'pnpm credda': apps/cli, from source. pnpm credda --help
The first credda investigate against a path records the repository. The checkout is copied into a disposable workspace and never modified.
# Point Credda at a checkout and describe the failure. pnpm credda investigate ./my-service "Checkout 500s when the billing address is missing" # Shells mangle long reports; Windows truncates at the first newline. pnpm credda investigate ./my-service @report.md gh issue view 184 --json body -q .body | pnpm credda investigate ./my-service - # Read what it found. pnpm credda status pnpm credda report <investigation-id> pnpm credda events <investigation-id> --follow
The CLI defaults to your host. The Action does not. --sandbox local runs a repository’s build and test commands as the same user, on the same machine, as Credda. Pass --sandbox docker for the plane the Action insists on. See security.
Running this dashboard
A read-only view over the Credda API, except that it can queue an investigation. An unreachable API is reported with the URL it tried.
npm install npm run dev # http://localhost:3000 # API base URL (defaults to http://localhost:4317) NEXT_PUBLIC_CREDDA_API_URL=http://localhost:4317
The API allows http://localhost:3000 by default. CREDDA_CORS_ORIGIN takes a comma-separated list.
Commands
Eighteen spellings, fifteen commands. Any unambiguous prefix of an investigation id works.
credda investigate <repo-path> <description | @file | -> [options]Reproduce a reported failure, find its cause, and on a model-backed provider write the patch and prove it.
The description is inline, from a file with @path, or piped in as -. The patch is written in Credda’s disposable checkout, never your working tree. It opens no pull request; the Action’s `open-pull-request` input does, on your own GITHUB_TOKEN with `contents: write` and `pull-requests: write`.
credda discover <repo-path> [--out <dir>] [--max-files <n>] [--json]Read a checkout and write the bug reports nobody filed. Executes nothing.
With --out each report becomes a file `credda investigate @<file>` reads. STATED rows are settled by reading the repository; every other row is an unrun candidate, and pattern-growth candidates are counted rather than listed.
credda sweep <repo-path> [--max-candidates <n>] [--cost-ceiling <usd>] [--open-pull-request] [options]Discover, investigate each candidate, and open a PR for verified fixes (opt-in).
The whole loop end to end. Model-backed per candidate; --max-candidates caps how many runs start (default 3) and --cost-ceiling stops before starting a run once that much spend is recorded. Without --open-pull-request it reports what it would propose and pushes nothing; with it, a pull request opens for each run that carries a verified change, on the same gate the Action’s `open-pull-request` input enforces. Credda never merges.
credda docscan <repo-path> [--confirmed-only] [--json]Check a checkout against its own documented examples. Opens nothing.
Reads README fenced blocks and JSDoc @example blocks, admits the examples that are a self-contained call of the documented package with a literal expected value, executes each, and lists the ones whose output contradicts the documented value. No model call, no API key, no install, no network. Execution is a short-lived `node -e` child in the checkout — process-level isolation, not the engine sandbox — so it is safe for a checkout you trust and not production-safe for an untrusted one.
credda resolve | credda fix <repo-path> <description | @file | -> [options]Aliases for 'credda investigate', kept because scripts and docs use them.
investigate is the current name. Permanent aliases, not deprecated.
credda doctor [<repo-path>] [--deep]Check that this environment can reproduce a bug.
With a <repo-path> it reports the install, test, build and typecheck commands a run would use, read off the repository rather than executed.
credda init [--global] [--force]Write an credda.config.json with documented defaults.
--global writes to $CREDDA_HOME instead of the working directory.
credda status [--repository <path-or-id>] [--state <state>] [--outcome <outcome>] [--ref <ref>] [--limit <n>] [--offset <n>] [--json]List recent investigations.
Defaults to the most recent 20.
credda report <investigation-id-or-prefix> [--json] [--markdown] [--patch]Show what an investigation established, and what it did not.
Bug, Evidence, Reproduction, Root Cause and Confidence. Change always prints, and says so when a run stopped at the diagnosis. Verification prints when recorded.
credda resolution <investigation-id-or-prefix> [--json] [--markdown] [--patch]Alias for 'credda report', kept because scripts and docs use it.
report is the current name. A permanent alias, not deprecated.
credda inspect <investigation-id-or-prefix>Show everything one run recorded, in full.
The reproduction, every hypothesis including refuted ones, and the evidence records.
credda events <investigation-id-or-prefix> [--since <n>] [--follow] [--json]Show the event timeline for an investigation.
--follow (-f) tails a run until it reaches a terminal state.
credda triage <issue-file> [--repo <path>]Say what Credda could not use in a report, or say nothing.
Runs no repository code, starts no container, installs nothing, makes no model call, establishes nothing. Silence is the most common outcome; the refusal harvest on /benchmark measures how often.
credda validations [--repository <path-or-id>] [--state <state>] [--outcome <outcome>] [--limit <n>] [--offset <n>]List change-scoped validation runs.
credda validation <validation-id-or-prefix> [--severity <s>] [--status <s>] [--limit <n>] [--offset <n>]Show one validation: its checks, and the findings they raised.
credda cancel <investigation-id-or-prefix> [--reason <text>]Stop a running investigation, or say why it cannot be stopped.
credda reap [--dry-run] [--max-age-hours <n>]Remove sandbox containers left behind by an interrupted run.
Options
The aliases accept the same set.
| Option | Effect | Default |
|---|---|---|
| --sandbox <local|native|docker> | Execution plane. local and native both run repository code directly on this host; native adds enforced memory, CPU and process limits, which is real containment but not an isolation boundary. Only a local credda invocation may select either; anything else must use docker. Never falls back silently. | local |
| --provider <auto|heuristic|openai-compatible> | auto uses ANTHROPIC_API_KEY then CREDDA_OPENAI_API_KEY; heuristic forces rule-based reasoning; openai-compatible targets an OpenAI-compatible endpoint (NVIDIA NIM by default). | auto |
| --budget-minutes <n> | Wall-clock budget for the investigation. | 20 |
| --max-turns <n> | Maximum model calls across all agent roles. | 120 |
| --out <file> | Also write this run’s result to <file> as JSON. | not written |
| --ref <ref> | Record where this report came from: an issue reference, a URL, or the ref `credda discover` prints. Stored on the run and shown by `credda report`. Nothing branches on it. | not recorded |
Accepted by every command.
| Option | Effect |
|---|---|
| -h, --help | Show help and exit. |
| --version | Print the credda version and exit. |
| --json | Machine-readable JSONL on stdout, and nothing else. |
| --quiet | Print only the final outcome line. |
| --verbose | Include debug-severity events. |
| --no-color | Disable ANSI colour. NO_COLOR is honoured too. |
Configuration precedence, highest first: the flag, the environment variable, credda.config.json (searched upward from the working directory, then $CREDDA_HOME), the built-in default. Create one with credda init.
Environment
| Variable | Effect |
|---|---|
| CREDDA_HOME | Database and evidence store. Defaults to ./.credda |
| ANTHROPIC_API_KEY | Enables the Anthropic provider. Without it reasoning is rule-based: a reproduction, rarely a diagnosis. Every report names its provider. |
| CREDDA_PROVIDER | 'auto', 'heuristic' or 'openai-compatible'. 'heuristic' forces the deterministic provider. |
| CREDDA_MODEL | Overrides the Anthropic model id. |
| CREDDA_OPENAI_API_KEY | Enables the openai-compatible provider (NVIDIA NIM by default). NVIDIA_API_KEY also works. |
| CREDDA_OPENAI_BASE_URL | OpenAI-compatible base URL. Defaults to NVIDIA NIM. |
| CREDDA_OPENAI_MODEL | Model id served by that endpoint. |
| CREDDA_OPENAI_RPM | Client-side pacing. Defaults to 40, NVIDIA's free-tier limit. |
| CREDDA_SANDBOX | 'local' (default), 'native' or 'docker'. local and native both run repository code directly on this host, so both are only for repositories you trust; native adds enforced memory, CPU and process limits but is not an isolation boundary. Only a local credda invocation may select either: a repository arriving any other way is refused them and must use docker. Never falls back silently. |
| CREDDA_SANDBOX_IMAGE | Overrides the image the docker plane builds or pulls. |
| CREDDA_LOG_LEVEL | debug | info | warn | error. Defaults to warn. |
| NO_COLOR | Set to any value to disable ANSI colour. |
| CREDDA_ASCII | Set to any value to draw with ASCII instead of box characters. |
| TERM | 'dumb' is treated the same as NO_COLOR. |
Exit codes
| Code | Meaning |
|---|---|
| 0 | Credda executed something and its finding is on record: REPRODUCED_AND_DIAGNOSED, REPRODUCED_NOT_DIAGNOSED, NO_CHANGE_REQUIRED or INCONCLUSIVE. Abstention is a success. |
| 1 | Internal error. Credda failed; no verdict was reached. |
| 2 | Usage error. A bad flag, a missing value, or an unreadable input. |
| 3 | PATCH_REJECTED. Independent verification rejected the change. Model-backed providers only (ADR 0019): a rule-based run stops at the diagnosis and cannot return 3. |
| 4 | Canceled by the operator (Ctrl-C). |
| 5 | NO_RUNNABLE_CHECK. Nothing runnable could be derived from the report, so nothing ran. A fact about the report, not your code, and kept out of 0 so "credda investigate ... && deploy" cannot read it as a pass. |
No exit code carries the confidence class. Read it with credda report <id> --json.
Further reading
- Verification is not testingA test written after a patch, by the process that wrote the patch, proves the test agrees with the patch. It does not prove the defect existed, or that it is gone.
- The false modification problemAn agent that edits working code is worse than an agent that fixes less. Abstention has to be a first-class outcome, with its own success criteria, and it has to be measured.
- Failure signatures: comparing failures, not logsDiffing raw command output answers the wrong question. The question is whether this is the failure that was reported. The dangerous answer is a real failure that is the wrong one.