Benchmark

Four corpora, and what each one actually measured.

How often a reported failure is reproduced on input we did not choose. The other half, finding defects nobody filed, is on what Credda found.

Every reading carries its artifact and date; none is averaged with another. Live against the upstream checkouts on 2026-09-20: 110 of 158 over the cases that were scored. The corpus holds 158 admitted cases, so 0 are admitted and not yet scored; admission is free and scoring is not, and no rate here divides by them. From the transcripts recorded earlier: 0 of 8. Both gradings are published, per case. Of those 110 live right failures, at most 76 are textually exact matches; the other 34 or more reached their match with the grader normalising one side — whitespace, grouping parentheses, a V8 message rename, or the expression re-spelled as the report pinned it. That is a bound rather than a split, because 15 of them cannot be checked from this card at all: it was written before the run recorded which spelling a match went through. Every relaxation is argued case by case and none was widened to produce this page — bench/external/normalization-audit-2026-09-20.json, reproduce with cd core && npx tsx scripts/replay-normalizations.ts. Beyond it the reproduce corpus now holds 357 admitted cases across JavaScript, Python, Elixir — section 10 — of which only the 158 above have been scored.

01The panel

Four readings, four artifacts, four things they do not say.

  1. In-house seeded suite

    measured 2026-08-29

    10 of 10

    Reproduced on demand, with a captured signature, over the reproducible cases.

    bench/scorecard.json, 2026-08-29

    Does not sayNothing about the field. Every case was written by the harness authors, against repositories they built.

  2. External corpus, executed live

    measured 2026-09-20

    110 of 158

    Right failures over the cases actually scored. Real closed issues, reports unedited, upstream checkouts.

    bench/external/scorecard.json, 2026-09-20

    Does not sayFixing. The provider was rule-based and model-free, so no verified fix was possible.

  3. Labelled corpus, constructed ground truth

    measured 2026-08-28

    7 of 7

    Reproduced the reported failure, over the placed defects. The only corpus with applications certified to contain nothing wrong, which is what makes its other readings possible.

    bench/labelled/scorecard.json, 2026-08-28

    Does not sayFinding anything. Every case arrives with a written report and is scored on reproducing it. The applications are small and constructed: this is known ground, not hard ground.

  4. Refusal harvest, unfiltered real inbound

    measured 2026-08-29

    280 of 727

    Issues that got a refusal naming something in the report. The frequency argument behind the paid half.

    bench/harvest/scorecard.json, 2026-08-29

    Does not sayWhether anybody wants one. Nothing was cloned or executed: this is what could be said about a report, not what a run said or a maintainer thought.

Ratios, never percentages. 280 of 727 is checkable against a corpus; 39 per cent is not.

02Results, external corpus

The grader’s output, unedited.

Right Failure Rate is the only rate that counts: a signature that is not the reported failure is a real failure standing in for a defect the run never executed.

credda-bench externalsource RECORDED · exit 0
$ credda-bench external

METRICS

  Scope                     REPRODUCE_AND_REPORT (ADR 0015). This corpus measures reproduction only.
  Observations              8 transcribed from earlier runs
  Execution plane           not recorded: these are transcripts
  Right Failure Rate        0% (0/8)   <- the headline, over the 8 case(s) whose defect is present
  Wrong failure captured    7
  Reproduction Executed     88% (7/8)
  Signature Captured        100% (7/7)   <- of the reproductions that executed
  False Successes           0   <- must be zero
  Unproven Successes        0   <- must be zero: a reproduction claimed over the wrong failure
  Recorded False Successes  100% (3/3) still detected
    outcome NO_CHANGE_REQUIRED        3
    outcome INCONCLUSIVE              5
  Errored                   0
  Runtime-only trees        0
  Total duration            28.8s

exit 0

GATE HELD The must be zero line is zero and the run exits 0. This grades the committed RECORDED transcripts, where reproduction succeeded 0 of 8 times; the live run reads 110 of 158. The <- the headline marker is the grader’s own label, left in place because the transcript is verbatim, and it names a different run from this page’s headline. The zero is narrow: 3 of these 8 transcripts did claim a success over a contradicting failure, and the corpus names them: camelcase#46, picomatch#49, node-semver#801. What cannot regress is that no new one appeared and that 3 of 3 recorded ones are still caught.

bench external · every case, both gradings158 rows

A captured failure that is not the reported one is not partial credit. The run holds a real signature it can present as evidence, and the defect the reporter described was never executed, which is worse than reproducing nothing, because it looks like progress.

Every case in the external corpus, with what the reporter described beside what each run captured, and the verdict from both gradings.
CaseWhat the reporter describedWhat the live run capturedLIVE2026-09-20RECORDED
TinyColor-103tinycolor("#red").toString() produces 'red'; the fix makes it produce '#000000'.`tinycolor("#red").toString()` still produces "red" (read red)RIGHT_FAILURENOT_GRADED
TinyColor-36tinycolor('rgb(0.9, 0.0, 0.0)').toRgbString() produces 'rgb(1, 0, 0)'; the fix makes it produce 'rgb(230, 0, 0)'.nothing, no failure signature capturedNO_FAILURE_OBSERVEDNOT_GRADED
bytes-31bytes('250Kg') produces 250; the fix makes it produce null.`bytes('250Kg')` still produces 250 (read 250)RIGHT_FAILURENOT_GRADED
bytes-61bytes.parse("string without any numbers") produces NaN; the fix makes it produce null.`bytes.parse("string without any numbers")` still produces NaN (read NaN)RIGHT_FAILURENOT_GRADED
camelcase-11camelcase('fooFoo bar') produces 'foofooBar'; the fix makes it produce 'fooFooBar'.nothing, no failure signature capturedNO_FAILURE_OBSERVEDNOT_GRADED
camelcase-4camelCase('', '') produces '-'; the fix makes it produce ''.`camelCase('', '')` still produces '-' (read -)RIGHT_FAILURENOT_GRADED
camelcase-46camelcase('A::a') returns 'a:-:a' instead of leaving the non-alphabetic separator alone.`camelcase("A::a")` still produces "a:-:a" (read a:-:a)RIGHT_FAILUREWRONG_FAILURE
camelcase-52camelcase('Hello1World') produces 'hello1world'; the fix makes it produce 'hello1World'.nothing, no failure signature capturedNO_FAILURE_OBSERVEDNOT_GRADED
camelcase-77camelcase('volume_3d') and camelcase('volume3d') both return 'volume3D'.`camelcase('volume_3d')` still produces volume3D (read volume3D)RIGHT_FAILUREWRONG_FAILURE
camelcase-98camelCase('-') produces '-'; the fix makes it produce ''.`camelCase('-')` still produces '-' (read -)RIGHT_FAILURENOT_GRADED
camelcase-keys-13camelcaseKeys([{'foo-bar': true}], {deep:true}) does not produce [{"fooBar": true}]; the fix makes it do so.nothing, no failure signature capturedNO_FAILURE_OBSERVEDNOT_GRADED
camelcase-keys-68camelcaseKeys({'4.2': 'foo'}) produces { '42': 'foo' }; the fix makes it produce { '4.2': 'foo' }.`camelcaseKeys({'4.2': 'foo'})` still produces {'42': 'foo'} (read { '42': 'foo' })RIGHT_FAILURENOT_GRADED
camelcase-keys-80camelcaseKeys(({ '_name': true })) produces { name: true }; the fix makes it produce { _name: true }.`camelcaseKeys(obj)` still produces { name: true } (read { name: true })RIGHT_FAILURENOT_GRADED
chalk-194Tagging a template literal whose interpolated expression evaluates to undefined or null throws TypeError: Cannot read property 'toString' of undefined out of chalkTag, instead of rendering the value.TypeError: Cannot read properties of undefined (reading 'toString')RIGHT_FAILURENOT_GRADED
cheerio-1101cheerio.load('')('<span><svg ><use xlink:href="#1" /></svg></span>').find('use').prop('xlink:href') does not produce undefined; the fix makes it do so.nothing, no failure signature capturedNO_FAILURE_OBSERVEDNOT_GRADED
cheerio-116($('<div class="foo">bar</div>')).toString() does not produce '<div class="foo">bar</div>'; the fix makes it do so.`html.toString()` still produces '<div class="foo">bar</div>' (read [object Object])RIGHT_FAILURENOT_GRADED
cheerio-915$('<xmp><h2></xmp>').html() produces '<h2></h2>'; the fix makes it produce '<h2>'.`$('<xmp><h2></xmp>').html()` still produces '<h2></h2>' (read <h2></h2>)RIGHT_FAILURENOT_GRADED
clsx-17clsx('stack', { pop: true, push: true }) produces 'stack'; the fix makes it produce 'stack pop push'.`clsx('stack', { pop: true, push: true })` still produces 'stack' (read stack)RIGHT_FAILURENOT_GRADED
color-convert-73convert.rgb.hcg.raw([250, 0, 255]) does not produce [298.8235294117647, 100, 0]; the fix makes it do so.`convert.hcg.rgb.raw(convert.rgb.hcg.raw([250, 0, 255]))` still produces [ 0, 255, 249.99999999999991 ] (read [ 0, 255, 249.99999999999991 ])RIGHT_FAILURENOT_GRADED
cookie-21cookie.parse('expires=Wed, 29 Jan 2014 17:43:25 GMT; Path=/') produces { Path: '/', expires: 'Wed' }; the fix makes it produce { Path: '/', expires: 'Wed, 29 Jan 2014 17:43:25 GMT' }.`cookie.parse('expires=Wed, 29 Jan 2014 17:43:25 GMT; Path=/')` still produces {Path: '/', expires: 'Wed'} (read { expires: 'Wed', Path: '/' })RIGHT_FAILURENOT_GRADED
cron-parser-239parser.parseExpression('* * * 2 *').stringify() produces '* * 1-29 2 *'; the fix makes it produce '* * * 2 *'.`parser.parseExpression('* * * 2 *').stringify()` still produces "* * 1-29 2 *" (read * * 1-29 2 *)RIGHT_FAILURENOT_GRADED
cron-parser-424(parser.parse('0 0 16 * 0-6', ({ currentDate: new Date('2026-01-01T00:00:00Z'), tz: 'UTC' }))).stringify() produces '0 0 16 * *'; the fix makes it produce '0 0 16 * 0-6'.TypeError: parser.parse is not a functionWRONG_FAILURENOT_GRADED
cron-parser-442(CronExpressionParser.parse('20 15 * *')).stringify(true) produces '* 0 20 15 * *'; the fix makes it produce '0 * 20 15 * *'.`interval.stringify(true)` still produces '* 0 20 15 * *' (read * 0 20 15 * *)RIGHT_FAILURENOT_GRADED
culori-118culori.parse('hsla(219, 34%, 46%, 1)') produces { alpha: 0.00392156862745098, h: 219, l: 0.46, mode: 'hsl', s: 0.34 }; the fix makes it produce { alpha: 1, h: 219, l: 0.46, mode: 'hsl', s: 0.34 }.nothing, no failure signature capturedNO_FAILURE_OBSERVEDNOT_GRADED
dayjs-2230dayjs(Date.parse('0001-01-01')).format('YYYY-MM-DD') produces '1-01-01'; the fix makes it produce '0001-01-01'.`dayjs(Date.parse('0001-01-01')).format('YYYY-MM-DD')` still produces '1-01-01' (read 1-01-01)RIGHT_FAILURENOT_GRADED
dayjs-244dayjs() instanceof dayjs produces false; the fix makes it produce true.`dayjs() instanceof dayjs` still produces false (read false)RIGHT_FAILURENOT_GRADED
dayjs-3015dayjs('2024-01-02').format('Y') does not produce 'Y'; the fix makes it do so.`dayjs('2024-01-02').format('Y')` still produces 'Y' (read +0000)RIGHT_FAILURENOT_GRADED
decamelize-21decamelize('ADDRESS1') produces 'addres_s1'; the fix makes it produce 'address1'.`decamelize('ADDRESS1')` still produces 'addres_s1' (read addres_s1)RIGHT_FAILURENOT_GRADED
deepmerge-150Symbol-keyed properties are dropped by merge().`z` still produces { value: 42, other: 33 } (read { value: 42, other: 33 })RIGHT_FAILUREWRONG_FAILURE
deepmerge-23RegExp values are not carried through a merge; result.test is undefined, so result.test.test() throws TypeError.TypeError: result.test.test is not a functionRIGHT_FAILURENOT_GRADED
dot-prop-27dotProp.has(({foo: undefined}), 'foo') produces false; the fix makes it produce true.`dotProp.has(a, 'foo')` still produces false (read false)RIGHT_FAILURENOT_GRADED
dot-prop-38dotProp.get((undefined), 'attributes.required', false) does not produce false; the fix makes it do so.`dotProp.get(fields, 'attributes.required', false)` still produces false (read undefined)RIGHT_FAILURENOT_GRADED
fast-xml-parser-317validate('<?xml version="1.0" encoding="utf-8"?><Content><?tibco-char 19?> content</Content>') does not produce true; the fix makes it do so.nothing, no failure signature capturedNOT_EXECUTEDNOT_GRADED
filenamify-13(filenamify(("This? This is very long filename that will lose its extension when passed into filenamify, which could cause issues.csv"))) produces 'This! This is very long filename that will lose its extension when passed into filenamify, which cou'; the fix makes it produce 'This! This is very long filename that will lose its extension when passed into filenamify, which cou.csv'.`safeName` still produces "This! This is very long filename that will lose its extension when passed into filenamify, which cou" (read This! This is very long filename that will lose its extension when passed into filenamify, which cou)RIGHT_FAILURENOT_GRADED
filesize-77filesize(1/8, { bits: true }) produces '1024 '; the fix makes it produce '1 b'.`filesize(1/8, { bits: true })` still produces '1024 ' (read 1024 )RIGHT_FAILURENOT_GRADED
filter-obj-20Object.getOwnPropertyDescriptor((includeKeys((Object.defineProperty({}, 'prop', { value: true, enumerable: true, writable: false, configurable: false })), ['prop'])), 'prop') produces { configurable: true, enumerable: true, value: true, writable: true }; the fix makes it produce { configurable: false, enumerable: true, value: true, writable: false }.`Object.getOwnPropertyDescriptor(objCopy, 'prop')` still produces { value: true, writable: true, enumerable: true, configurable: true } (read { value: true, writable: true, enumerable: true, configurable: true })RIGHT_FAILURENOT_GRADED
immutable-1040fromJS({ a: 1, b: 2 }).take(Number.MAX_SAFE_INTEGER).toJS() produces { a: 1 }; the fix makes it produce { a: 1, b: 2 }.`fromJS({ a: 1, b: 2 }).take(Number.MAX_SAFE_INTEGER).toJS()` still produces { a: 1 } (read { a: 1 })RIGHT_FAILURENOT_GRADED
immutable-1247(new (Record({ a: 1 }))()).wasAltered() produces true; the fix makes it produce false.`r.wasAltered()` still produces true (read true)RIGHT_FAILURENOT_GRADED
immutable-240Immutable.List([1,2,3]).remove(0).keys().next() produces { done: false, value: -1 }; the fix makes it produce { done: false, value: 0 }.`Immutable.List([1,2,3]).remove(0).keys().next()` still produces { value: -1, done: false } (read { value: -1, done: false })RIGHT_FAILURENOT_GRADED
immutable-406Immutable.List([1, 2, 3]).merge(Immutable.List([4, 5])).toJSON() produces [ 4, 5, 3 ]; the fix makes it produce [ 1, 2, 3, 4, 5 ].`Immutable.List([1, 2, 3]).merge(Immutable.List([4, 5])).toJSON()` still produces [4, 5, 3] (read [ 4, 5, 3 ])RIGHT_FAILURENOT_GRADED
immutable-480Immutable.Map({a: 1, b: 2}).valueSeq().filter(function (n) {return n > 1}).take(10).toJS() does not produce [ 2 ]; the fix makes it do so.nothing, no failure signature capturedNO_FAILURE_OBSERVEDNOT_GRADED
immutable-703Immutable.is((Immutable.fromJS([1, 2])).get(((Immutable.fromJS([1, 2])).lastIndexOf(((Immutable.fromJS([1, 2])).max())))), ((Immutable.fromJS([1, 2])).max())) does not produce true; the fix makes it do so.nothing, no failure signature capturedNO_FAILURE_OBSERVEDNOT_GRADED
immutable-86(Immutable.Vector(1, 2, 3)).splice().length produces NaN; the fix makes it produce 3.`Immutable.Vector(1, 2, 3).splice().length` still produces NaN (read NaN)RIGHT_FAILURENOT_GRADED
is-109is.numericString(' ') produces true; the fix makes it produce false.`is.numericString(' ')` still produces true (read true)RIGHT_FAILURENOT_GRADED
is-number-3isNumber(' ') produces true; the fix makes it produce false.`isNumber(' ')` still produces true (read true)RIGHT_FAILURENOT_GRADED
joi-121Joi.validate([], Joi.types.Object()) produces null; the fix makes it produce { [Error: the value of <root> must be an object] _errors: [ { message: 'the value of <root> must be an object', path: null } ], _object: [] }.`Joi.validate([], Joi.types.Object())` still produces null (read null)RIGHT_FAILURENOT_GRADED
joi-2176joi.types().func produces undefined; the fix makes it produce { '$_root': { ValidationError: [class (anonymous) extends Error], [Symbol(@hapi/lab/coverage/initialize)]: [Function: trace], _types: Set(12) { 'alternatives', 'any', 'array', 'binary', 'boolean', 'date', 'function', 'link', 'number', 'object', 'string', 'symbol' }, allow: [Function (anonymous)], alt: [Function (anonymous)], alternatives: [Function (anonymous)], any: [Function (anonymous)], array: [Function (anonymous)], assert: [Function: assert], attempt: [Function: attempt], binary: [Function (anonymous)], bool: [Function (anonymous)], boolean: [Function (anonymous)], build: [Function: build], cache: { provision: [Function: provision] }, checkPreferences: [Function: checkPreferences], compile: [Function: compile], custom: [Function (anonymous)], date: [Function (anonymous)], defaults: [Function: defaults], disallow: [Function (anonymous)], equal: [Function (anonymous)], exist: [Function (anonymous)], expression: [Function: expression], extend: [Function: extend], forbidden: [Function (anonymous)], func: [Function (anonymous)], function: [Function (anonymous)], in: [Function: in], invalid: [Function (anonymous)], isExpression: [Function: isTemplate], isRef: [Function (anonymous)], isSchema: [Function (anonymous)], link: [Function (anonymous)], not: [Function (anonymous)], number: [Function (anonymous)], object: [Function (anonymous)], only: [Function (anonymous)], optional: [Function (anonymous)], options: [Function (anonymous)], override: Symbol(override), preferences: [Function (anonymous)], prefs: [Function (anonymous)], ref: [Function: ref], required: [Function (anonymous)], string: [Function (anonymous)], strip: [Function (anonymous)], symbol: [Function (anonymous)], trace: [Function: trace], types: [Function: types], untrace: [Function (anonymous)], valid: [Function (anonymous)], version: '16.1.7', when: [Function (anonymous)], x: [Function: expression] }, '$_super': { default: [Function: bound default] }, '$_temp': { ruleset: null, whens: {} }, '$_terms': { alterations: null, dependencies: null, examples: null, externals: null, keys: null, metas: [], notes: [], patterns: null, renames: null, shared: null, tags: [], whens: null }, _cache: null, _flags: {}, _ids: { _byId: Map(0) {}, _byKey: Map(0) {}, _schemaChain: false }, _invalids: null, _preferences: null, _refs: { refs: [] }, _rules: [], _singleRules: Map(0) {}, _valids: null, type: 'function' }.nothing, no failure signature capturedNO_FAILURE_OBSERVEDNOT_GRADED
joi-2404Joi.array().ordered(Joi.string().required(), Joi.number().default(0), Joi.number().default(6)).required().validate(['test']).value produces [ 'test' ]; the fix makes it produce [ 'test', 0, 6 ].Error: Cannot find module '@hapi/joi'WRONG_FAILURENOT_GRADED
js-yaml-117yaml.dump(-0.0) produces '0\n'; the fix makes it produce '-0.0\n'.`yaml.dump(-0.0)` still produces '0\n' (read 0 )RIGHT_FAILURENOT_GRADED
js-yaml-220A float written in scientific notation does not survive a dump/load round-trip: yaml.safeLoad(yaml.safeDump({foo: 5e-324})) gives {foo: "5e-324"}, a string, rather than the number 5e-324.`yaml.safeLoad('foo: 5e-324').foo` still produces "5e-324" (read 5e-324)WRONG_FAILURENOT_GRADED
js-yaml-303yaml.load("'''foo''\n'") produces "'foo "; the fix makes it produce "'foo' ".`yaml.load("'''foo''\n'")` still produces "'foo " (read 'foo )RIGHT_FAILURENOT_GRADED
js-yaml-321yaml.safeLoad('[,,]') returns [null, null]; the elided entries of a flow sequence are parsed as null nodes instead of raising a YAMLException for invalid YAML.`yaml.safeLoad('[,,]')` still produces [null, null] (read [ null, null ])RIGHT_FAILURENOT_GRADED
js-yaml-784loadAll('a:\n---\nx: 1\n') produces [ {}, { x: 1 } ]; the fix makes it produce [ { a: null }, { x: 1 } ].`loadAll('a:\n---\nx: 1\n')` still produces [ {}, { x: 1 } ] (read [ {}, { x: 1 } ])WRONG_FAILURENOT_GRADED
lodash-1012_.snakeCase('enable24HFormat') produces 'enable_24_hf_ormat'; the fix makes it produce 'enable_24_h_format'.nothing, no failure signature capturedNO_FAILURE_OBSERVEDNOT_GRADED
lodash-1038_.difference(undefined, [1]) produces [ 1 ]; the fix makes it produce [].nothing, no failure signature capturedNO_FAILURE_OBSERVEDNOT_GRADED
lodash-1061_.merge({set:{}}, {set:[]}) produces { set: [ undefined ] }; the fix makes it produce { set: [] }.`_.merge({}, [])` still produces {} (read {})WRONG_FAILURENOT_GRADED
lodash-379_.map( ([ [1,2,3], [4,5,6], [7,8,9] ]), _.max ) produces [ 3, -Infinity, -Infinity ]; the fix makes it produce [ 3, 6, 9 ].`_.map( a, _.max )` still produces [3, -Infinity, -Infinity] (read [ 3, -Infinity, -Infinity ])RIGHT_FAILURENOT_GRADED
lodash-69_.merge({a:1}, {a:2}, {a:3}, {a:4}) produces { a: 3 }; the fix makes it produce { a: 4 }.`_.merge({a:1}, {a:2}, {a:3}, {a:4})` still produces {a:3} (read { a: 3 })RIGHT_FAILURENOT_GRADED
luxon-1058Duration.fromISO('PT9.5H').toObject() produces {}; the fix makes it produce { hours: 9.5 }.`Duration.fromISO('PT9.5H').toObject()` still produces {} (read {})RIGHT_FAILURENOT_GRADED
luxon-1068(new luxon.FixedOffsetZone('CDT')).isValid produces true; the fix makes it produce false.`zone.isValid` still produces true (read true)RIGHT_FAILURENOT_GRADED
luxon-1070luxon.Duration.fromObject({hour: 2.4}).toFormat('hh:mm') produces '02:23'; the fix makes it produce '02:24'.`luxon.Duration.fromObject({hour: 2.4}).toFormat('hh:mm')` still produces "02:23" (read 02:23)RIGHT_FAILURENOT_GRADED
luxon-709Interval.fromDateTimes(DateTime.fromISO('2016-05-25T00:00:00'), DateTime.fromISO('2016-05-25T00:00:00')).hasSame('day') does not produce true; the fix makes it do so.`interval.hasSame('day')` still produces false (read false)RIGHT_FAILURENOT_GRADED
luxon-882luxon.Duration.fromISO("PT-1.5S").toISO() produces 'PT-0.5S'; the fix makes it produce 'PT-1.5S'.`duration.toISO()` still produces "PT-0.5S" (read PT-0.5S)RIGHT_FAILURENOT_GRADED
matcher-13matcher.isMatch('rainbow', '!unicorn') does not produce true; the fix makes it do so.`matcher.isMatch('rainbow', '!unicorn')` still produces true (read false)RIGHT_FAILURENOT_GRADED
mathjs-291math.format((1e+27), {notation: 'fixed'}) produces '1e+27'; the fix makes it produce '1000000000000000000000000000'.`math.format(x, {notation: 'fixed'})` still produces "1e+27" (read 1e+27)RIGHT_FAILURENOT_GRADED
mathjs-2936mod(0.15, 0.05) produces 0.04999999999999999; the fix makes it produce -2.7755575615628914e-17.nothing, no failure signature capturedNO_FAILURE_OBSERVEDNOT_GRADED
mathjs-2964distance([8,0],[10,0],[5,0]) produces 3.5777087639996634; the fix makes it produce 0.nothing, no failure signature capturedNO_FAILURE_OBSERVEDNOT_GRADED
mathjs-3100round(0.145*100) produces 14; the fix makes it produce 15.nothing, no failure signature capturedNO_FAILURE_OBSERVEDNOT_GRADED
mathjs-680math.eval('W<=5*Z', {W:'50',Z:'300'}) produces false; the fix makes it produce true.`math.eval('W<=5*Z', {W:'50',Z:'300'})` still produces false (read false)RIGHT_FAILURENOT_GRADED
micromatch-11(mm.makeRe(('maiden/{**/*.,*.}'), ({}))).test('maiden/code') produces true; the fix makes it produce false.`re.test('maiden/code')` still produces true (read true)RIGHT_FAILURENOT_GRADED
micromatch-24mm.isMatch('markup/modules/exampleModule/assets/image.png', 'markup/modules/**/assets/**/*.*') produces false; the fix makes it produce true.`mm.isMatch('markup/modules/exampleModule/assets/image.png', 'markup/modules/**/assets/**/*.*')` still produces false (read false)RIGHT_FAILURENOT_GRADED
micromatch-45micromatch.isMatch('coffee+/src/glimini.js', 'coffee+/src/**') produces false; the fix makes it produce true.`micromatch.isMatch('coffee+/src/glimini.js', 'coffee+/src/**')` still produces false (read false)RIGHT_FAILURENOT_GRADED
micromatch-91micromatch(["bar/bar"], ["foo/**", "!foo/baz"]) does not produce []; the fix makes it do so.`micromatch(['bar/bar'], ['foo/**', '!foo/baz'])` still produces ["bar/bar"] (read [ 'bar/bar' ])WRONG_FAILURENOT_GRADED
micromatch-96mm.not(["C:\\bla\\bar.xml"], ["*.xml"], {basename: true, unixify: true}) produces [ 'C:\\bla\\bar.xml' ]; the fix makes it produce [].`mm.not(["C:\\bla\\bar.xml"], ["*.xml"], {basename: true, unixify: true})` still produces [ 'C:\\bla\\bar.xml' ] (read [ 'C:\\bla\\bar.xml' ])RIGHT_FAILURENOT_GRADED
minimatch-215makeRe(('some/path/**')).test(('some/path-but-different')) produces true; the fix makes it produce false.`makeRe('some/path/**').test('some/path-but-different')` still produces true (read true)RIGHT_FAILURENOT_GRADED
minimatch-5minimatch('js/lib/.svn/tmp', '**/.svn/**') produces false; the fix makes it produce true.`minimatch('js/lib/.svn/tmp', '**/.svn/**')` still produces false (read false)RIGHT_FAILURENOT_GRADED
minimist-30minimist(['--aa=false', '--bb=false'], {boolean: ['a', 'bb'], alias: {a: 'aa', bb: 'b'}}) produces { _: [], a: 'false', aa: 'false', b: false, bb: false }; the fix makes it produce { _: [], a: false, aa: false, b: false, bb: false }.`minimist(['--aa=false','--bb=false'], {boolean:['a','bb'], alias:{a:'aa', bb:'b'}})` still produces { _: [], a: 'false', aa: 'false', bb: false, b: false } (read { _: [], a: 'false', aa: 'false', bb: false, b: false })RIGHT_FAILURENOT_GRADED
moment-1075moment('1371065286', ['X']).isValid() produces false; the fix makes it produce true.`moment('1371065286', ['X']).isValid()` still produces false (read false)RIGHT_FAILURENOT_GRADED
moment-1083moment('2013-09-13 7:26 am').format('hh:mm a') produces '12:00 am'; the fix makes it produce '07:26 am'.nothing, no failure signature capturedNO_FAILURE_OBSERVEDNOT_GRADED
moment-1290moment("2013-11-21T10:10:56 Z").calendar() produces 'Invalid date'; the fix makes it produce '11/21/2013'.nothing, no failure signature capturedNO_FAILURE_OBSERVEDNOT_GRADED
moment-323moment.utc(1338499506000) produces { _d: Invalid Date, _isUTC: true }; the fix makes it produce { _d: 2012-05-31T21:25:06.000Z, _isUTC: true }.`moment.utc(1338499506000)` still produces { _d: Invalid Date, _isUTC: true } (read { _d: Invalid Date, _isUTC: true })RIGHT_FAILURENOT_GRADED
moment-92moment("12:00", "HH:mm").diff(moment("08:00", "HH:mm"), "hours") produces -8; the fix makes it produce 4.`end.diff(start, "hours")` still produces -8 (read -8)RIGHT_FAILURENOT_GRADED
ms-103ms(-1 * 60 * 1000, { long: true }) produces '-60000 ms'; the fix makes it produce '-1 minute'.nothing, no failure signature capturedNO_FAILURE_OBSERVEDNOT_GRADED
ms-22ms('1m') produces NaN; the fix makes it produce 60000.`ms('1m')` still produces NaN (read NaN)RIGHT_FAILURENOT_GRADED
ms-70ms(ms(-123)) produces undefined; the fix makes it produce -123.nothing, no failure signature capturedNO_FAILURE_OBSERVEDNOT_GRADED
mustache-330Mustache.render("{{#nums}}{{.}}, {{/nums}}", {nums: [0, 1, 2]}) produces '[object Object], 1, 2, '; the fix makes it produce '0, 1, 2, '.`Mustache.render("{{#nums}}{{.}}, {{/nums}}", {nums: [0, 1, 2]})` still produces "[object Object], 1, 2, " (read [object Object], 1, 2, )RIGHT_FAILURENOT_GRADED
normalize-url-149normalizeUrl('http://sindresorhus.com/?url=http://example.com', { sortQueryParameters: true }) produces 'http://sindresorhus.com/?url=http%3A%2F%2Fexample.com'; the fix makes it produce 'http://sindresorhus.com/?url=http://example.com'.`normalizeUrl('http://sindresorhus.com/?url=http://example.com', { sortQueryParameters: true })` still produces 'http://sindresorhus.com/?url=http%3A%2F%2Fexample.com' (read http://sindresorhus.com/?url=http%3A%2F%2Fexample.com)RIGHT_FAILURENOT_GRADED
normalize-url-187normalizeUrl('sindresorhus.com:123', { removeExplicitPort: true }) does not produce 'http://sindresorhus.com'; the fix makes it do so.`normalizeUrl('sindresorhus.com:123', { removeExplicitPort: true })` still produces 'http://sindresorhus.com' (read sindresorhus.com:123)RIGHT_FAILURENOT_GRADED
object-inspect-6oi(new Number(5)) produces '{}'; the fix makes it produce 'Object(5)'.`oi(new Number(5))` still produces '{}' (read {})RIGHT_FAILURENOT_GRADED
path-to-regexp-148path2regexp('', [], {end: false}).exec('a/b') produces null; the fix makes it produce [ '', groups: undefined, index: 0, input: 'a/b' ].`re2.exec('a/b')` still produces null (read null)RIGHT_FAILURENOT_GRADED
pathe-18dirname('test.mjs') produces '/'; the fix makes it produce '.'.nothing, no failure signature capturedNOT_EXECUTEDNOT_GRADED
picomatch-142pm('test(/utils/**)')('test/utils') produces false; the fix makes it produce true.`pm('test(/utils/**)')('test/utils')` still produces false (read false)RIGHT_FAILURENOT_GRADED
picomatch-187pm.isMatch('a', '[!abc]') returns true; POSIX bracket negation is inverted, so the class matches exactly the characters it should exclude.`pm.isMatch('a', '[!abc]')` still produces true (read true)RIGHT_FAILURENOT_GRADED
picomatch-2((picomatch('../*.js'))('../.test.js')) produces true; the fix makes it produce false.`result` still produces true (read true)RIGHT_FAILURENOT_GRADED
picomatch-49picomatch.parse('{foo}').output returns '(foo)'; a single-item brace should stay literal.`picomatch.parse('{foo}').output` still produces '(foo)' (read (foo))RIGHT_FAILUREWRONG_FAILURE
pluralize-119pluralize('passerby') produces 'passerbies'; the fix makes it produce 'passersby'.nothing, no failure signature capturedNO_FAILURE_OBSERVEDNOT_GRADED
pluralize-123pluralize.singular("whiskies") produces 'whiskey'; the fix makes it produce 'whisky'.nothing, no failure signature capturedNO_FAILURE_OBSERVEDNOT_GRADED
pluralize-21pluralize("you") produces 'yous'; the fix makes it produce 'you'.`pluralize("you")` still produces "yous" (read yous)RIGHT_FAILURENOT_GRADED
pluralize-22pluralize("olives", 1) produces 'olife'; the fix makes it produce 'olive'.`pluralize("olives", 1)` still produces "olife" (read olife)RIGHT_FAILURENOT_GRADED
pluralize-28pluralize('is', 1) produces 'i'; the fix makes it produce 'is'.nothing, no failure signature capturedNO_FAILURE_OBSERVEDNOT_GRADED
pretty-ms-7prettyMs(1337000000, {verbose: true}) does not produce '15 days 11 hours 23 minutes 20 seconds'; the fix makes it do so.`prettyMs(1337000000, {verbose: true})` still produces '15 days 11 hours 23 minutes 20 seconds' (read 15d 11h 23m 20s)RIGHT_FAILURENOT_GRADED
qs-357qs.parse({ color: 'a,b' }, { comma: true }) does not produce { color: [ 'a', 'b' ] }; the fix makes it do so.`qs.parse('color=a,b', { comma: true })` still produces { color: [ 'a', 'b' ] } (read { color: [ 'a', 'b' ] })WRONG_FAILURENOT_GRADED
qs-37qs.parse("a[]=&a[]=b&a[]=c") produces { a: [ 'b', 'c' ] }; the fix makes it produce { a: [ '', 'b', 'c' ] }.`qs.parse("a[]=&a[]=b&a[]=c")` still produces { a: [ 'b', 'c' ] } (read { a: [ 'b', 'c' ] })RIGHT_FAILURENOT_GRADED
qs-390qs.stringify({"foo(ref)": "BAR"}, {format: "RFC1738"}) produces 'foo%28ref%29=BAR'; the fix makes it produce 'foo(ref)=BAR'.nothing, no failure signature capturedNO_FAILURE_OBSERVEDNOT_GRADED
qs-45Qs.parse('a=b&a[]=b') does not produce { a: ['b', 'b'] }; the fix makes it do so.`Qs.parse('a=b&a[]=b')` still produces { a: ['b', 'b'] } (read { a: 'b' })RIGHT_FAILURENOT_GRADED
qs-514qs.parse("a=1&a=2&b[]=1&b[]=2", {duplicates: "last"}) produces { a: '2', b: [ '2' ] }; the fix makes it produce { a: '2', b: [ '1', '2' ] }.`qs.parse("a=1&a=2&b[]=1&b[]=2", {duplicates: "last"})` still produces { a: '2', b: [ '2' ] } (read { a: '2', b: [ '2' ] })RIGHT_FAILURENOT_GRADED
query-string-1qs.parse('foo=c++') produces { foo: 'c++' }; the fix makes it produce { foo: 'c ' }.`qs.parse('foo=c++')` still produces {foo: 'c++'} (read { foo: 'c++' })RIGHT_FAILURENOT_GRADED
query-string-296With arrayFormat: 'comma', a single value that itself contains commas does not survive a stringify/parse round-trip: the percent-encoded commas are split, so { not_important: ["I, am, one, single, value"] } comes back as five elements.nothing, no failure signature capturedNO_FAILURE_OBSERVEDNOT_GRADED
query-string-302queryString.stringify({ a: ["asd", null, "123", ""] }, { arrayFormat: "comma", sort: false }) produces 'a=asd,123'; the fix makes it produce 'a=asd,,123,'.`queryString.stringify( { a: data }, { arrayFormat: "comma", sort: false, } )` still produces a=asd,123 (read a=asd,123)RIGHT_FAILURENOT_GRADED
query-string-346queryString.stringifyUrl({ url: 'https://foo.bar', query: { top: 'foo' }, fragmentIdentifier: '/bar/hello' }) produces 'https://foo.bar?top=foo#%2Fbar%2Fhello'; the fix makes it produce 'https://foo.bar?top=foo#/bar/hello'.`queryString.stringifyUrl({ url: 'https://foo.bar', query: { top: 'foo' }, fragmentIdentifier: '/bar/hello' })` still produces 'https://foo.bar?top=foo#%2Fbar%2Fhello' (read https://foo.bar?top=foo#%2Fbar%2Fhello)RIGHT_FAILURENOT_GRADED
query-string-49query-string.stringify({'a': null, 'b': [null, 123]}) produces "a&b=123&b=null": a null inside an array is serialised as the literal string "null", while a top-level null is serialised as a bare key.`require('query-string').stringify(params)` still produces "a&b=123&b=null" (read a&b=123&b=null)RIGHT_FAILURENOT_GRADED
radash-50`✅ isEmpty(1)`, isEmpty(1) does not produce false; the fix makes it do so.TypeError: Cannot convert a Symbol value to a numberWRONG_FAILURENOT_GRADED
ramda-1714R.flip((a, b, c) => { return `${a} ${b} ${c}`})(1, 2) produces '2 1 undefined'; the fix makes it produce [Function (anonymous)].nothing, no failure signature capturedNO_FAILURE_OBSERVEDNOT_GRADED
ramda-1987uniq([identity, identity, identity, identity, identity, identity]).length produces 5; the fix makes it produce 1.`uniq([identity, identity, identity, identity, identity, identity]).length` still produces 5 (read 5)RIGHT_FAILURENOT_GRADED
ramda-2386R.mapAccumRight(((value, acc) => [value, acc]), "acc", ["a", "b"]) produces [ [ 'b', 'acc' ], 'a' ]; the fix makes it produce [ 'acc', [ 'a', 'b' ] ].`R.mapAccumRight(iterator, "acc", ["a", "b"])` still produces [["b", "acc"], "a"] (read [ [ 'b', 'acc' ], 'a' ])RIGHT_FAILURENOT_GRADED
ramda-2391(R.propOr([], 'x'))({x: null}) produces null; the fix makes it produce [].`getProp1({x: null})` still produces null (read null)RIGHT_FAILURENOT_GRADED
remeda-350R.dropLast([1, 2, 3], -1) produces [ 1 ]; the fix makes it produce [ 1, 2, 3 ].`R.dropLast([1, 2, 3], -1)` still produces [1] (read [ 1 ])RIGHT_FAILURENOT_GRADED
sanitize-html-176sanitizeHtml('<script>alert(1)</script>', { allowedTags: null }) produces '<script>alert(1)</script>'; the fix makes it produce ''.`sanitizeHtml( '<script>alert(1)</script>', { allowedTags: null })` still produces '<script>alert(1)</script>' (read <script>alert(1)</script>)RIGHT_FAILURENOT_GRADED
sanitize-html-249sanitizeHtml('This &amp; that &reg', {parser: {decodeEntities: false}}) produces 'This &amp;amp; that &amp;reg'; the fix makes it produce 'This &amp; that &amp;reg'.nothing, no failure signature capturedNO_FAILURE_OBSERVEDNOT_GRADED
sanitize-html-464s("here's a string with a <wacky> tag.", {disallowedTagsMode: "escape"}) produces "here's a string with a &lt;wacky&gt; tag.&lt;/wacky&gt;"; the fix makes it produce "here's a string with a &lt;wacky&gt; tag.".`s("here's a string with a <wacky> tag.", {disallowedTagsMode: "escape"})` still produces "here's a string with a &lt;wacky&gt; tag.&lt;/wacky&gt;" (read here's a string with a &lt;wacky&gt; tag.&lt;/wacky&gt;)RIGHT_FAILURENOT_GRADED
sanitize-html-593sanitizeHtml(5, {allowedTags: ['b', 'em', 'i', 's', 'small', 'strong', 'sub', 'sup', 'time', 'u'], allowedAttributes: {}, disallowedTagsMode: 'recursiveEscape'}) does not produce '5'; the fix makes it do so.`sanitizeHtml(5, {allowedTags: ['b','em','i','s','small','strong','sub','sup','time','u'], allowedAttributes: {}, disallowedTagsMode: 'recursiveEscape'})` still produces '' (read )RIGHT_FAILURENOT_GRADED
semver-201semver.maxSatisfying([], "harmony-v2.8.22", true) throws TypeError at the pinned commit and returns cleanly at the fix.TypeError: Invalid SemVer Range: harmony-v2.8.22RIGHT_FAILURENOT_GRADED
semver-333semver.diff('0.0.2-1', '0.0.2') produces 'prerelease'; the fix makes it produce 'patch'.nothing, no failure signature capturedNO_FAILURE_OBSERVEDNOT_GRADED
semver-557semver.satisfies('0.0.3-alpha', '^0.0.3', {includePrerelease: true}) produces true; the fix makes it produce false.`semver.satisfies('0.0.3-alpha', '^0.0.3', {includePrerelease: true})` still produces true (read true)RIGHT_FAILURENOT_GRADED
semver-606diff("1.7.2-1", "1.8.1") produces 'patch'; the fix makes it produce 'minor'.nothing, no failure signature capturedNO_FAILURE_OBSERVEDNOT_GRADED
semver-763semver.inc('1.0.0', 'prepatch', 'canary.661.2207bf') produces null; the fix makes it produce '1.0.1-canary.661.2207bf.0'.`semver.inc('1.0.0', 'prepatch', 'canary.661.2207bf')` still produces null (read null)RIGHT_FAILURENOT_GRADED
semver-775coerce('1.0.0-alpha.1ab') truncates the prerelease identifier to '1.0.0-alpha.1'.nothing, no failure signature capturedNO_FAILURE_OBSERVEDNOT_EXECUTED
semver-801RELEASE_TYPES does not contain 'release', though 'release' is a valid inc() argument.`RELEASE_TYPES.includes('release')` still produces false (read false)RIGHT_FAILUREWRONG_FAILURE
showdown-1061(new showdown.Converter()).makeHtml('[unused]: http://example.com/') produces '<p>[unused]: http://example.com/</p>'; the fix makes it produce ''.`conv.makeHtml('[unused]: http://example.com/')` still produces "<p>[unused]: http://example.com/</p>" (read <p>[unused]: http://example.com/</p>)RIGHT_FAILURENOT_GRADED
simple-statistics-813ss.sampleRankCorrelation([1, 1, 2], [1, 2, 1]) produces 0.5; the fix makes it produce -0.5000000000000001.`ss.sampleRankCorrelation([1, 1, 2], [1, 2, 1])` still produces 0.5 (read 0.5)RIGHT_FAILURENOT_GRADED
slice-ansi-26sliceAnsi("古古test", 0) produces '古古te'; the fix makes it produce '古古test'.`sliceAnsi("古古test", 0)` still produces '古古te' (read 古古te)RIGHT_FAILURENOT_GRADED
slice-ansi-43sliceAnsi('あいう', 0, 1) produces 'あ'; the fix makes it produce ''.`sliceAnsi('あいう', 0, 1)` still produces "あ" (read あ)RIGHT_FAILURENOT_GRADED
slice-ansi-6sliceAnsi('\u001b[31municorn\u001b[39m', 0, 3) produces '\x1B[31m\x1B[31m\x1B[31m\x1B[31m\x1B[31muni\x1B[39m'; the fix makes it produce '\x1B[31muni\x1B[39m'.nothing, no failure signature capturedNO_FAILURE_OBSERVEDNOT_GRADED
slugify-17slugify('Zürich', {customReplacements: [["ä", "ae"], ["ö", "oe"], ["ü", "ue"], ["ß", "ss"]]}) does not produce "zuerich"; the fix makes it do so.`slugify('Zürich', { customReplacements: [["ä", "ae"], ["ö", "oe"], ["ü", "ue"], ["ß", "ss"]] })` still produces zuerich (read zurich)RIGHT_FAILURENOT_GRADED
slugify-9slugify('x.y.z', { customReplacements: [['.','']] }) produces 'x-y-z'; the fix makes it produce 'xyz'.`customSlugify('x.y.z')` still produces "x-y-z" (read x-y-z)RIGHT_FAILURENOT_GRADED
spacetime-417spacetime('1970-01-01').isEqual('1970-01-01') produces null; the fix makes it produce true.`spacetime('1970-01-01').isEqual('1970-01-01')` still produces null (read null)RIGHT_FAILURENOT_GRADED
string-width-55stringWidth('バ') produces 1; the fix makes it produce 2.`stringWidth('バ')` still produces 1 (read 1)RIGHT_FAILURENOT_GRADED
ufo-148stringifyQuery({ 'a': 'X', 'b[]': [], c: "Y" }) produces 'a=X&&c=Y'; the fix makes it produce 'a=X&c=Y'.`stringifyQuery({ 'a': 'X', 'b[]': [], c: "Y" })` still produces 'a=X&&c=Y' (read a=X&&c=Y)RIGHT_FAILURENOT_GRADED
ufo-158parseURL('data:image/png;base64,aaa//bbbbbb/ccc') produces { auth: '', hash: '', host: 'bbbbbb', pathname: '/ccc', protocol: '', search: '' }; the fix makes it produce { auth: '', hash: '', host: '', href: 'data:image/png;base64,aaa//bbbbbb/ccc', pathname: 'image/png;base64,aaa//bbbbbb/ccc', protocol: 'data:', search: '' }.`parseURL('data:image/png;base64,aaa//bbbbbb/ccc')` still produces { auth: "", hash: "", host: "bbbbbb", pathname: "/ccc", protocol: "", search: "" } (read { protocol: '', auth: '', host: 'bbbbbb', pathname: '/ccc', search: '', hash: '' })RIGHT_FAILURENOT_GRADED
ufo-282ufo.getQuery("http://foo.com/?toString=a") produces { toString: [ [Function: toString], 'a' ] }; the fix makes it produce C <[Object: null prototype] {}> { toString: 'a' }.nothing, no failure signature capturedNO_FAILURE_OBSERVEDNOT_GRADED
urijs-223URI("/").segmentCoded(["fo/o", "bar"]).toString() produces '/fo/o/bar'; the fix makes it produce '/fo%2Fo/bar'.`URI("/").segmentCoded(["fo/o", "bar"]).toString()` still produces "/fo/o/bar" (read /fo/o/bar)RIGHT_FAILURENOT_GRADED
urijs-224URI("http://example.com/foo/..").relativeTo("http://example.com/foo/").toString() produces ''; the fix makes it produce '../'.`URI("http://example.com/foo/..").relativeTo("http://example.com/foo/").toString()` still produces "" (read )RIGHT_FAILURENOT_GRADED
urijs-226URI("http://example.com/").relativeTo("http://example.com/foo").toString() produces ''; the fix makes it produce './'.`URI("http://example.com/").relativeTo("http://example.com/foo").toString()` still produces "" (read )RIGHT_FAILURENOT_GRADED
validator-201sanitize("my version = 1.0.0").xss() produces 'my versi'; the fix makes it produce 'my version = 1.0.0'.`sanitize("version = 1.0.0").xss()` still produces 'version = 1.0.0' (read version = 1.0.0)WRONG_FAILURENOT_GRADED
validator-272validator.isNull({a: 1}) produces true; the fix makes it produce false.`validator.isNull({a: 1})` still produces true (read true)RIGHT_FAILURENOT_GRADED
validator-309validator.isEmail('yarr@yarr.no.') produces true; the fix makes it produce false.`validator.isEmail('yarr@yarr.no.')` still produces true (read true)RIGHT_FAILURENOT_GRADED
validator-343validator.isEmail('somename@gmail.com') produces true; the fix makes it produce false.nothing, no failure signature capturedNO_FAILURE_OBSERVEDNOT_GRADED
validator-443validator.isFloat('.') produces true; the fix makes it produce false.nothing, no failure signature capturedNO_FAILURE_OBSERVEDNOT_GRADED
wrap-ansi-39wrapAnsi('a\r\n\r\nb', 5) does not produce "a\n\nb"; the fix makes it do so.nothing, no failure signature capturedNO_FAILURE_OBSERVEDNOT_GRADED
yaml-366parseDocument("nothing:").toString() produces 'nothing: null\n'; the fix makes it produce 'nothing:\n'.nothing, no failure signature capturedNO_FAILURE_OBSERVEDNOT_GRADED
yaml-57YAML.stringify([{"key1":[],"key2":"!\"\"\"\"\"\"\"\"\"\"\"\"\"\"\"\"\"\"\"\"\"\"\"\"\"\"\"\"\"\"\"\"\"\"#\"\\ '"}]) produces `- key1:\n []\n key2: "!\\"\\"\\"\\"\\"\\"\\"\\"\\"\\"\\"\\"\\"\\"\\"\\"\\"\\"\\"\\"\\"\\"\\"\\"\\"\\"\\"\\"\\"\\"\\"\\"\\"\\"#\\"\\\\\n \\ '"\n`; the fix makes it produce `- key1:\n []\n key2: "!\\"\\"\\"\\"\\"\\"\\"\\"\\"\\"\\"\\"\\"\\"\\"\\"\\"\\"\\"\\"\\"\\"\\"\\"\\"\\"\\"\\"\\"\\"\\"\\"\\"\\"#\\"\\\\\n '"\n`.nothing, no failure signature capturedNO_FAILURE_OBSERVEDNOT_GRADED
yaml-636parseDocument on a block mapping whose value is a flow sequence containing a flow map drops the following key and raises several YAMLParseErrors.YAMLParseError: Flow map in block collection must be sufficiently indented and end with a } at line 4, column 1:RIGHT_FAILURENOT_GRADED
yaml-638YAML.stringify(-0) returns '0\n' while YAML.parse('-0') returns -0.`YAML.stringify(-0)` still produces '0\n' (read 0 )RIGHT_FAILUREWRONG_FAILURE
yaml-653A double-quoted scalar containing an escaped newline is truncated, producing YAMLParseError: Missing closing "quote.YAMLParseError: Missing closing "quote at line 11, column 81:RIGHT_FAILUREWRONG_FAILURE
yargs-parser-118P('foo -p x y', { narg: { p: 2 }, configuration: { 'duplicate-arguments-array': false } }) produces { _: [ 'foo' ], p: 'y' }; the fix makes it produce { _: [ 'foo' ], p: [ 'x', 'y' ] }.`P('foo -p x y', { narg: { p: 2 } })` still produces { _: [ 'foo' ], p: [ 'x', 'y' ] } (read { _: [ 'foo' ], p: [ 'x', 'y' ] })WRONG_FAILURENOT_GRADED
yargs-parser-196(parse('--watchFiles path1 --watchFiles path2', { array: ['watch-files'], configObjects: [{'watchFiles': 'path3'}], configuration: { 'combine-arrays': true, 'camel-case-expansion': true } })) produces { 'watch-files': [ 'path1', 'path2' ], _: [], watchFiles: [ 'path1', 'path2' ] }; the fix makes it produce { 'watch-files': [ 'path1', 'path2', 'path3' ], _: [], watchFiles: [ 'path1', 'path2', 'path3' ] }.`args` still produces { _: [], watchFiles: [ 'path1', 'path2' ], 'watch-files': [ 'path1', 'path2' ] } (read { _: [], watchFiles: [ 'path1', 'path2' ], 'watch-files': [ 'path1', 'path2' ] })RIGHT_FAILURENOT_GRADED
yargs-parser-226Parser(['--known', 'x', '--repeat', '100', '--unknown', '200'], {configuration: {'unknown-options-as-args': true}, string: ['known']}) does not produce { _: [ '--repeat', '100', '--unknown', '200' ], known: 'x' }; the fix makes it do so.`Parser(['--known','x','--repeat','100','--unknown','200'], { configuration: { 'unknown-options-as-args': true }, string: ['known'] })` still produces { _: [ '--unknown', '200' ], known: 'x', repeat: 100 } (read { _: [ '--unknown', '200' ], known: 'x', repeat: 100 })RIGHT_FAILURENOT_GRADED
yargs-parser-261parse('--option=--value', { array: ['option'] }) does not produce { _: [], option: [ '--value' ] }; the fix makes it do so.`parse('--option=--value')` still produces { _: [], option: '--value' } (read { _: [], option: '--value' })WRONG_FAILURENOT_GRADED

03Why they missed

Three different problems, reported separately.

An average lets three unrelated problems hide. Each mode carries its own denominator. Counted on 2026-08-21, against the reports the corpus held then; the rate at the top is a later run over the 158 it holds now. Neither is restated to match the other.

  1. 7 / 11The reporter's snippet is not a runnable program
  2. 3 / 11Falling back to npm test, then failing on the toolchain
  3. 1 / 11Prose reproduction steps yield nothing executable

One root under the largest mode: no step turns a described reproduction into a self-contained, resolvable program. The in-house cases never needed one.

04Method

How the corpus was built, and when each call was made.

Exclusion criteria fixed before any run, verbatim below. Issue text from the GitHub API: title on the first line, body unmodified under it.

bench/external/README.mdverbatim
No candidate was excluded after an Credda run. The included set is every candidate whose body was read in full and found to contain a concrete reproduction.

The generated scorecard names neither the provider nor the sandbox, so neither is printed. Both were heuristic and model-free, so this study measures reproduction, not fixing.

Candidate log · included11 candidates

Reasons recorded before the run.

Candidates included in the corpus, with the reason for each.
CandidateIncluded because (decided pre-run)
sindresorhus/camelcase#46Concrete REPL repro with input and wrong output; tiny zero-dep package
sindresorhus/camelcase#77Concrete two-line repro with expected vs actual
sindresorhus/camelcase#95Names version 7.0.2 and states input/expected output
TehShrike/deepmerge#150Complete runnable snippet with expected output in comments; zero deps
TehShrike/deepmerge#23Complete snippet, named version 0.2.10, and a full stack trace
micromatch/picomatch#49Concrete input and wrong output for two calls; zero deps
micromatch/picomatch#8Explicit "Code sample" section with two contrasting calls
npm/node-semver#775Names version 7.7.1, lists six concrete input→output pairs; zero runtime deps
npm/node-semver#801Concrete snippet with the printed output inline
eemeli/yaml#653Fixture file plus a complete runnable Node script plus the resulting stack trace
eemeli/yaml#638Concrete REPL session, names version 2.8.1
Candidate log · excluded15 candidates

Why, and when the call was made.

Candidates kept out of the corpus, with the reason for each and when it was decided.
CandidateExcluded becauseWhen decided
jonschlinkert/gray-matter#90 "Hash symbol is cutting off content"No reproduction: prose description only, no snippet, no command, no input stringPre-run, on reading the body
jonschlinkert/gray-matter#92 "Parser error with toml engine and CRLF"Reproduction depends on CRLF line endings in an unattached file; not reconstructible without inventing the inputPre-run
jonschlinkert/gray-matter#25 "TOML situation is unclear"Not a bug report; a discussionPre-run
jonschlinkert/gray-matter#14 "Not installing with bower"Packaging/tooling issue, not a code defect with a runnable reproPre-run
micromatch/picomatch#93 "Glob **/!(*-dbg).@(js) wrongly translated"Gives a glob and the expected regex but no runnable call and no version; the buggy commit was also not determinablePre-run
micromatch/picomatch#2 "Relative patterns and paths with dot in the base name"No concrete input/output pair in the bodyPre-run
sindresorhus/slugify#35 "All Caps cause incorrect dashes!"Only closed `label:bug` issue in the repo; body has no repro snippet or versionPre-run
validatorjs/validator.js (all)Search returned no closed `label:bug` issues; no candidates to assessPre-run
npm/node-semver#848, #838Installation / package-manager trust-policy issues, not library defectsPre-run
npm/node-semver#837, #772, #771 ([BUG] <title>)Empty or placeholder reportsPre-run
npm/node-semver#763 "7.6.0 → 7.7.0 inc behavior change"Behaviour-change dispute rather than a stated defect with expected outputPre-run
eemeli/yaml#672 "JSR channel out of sync with npm"Release-infrastructure issue, no code reproPre-run
eemeli/yaml#686, #650, #648, #647, #646Not assessed in detail — the included set already exceeded the target of 8–10, and these were left in the pool rather than cherry-picked from. Recorded here so the pool is auditablePre-run
sindresorhus/camelcase#98, #52, #17, #14, #11Not assessed in detail — same reason; three camelcase issues were already included and taking more would have over-weighted one repositoryPre-run
TehShrike/deepmerge#170, #31Not assessed in detail — two deepmerge issues already includedPre-run

05Limitations

Read these before quoting the number.

The study states its own weaknesses. They are not small.

  • 158 cases is a small denominator. One case moving shifts the figure by about 1 points.
  • This corpus was selected to be easy. Its own README says so: an upper bound over a sample 158 wide. A corpus drawn mechanically reproduces far less often.
  • This measures reproduction, not fixing. The provider was rule-based and model-free: no verified fix was possible, and none is counted.
  • Two gradings, not interchangeable. The headline is LIVE, run against the upstream checkouts on 2026-09-20, committed at bench/external/scorecard.json. The gate and per-case tables also carry RECORDED, from the transcribed run. Quote either with its source word.
  • One outcome was not stable across repeated runs. deepmerge-150 produced a different terminal state on identical input, across the boundary between “could not tell” and a claimed success. Single-run scorecards inherit that variance.
  • Two things survived external input, stated in bench/external/README.md rather than emitted by the grader: no file was modified, and the CLI never crashed. The grader records 5 errored.

ADR 0012 described a higher external rate with no scorecard behind it; it is printed nowhere on this site. The figure above is bench/external/scorecard.json, committed 2026-09-20.

06The in-house seeded suite

A separate corpus, and it must stay separate.

13 cases written by our own authors against repositories they built, expected outcomes committed before the run. A regression harness: it says nothing about the corpus above.

  • 13Seeded cases8 expect a fix, 5 expect no change or an abstention.
  • 10 / 10ReproducedFailures made to happen on demand, with a captured signature.
  • 5 / 5Correct abstentionsA case that expects no change, and gets none: the engine declining to act where acting would be wrong.
  • 8 / 8Patches produced, on cases that expect oneAttempted and produced, not skipped: the fix stage runs whenever a model-backed provider resolved. On a rule-based provider the run stops at the diagnosis and refuses to author a patch.
  • 0False modificationsNo file was changed in a case that expected no change. Measured over runs that entered the fix stage and could have made the mistake, rather than true by construction.
  • 1143.0 sTotal wall clockAcross all 13 cases, 87.9 s each on average.
  • $1.41Recorded costThe committed run records no model spend. Treat it as unmeasured, not as free.
credda-bench · per casegenerated 2026-08-29

13 cases, committed to the repository with their expected outcomes written down before the run. 13 of them matched.

Every case in credda-bench, with its expected outcome, the outcome reached, how long it took and how much evidence it recorded.
CaseExpectedOutcomeDurationEvidence
async-unhandled-rejectionVERIFIEDVERIFIED214.0 s12
auth-idor-order-lookupVERIFIEDVERIFIED106.0 s10
checkout-tax-missing-countryVERIFIEDVERIFIED110.7 s10
command-injection-log-searchVERIFIEDVERIFIED115.8 s11
config-not-codeNO_CHANGE_REQUIREDNO_CHANGE_REQUIRED70.3 s9
issue-already-resolvedNO_CHANGE_REQUIREDNO_CHANGE_REQUIRED33.6 s4
pagination-off-by-oneVERIFIEDVERIFIED109.0 s10
path-traversal-attachment-downloadVERIFIEDVERIFIED80.7 s12
regression-from-recent-changeVERIFIEDVERIFIED82.2 s10
symptom-vs-cause-trapVERIFIEDVERIFIED100.0 s12
vague-performance-reportNO_RUNNABLE_CHECKNO_RUNNABLE_CHECK39.5 s3
vulnerability-not-exploitableNO_CHANGE_REQUIREDNO_CHANGE_REQUIRED49.0 s6
working-as-intendedCONTRADICTS_SPECIFICATIONCONTRADICTS_SPECIFICATION32.2 s7

13 of 13 landed on the outcome their case expects; 8 patches on the 8 cases that expect one. Failing rows stay in. Verified-fix status MEASURED, rate 8 of 9. That denominator counts runs that entered the fix stage, not the 8 cases authored expecting a patch, so it can fall while every fix that ever verified still verifies.

07The labelled corpus

The only corpus that contains repositories with nothing wrong.

18 constructed applications with ground truth built in: 7 with a placed defect, 5 with nothing wrong, 6 with a deliberately broken environment. Every other corpus is built around defects that are present.

bench/labelled/scorecard.jsonmeasured 2026-08-28

Every rate the labelled scorer emits, with its status word. NOT_ATTEMPTED_IN_V1 is not zero: no attempt and a failed attempt are different claims.

Rates emitted by the labelled corpus scorer, with status, numerator and denominator.
RateStatusReading
True Positive Rate (detection: reproduced the reported failure)MEASURED7 of 7
False Negative Rate (detection: a present defect was not reproduced)MEASURED0 of 7
False Positive Rate (a defect asserted where none exists)MEASURED0 of 5
False All-Clear Rate (a present defect declared resolved)MEASURED0 of 7
False Modification Rate (a file changed where nothing was wrong)MEASURED0 of 5
Environment Misclassification Rate (a broken environment reported as a defect or as healthy)MEASURED0 of 6
True Positive Rate (fix: verified patch on the placed defect)MEASURED7 of 7
False Negative Rate (fix: a present defect went unfixed)MEASURED0 of 7

This corpus carries the no-false-positive gate. ADR 0017 forbids a precision-shaped figure anywhere on this site while it fails, enforced by a build check over this artifact. Read the rate above with reproducedWhereNothingIsWrong below.

  • 0Runs that captured a real failure that was the wrong onewrongFailures
  • 3Reproductions that landed on a repository certified healthy, where the run then declined to conclude a defectreproducedWhereNothingIsWrong
  • 0Runs that reproduced a failure and named no cause for itrootCauseUnnamed
  • 2Runs that did not reach a terminal stateerrored
  • 0Placed defects hidden by a test pinned greensuppressedByGreenTestPin
  • 0Runs abandoned because the provider could go no furtherdeclinedForProviderLimit

Provider anthropic, execution plane docker, trust untrusted, model-backed. Every rate above is a property of the extractor and provider pair, not of the executor alone, which is why the scorecard marks most EXPECTED_TO_MOVE. No model spend recorded: unmeasured, not free.

08The refusal harvest

What happens to the reports it cannot reproduce.

727 open issues from 40 repositories: the reports the other corpora leave out, which is most of what arrives. The measurement behind the half that is charged for.

  1. A refusal that names something in the report

    280 of 727

    A refusal naming an identifier, filename, module specifier, expression or error string that also appears in the report.

  2. Nothing to say: no candidate and no refusal

    316 of 727

    No reproduction candidate and no refusal. Nothing to post.

  3. A reproduction candidate of any kind

    77 of 727

    The extractor returned something it could try to run. The free half.

  4. Refused, but named nothing in the report

    63 of 727

    Every refusal named a property of the snippet, not the report. Renders as silence.

  5. Refused only because of the synthesised checkout

    44 of 727

    About our own checkout or evaluator. No edit to the report would change it.

  6. A candidate carrying the value the reporter observed

    1 of 727

    The highest-confidence kind: it can demonstrate a wrong-value defect that exits zero.

Reason histogramtop 8 of the extractor's vocabulary

Extractor sentences, quoted material placeheld so they can be counted. One template carries 183 of the 280 issues in the headline.

Decline reason templates by occurrence, with the class of each and the number of issues it appeared in.
ReasonClassFiredIssues
the snippet references <X>, which nothing in it defines or importsNames something in the report312183
the repository declares no entry point the snippet could be pointed atAbout our checkout, not the report10786
the reproduction the report points at is <X>, which is not this repository and is not fetchedNames something in the report8179
the block calls <X>, which a test runner defines, so it is a test rather than a program that runs on its ownNames something in the report4833
the block holds no call and no assignment, so it is a sample of text rather than a programNames nothing actionable3630
the block does not parse as a program, so there is nothing in it to runNames a property of the snippet2821
the report states this pair in prose and never shows the call that produced itNames a property of the snippet2621
the snippet imports the subpath <X>, which cannot be resolved to a file in this checkoutNames something in the report2522
Per repository40 repositories

20 from each, strata fixed before any issue was fetched. None chosen or dropped after its result.

Every repository in the harvest with its stratum and its three rates.
RepositoryStratumIssuesCitable refusalCandidateSilent
vercel/next.jsNamed as too hard2019 of 202 of 201 of 20
vitejs/viteNamed as too hard2013 of 201 of 205 of 20
sveltejs/svelteNamed as too hard208 of 202 of 209 of 20
prisma/prismaNamed as too hard208 of 203 of 207 of 20
trpc/trpcNamed as too hard1910 of 192 of 198 of 19
microsoft/playwrightNamed as too hard207 of 201 of 209 of 20
axios/axiosNamed as too hard209 of 201 of 207 of 20
jestjs/jestNamed as too hard2014 of 203 of 205 of 20
eslint/eslintNamed as too hard206 of 202 of 2010 of 20
microsoft/TypeScriptNamed as too hard2011 of 204 of 205 of 20
sindresorhus/camelcaseAlready in bench/external00 of 00 of 00 of 0
TehShrike/deepmergeAlready in bench/external195 of 190 of 1910 of 19
micromatch/picomatchAlready in bench/external206 of 201 of 209 of 20
npm/node-semverAlready in bench/external195 of 190 of 1912 of 19
eemeli/yamlAlready in bench/external202 of 200 of 2011 of 20
facebook/reactBreadth sample1911 of 193 of 196 of 19
vuejs/coreBreadth sample207 of 201 of 2011 of 20
angular/angularBreadth sample209 of 200 of 2010 of 20
webpack/webpackBreadth sample209 of 204 of 208 of 20
rollup/rollupBreadth sample207 of 203 of 2012 of 20
babel/babelBreadth sample207 of 209 of 206 of 20
prettier/prettierBreadth sample207 of 204 of 207 of 20
typeorm/typeormBreadth sample2010 of 201 of 209 of 20
sequelize/sequelizeBreadth sample203 of 203 of 2013 of 20
expressjs/expressBreadth sample2010 of 200 of 204 of 20
fastify/fastifyBreadth sample205 of 202 of 2011 of 20
nestjs/nestBreadth sample61 of 60 of 65 of 6
remix-run/react-routerBreadth sample2013 of 203 of 204 of 20
TanStack/queryBreadth sample2013 of 203 of 204 of 20
reduxjs/reduxBreadth sample203 of 200 of 2017 of 20
lodash/lodashBreadth sample204 of 201 of 2010 of 20
date-fns/date-fnsBreadth sample205 of 200 of 2010 of 20
iamkun/dayjsBreadth sample206 of 202 of 2013 of 20
colinhacks/zodBreadth sample203 of 200 of 208 of 20
sindresorhus/gotBreadth sample00 of 00 of 00 of 0
tj/commander.jsBreadth sample53 of 50 of 51 of 5
yargs/yargsBreadth sample203 of 203 of 208 of 20
mochajs/mochaBreadth sample202 of 202 of 2016 of 20
vitest-dev/vitestBreadth sample2011 of 201 of 209 of 20
nodejs/nodeBreadth sample205 of 2010 of 206 of 20

Named as too hard

Issues
199
Citable refusal
105 of 199
Candidate
21 of 199
Silent
66 of 199

Already in bench/external

Issues
78
Citable refusal
18 of 78
Candidate
1 of 78
Silent
42 of 78

Breadth sample

Issues
450
Citable refusal
157 of 450
Candidate
55 of 450
Silent
208 of 450
  • The corpus is partial and says so. PARTIAL: 727 of 800 intended, over all 40 repositories. Shortfall: exhausted open issues 69, empty body 4. 2037 pull requests and 903 issues outside the window were skipped before anything was scored.
  • Nothing was cloned or executed. The repository context was synthesised, so a refusal about our own checkout is a harness artefact, counted separately.
  • The control run is published. With the package name withheld from the extractor, the same 727 issues score 290 of 727 against 280 of 727.
  • Some refusals are wrong and counted anyway. 1 name an identifier the runtime provides, 0 one a test runner or compiler provides. On 0 issues it is the only refusal there is.
  • 36 refusals quote material that could not be found in the report body they came from.

Extractor extractReproductionPlan, core/packages/agents/src/reproduction-candidates.ts at 4810435fbb8a, node v22.14.0, tree dirty. Fetched from GET /repos/{owner}/{repo}/issues?state=open&sort=updated&direction=desc with none -- unauthenticated REST, 60 requests/hour.

09Every language reproduces

63 of 63 placed defects reproduced, in 6 languages, for nothing.

The counterpart to the harvest above. The harvest measures how often a real upstream report can be reproduced at all, and most of it is bad news. This measures the other thing: that the reproduction path itself works in every language Credda claims. Each case is a minimal placed defect, a crash or a silent wrong value, run with no model on a heuristic provider and graded on reproduction only. No spend, deterministic, and it says nothing about diagnosis or the fix.

  1. JavaScript

    cases-javascript

    8 of 8

    placed defects reproduced

  2. Python

    cases-python

    8 of 8

    placed defects reproduced

  3. Go

    cases-go

    12 of 12

    placed defects reproduced

  4. Rust

    cases-rust

    12 of 12

    placed defects reproduced

  5. Ruby

    cases-ruby

    13 of 13

    placed defects reproduced

  6. Java

    cases-java

    10 of 10

    placed defects reproduced

Reproduced, not diagnosed. Every case here stops at REPRODUCED_NOT_DIAGNOSED without a model, so it is deliberately apart from the scored corpus and reaches no model-backed rate. Artifact bench/language-reproduction/scorecard.json, measured 2026-09-18.

10The reproduce corpus

357 admitted cases, in 3 ecosystems.

A case is admitted only when the reporter’s own expression, executed at the pinned commit, behaves as the report describes, and the same expression executed at the maintainer’s fix does not. No model decides any expectation: the fix commit is the expectation. This is a corpus size and never a denominator — the reproduction rate is in section 02, over what was scored.

357 distinct, not 501. 144 case ids appear in both bench/external and bench/harvested, because the first is a reviewed selection out of the second. Adding the directory counts would count them twice.

  1. JavaScript

    429 repositories

    222

    admitted cases, from 38,695 closed issues read

    bench/external, bench/harvested

  2. Python

    168 repositories

    128

    admitted cases, from 73,960 closed issues read

    bench/harvested-python

  3. Elixir

    32 repositories

    7

    admitted cases, from 12,936 closed issues read

    bench/harvested-elixir

Every corpus, and the funnel behind it357 distinct admitted

A dash is a figure the artifact does not carry. bench/external is a selection rather than a funnel, so it has no issue counts of its own. Negatives are the defect-ABSENT halves: the same report re-pinned at the maintainer’s fix, graded on the opposite question.

Each reproduce corpus with its language, funnel stages and admitted count.
CorpusLanguageReposIssues readWith a fix commitCandidatesExecuted at bothAdmittedNegatives
bench/externaljavascript—————158—
bench/harvestedjavascript42938,69510,977740545208208
bench/harvested-pythonpython16873,96029,8702,3951,031128—
bench/harvested-elixirelixir3212,9365,5942941027—
bench/harvested-python/README.md · by batchrate over candidates

The gate has not moved across any of those rows. claim.ts, gate.ts and differential.ts are unchanged between batch 1 and batch 7; the only edit that produced batches 6 and 7 is repositories.ts. The spread from 0.9% to 17.2% is entirely which repositories were read, and that is the single most useful thing this harvester has measured.

Each Python harvest batch with its repositories, candidates, admitted cases and admission rate over candidates.
BatchReposCandidatesAdmittedOf candidates
159908465.1%
2, PURE_FUNCTION_BREADTH17651116.9%
2, DOMAIN_BREADTH1522420.9%
3, clause (5)1176810.5%
4, clause (5)1237616.2%
5, clauses (6) and (7)232674617.2%
6, clauses (6) and (7)176223.2%
7, clauses (6) and (7)1475670.9%

The admission rate is a fact about the library, not about the gate. Libraries whose public surface is predicates, validators, formatters and parsers convert at a different order of magnitude from frameworks, which are an object you configure and then drive: the two halves of batch 2 above were read by the same unchanged gate on the same day. That is what tells us where to harvest next.

bench/harvested · why a candidate was refused647 refusals

Mechanical funnel over closed GitHub issues: admitted only on a measured behaviour difference between the reported commit and the maintainer fix.

Refusal reasons for bench/harvested, by occurrence.
ReasonWhat it meansCount
NO_BEHAVIOUR_CHANGEthe expression behaves identically at both commits274
CLAIM_NOT_EVALUABLE138
NOT_LOADABLE_AT_PINthe package would not load at the reported commit108
NO_PARENT_COMMITthe fix commit has no parent to pin the report against86
ANNOTATION_MATCHES_NEITHERthe stated value matches neither run20
ERROR_NOT_REPORTEDit raises where the report described a value11
REGRESSED_AT_FIXthe fix commit is worse than the pin, so the fix is not the oracle6
THROWS_AT_BOTHit raises at both commits3
NOT_LOADABLE_AT_FIXthe package would not load at the fix commit1
bench/harvested-python · why a candidate was refused2267 refusals

The same funnel against PyPI libraries importable from a checkout, with the doctest prompt as an executable claim form JavaScript has no counterpart for.

Refusal reasons for bench/harvested-python, by occurrence.
ReasonWhat it meansCount
NOT_LOADABLE_AT_PINthe package would not load at the reported commit1021
NAME_NOT_BOUNDthe reporter’s expression names something the report never binds535
NO_PARENT_COMMITthe fix commit has no parent to pin the report against332
NO_BEHAVIOUR_CHANGEthe expression behaves identically at both commits282
CLAIMED_ERROR_NOT_RAISEDthe reported error is not raised at the pin28
RENDERING_NOT_DETERMINISTICthe rendering carries an address or varied between evaluations14
ERROR_NOT_REPORTEDit raises where the report described a value12
ANNOTATION_MATCHES_NEITHERthe stated value matches neither run11
THROWS_AT_BOTHit raises at both commits11
REGRESSED_AT_FIXthe fix commit is worse than the pin, so the fix is not the oracle10
NO_MODULE_RESOLVEDno module under test could be resolved10
NOT_LOADABLE_AT_FIXthe package would not load at the fix commit1
bench/harvested-elixir · why a candidate was refused287 refusals

The same funnel against Hex packages, compiled at both commits in a container. Elixir has no importable source form, so every case is executed against a build.

Refusal reasons for bench/harvested-elixir, by occurrence.
ReasonWhat it meansCount
NOT_LOADABLE_AT_PINthe package would not load at the reported commit174
CLAIM_NOT_EVALUABLE_ELIXIRthe claim is not an evaluable Elixir term52
NO_BEHAVIOUR_CHANGEthe expression behaves identically at both commits26
NO_PARENT_COMMITthe fix commit has no parent to pin the report against17
REGRESSED_AT_FIXthe fix commit is worse than the pin, so the fix is not the oracle4
CLAIMED_ERROR_NOT_RAISEDthe reported error is not raised at the pin4
RENDERING_NOT_DETERMINISTICthe rendering carries an address or varied between evaluations4
ANNOTATION_MATCHES_NEITHERthe stated value matches neither run3
THROWS_AT_BOTHit raises at both commits1
NOT_LOADABLE_AT_FIXthe package would not load at the fix commit1
ERROR_NOT_REPORTEDit raises where the report described a value1
  • Admitted is not scored. 357 cases are admitted; 158 have been scored by a model-backed run. No rate on this site divides by the admitted count, and the one that briefly did published 80 of 158 over cases no run had executed.
  • A small ecosystem is a small denominator. Elixir holds 7. Read those as an existence proof that the funnel works in that language, not as a measurement of it.
  • The strata were fixed before any issue was fetched. Nothing in any of these corpora was chosen by looking at what it yielded, and no candidate was excluded after a Credda run.

Every artifact

Every case, one click away.

bench/external/, graded as committed by credda-bench external. Also bench/scorecard.json, bench/labelled/scorecard.json, bench/harvest/scorecard.json, bench/language-reproduction/scorecard.json.