Evidencesimple-statistics-813
sampleRankCorrelation depends on row order when values are tied
simple-statistics#813, at commit 13530b2. A closed issue from a repository Credda did not choose.
LIVE2026-09-20
RIGHT_FAILUREexecuted against the upstream checkout.
- Outcome
- PATCH_REJECTED
- Wall time
- 163.8s
- Checks
- 5 passed of 5 applicable
RECORDED
NOT_GRADEDgraded from the transcript committed with this case.
- Outcome
- not recorded
- Checks
- none run
- Repository
- simple-statistics/simple-statistics
- Issue
- #813
- Pinned commit
- 13530b2557022647f8164d01ed9c9deeb14dbf1f
01The signal
The report, exactly as it was filed.
Nothing paraphrased or cleaned up. The mess is the thing under test.
sampleRankCorrelation depends on row order when values are tied
`sampleRankCorrelation` returns a different value for the same paired data depending on the order the pairs are listed in, whenever either input contains tied values. Reordering rows can change the magnitude and flip the sign.
**Version:** simple-statistics 7.9.3 (npm), Node v24.18.0, Windows 11.
### Reproduction
```js
import * as ss from "simple-statistics";
// The same three paired observations {(1,1), (1,2), (2,1)},
// with the first two rows swapped.
ss.sampleRankCorrelation([1, 1, 2], [1, 2, 1]); // => 0.5
ss.sampleRankCorrelation([1, 1, 2], [2, 1, 1]); // => -0.5
```
Three pairs is the minimum case. It gets worse with more data — over all 720 orderings of one 6-pair dataset it returns **16 distinct values**, from −0.257 to +0.657:
```js
const x = [1, 1, 2, 2, 3, 3];
const y = [1, 2, 2, 3, 1, 3];
// every permutation of the same 6 pairs:
// simple-statistics -> 16 distinct results, min -0.257, max +0.657
// Spearman with midranks -> one value, 0.25, for all 720
```
Two more cases that fall out of the same cause:
```js
// No monotonic association at all, by symmetry — every increase in one
// variable is matched by a decrease in the other.
ss.sampleRankCorrelation([1, 1, 2, 2], [1, 2, 1, 2]); // => 0.7999999999999999, expected 0
// y is constant, so its ranks have no variance and rho is undefined.
ss.sampleRankCorrelation([1, 2, 3, 4], [5, 5, 5, 5]); // => 1.0000000000000002, expected NaN
```
The last one seems the most likely to cause real harm: a constant column reports near-perfect monotonic association rather than signalling that the statistic does not apply.
### Cause
`src/sample_rank_correlation.js` lines 26–29 assign each value its position in the sorted array as its rank:
```js
for (let i = 0; i < xIndexes.length; i++) {
xRanks[xIndexes[i]] = i;
yRanks[yIndexes[i]] = i;
}
```
Tied values therefore get distinct consecutive ranks, and because `Array.prototype.sort` is stable the tiebreak is the value's original index. That makes the rank of a tied value a function of where it sits in the input array. Spearman's rho is defined on midranks, where tied values share the average of the ranks they span; with midranks the statistic depends only on the pairs.
### What I ruled out
- **Not floating point.** The differences are large and the sign changes; the control case with no ties anywhere behaves correctly and is order-invariant.
- **Not `sampleCorrelation`.** It receives the already-wrong rank vectors. Feeding it midranks computed independently gives the order-invariant answer.
- **Not previously caught by the suite**, which is why it has survived. There is no `test/sample_rank_correlation.test.js`; the only coverage is two cases inside `test/sample_correlation.test.js` (a monotonic `y = ±x²` case and the R-validated case), and both use strictly distinct values, so the tie path has no coverage at all.
Ties are the normal case for ordinal, Likert, binned, rounded, or low-cardinality data — which is much of what rank correlation is reached for in the first place.
I have a fix ready that ranks both inputs with midranks and adds the missing test file; opening it as a PR now. Results for tie-free inputs are unchanged and the existing R-validated case still passes.- Repository
- simple-statistics/simple-statistics
- Issue
- #813
- Commit
- 13530b2557022647f8164d01ed9c9deeb14dbf1f
- Why this commit
- The first parent of the fix commit 0779f22eebbbaeb4e6b4b37a477c27e4ce933cdc, which GitHub binds to this issue via CLOSED_EVENT_PR. Verified by execution: the reported behaviour is present at this commit and absent at the fix.
- How the text was obtained
- Fetched verbatim via the GitHub GraphQL API. Title on the first line, body unmodified below it. Nothing was paraphrased, cleaned up, or supplemented.
- Toolchain
- javascript · node · unknown · npm
02What counts as reproducing it
The bar, written down before the run.
- Symptom
- ss.sampleRankCorrelation([1, 1, 2], [1, 2, 1]) produces 0.5; the fix makes it produce -0.5000000000000001.
- Expression
- ss.sampleRankCorrelation([1, 1, 2], [1, 2, 1])
- Reported output
- 0.5
- Where that came from
- Read mechanically from the report's fenced code, SAME_LINE form: `ss.sampleRankCorrelation([1, 1, 2], [1, 2, 1]); // => 0.5`.
03What happened
The live run reproduced the reported failure.
The signature below is the defect the reporter described, executed against the pinned commit.
`ss.sampleRankCorrelation([1, 1, 2], [1, 2, 1])` still produces 0.5 (read 0.5)
The LIVE grading as emitted. A check that did not apply is never shown as a pass.
| Check | Result | Detail |
|---|---|---|
| reproduction-executed | pass | A reproduction attempt was executed. |
| signature-captured | pass | `ss.sampleRankCorrelation([1, 1, 2], [1, 2, 1])` still produces 0.5 (read 0.5) |
| right-failure | pass | Reproduced the reported failure: ss.sampleRankCorrelation([1, 1, 2], [1, 2, 1]) produces 0.5; the fix makes it produce -0.5000000000000001. |
| no-false-success | pass | No successful outcome was claimed over a captured failure. |
| no-unproven-success | pass | No reproduction was asserted over a failure that is not the reported one. |
bench/external/scorecard.json, the run of 2026-09-20 against all 158 upstream checkouts.
The same case, graded from the transcript recorded .
The grading the benchmark gate runs on. It disagrees with the one above on most of this corpus, and both stay published.
Check it yourself
Everything here is downstream of a public commit.
Clone it, check out 13530b2, run the report through the CLI the way the study did.
git clone https://github.com/simple-statistics/simple-statistics git checkout 13530b2557022647f8164d01ed9c9deeb14dbf1f npm install CREDDA_PROVIDER=heuristic \ npx tsx apps/cli/src/main.ts fix <repo-path> @<issue-file> --no-color