Pick any two SWE-bench Verified submissions. The leaderboard shows them in rank order — but with only 500 tasks, most of those gaps are statistical noise. This checks, using the same per-instance results the leaderboard is built from.
Method. When two systems run the same 500 tasks, only the tasks where exactly one succeeds (the discordant pairs) carry information about which is better. The exact McNemar test asks: given b wins for A and c for B, is the split distinguishable from a coin flip? A gap below the benchmark's resolving power (~5–6 points here) can't be called real, no matter how confidently the leaderboard orders it.
Caveats (they make the leaderboard look better than it is). Per-pair p-values are unadjusted for multiple comparisons, and comparing an observed ranking involves selection effects. Both biases favor the leaderboard — so if anything, more gaps are noise than shown here.
Reproduce from public data: pip install evalstats · data from the SWE-bench/experiments repo · snapshot 2026-07-01.