Is that SWE-bench gap real?

Pick any two SWE-bench Verified submissions. The leaderboard shows them in rank order — but with only 500 tasks, most of those gaps are statistical noise. This checks, using the same per-instance results the leaderboard is built from.

gap discordant pairs McNemar p

Method. When two systems run the same 500 tasks, only the tasks where exactly one succeeds (the discordant pairs) carry information about which is better. The exact McNemar test asks: given b wins for A and c for B, is the split distinguishable from a coin flip? A gap below the benchmark's resolving power (~5–6 points here) can't be called real, no matter how confidently the leaderboard orders it.

Caveats (they make the leaderboard look better than it is). Per-pair p-values are unadjusted for multiple comparisons, and comparing an observed ranking involves selection effects. Both biases favor the leaderboard — so if anything, more gaps are noise than shown here.

Reproduce from public data: pip install evalstats  ·  data from the SWE-bench/experiments repo  ·  snapshot 2026-07-01.