Tracked benchmark
SWE-bench Verified
Official source retained
Comparable rows
0
9 routed models awaiting a shared run
Leaderboard use
Source only
Not mixed into model scores
Evidence boundaryNo apples-to-oranges ranking

No catalog-wide comparable run for the current retained models. Individual submissions remain useful evidence, but they do not become a catalog ranking until the current models share a comparable agent harness and benchmark version.

Open SWE-bench Verified

Looking for the full sortable table? The Leaderboard ranks every routed model with public evidence and Core judgment kept separate.