- Tracked benchmark
- SWE-bench Verified Official source retained
- Comparable rows
- 0 9 routed models awaiting a shared run
- Leaderboard use
- Source only Not mixed into model scores
Evidence boundaryNo apples-to-oranges ranking
No catalog-wide comparable run for the current retained models. Individual submissions remain useful evidence, but they do not become a catalog ranking until the current models share a comparable agent harness and benchmark version.
Open SWE-bench Verified ↗Looking for the full sortable table? The Leaderboard ranks every routed model with public evidence and Core judgment kept separate.