| Model + harness | Index | DeepSWE | Terminal-Bench v2 | SWE-Atlas | $ / task | Time |
|---|---|---|---|---|---|---|
| GPT-5.6 SolCodex · max | $7.08 | 10.2m | ||||
| Claude Fable 5Claude Code · max | $11.71 | 23.4m | ||||
| Grok 4.5Grok Build · high | $2.59 | 16.5m | ||||
| GPT-5.6 TerraCodex · max | $2.76 | 8.4m | ||||
| Kimi K3Kimi Code CLI · max | $3.18 | 23.8m | ||||
| Claude Opus 4.8Claude Code · max | $7.70 | 23.1m | ||||
| GPT-5.6 LunaCodex · max | $1.57 | 8.0m | ||||
| Muse Spark 1.1OpenCode · xhigh | $1.43 | 12.6m |
Equal-weight composite of DeepSWE, Terminal-Bench v2, and SWE-Atlas-QnA; every result is tied to its evaluated harness.
Looking for the full sortable table? The Leaderboard ranks every routed model with public evidence and Core judgment kept separate.