Measured fields are split6–612 fields name a winner
Where measured quality diverges
Zero-based bars preserve magnitude; the exact raw values remain in the chart table and full comparison below.
| Claude Fable 5by Anthropic | Kimi K3by Moonshot AI | |
|---|---|---|
| Evaluation | ||
| Core fitHigher is better | 64 | 61 |
| AA Intelligence IndexHigher is better | 60 | 57 |
| Class | Frontier | Frontier |
| Rated effort | max | max |
| Coding agentsClaude Code vs Kimi Code CLI | ||
| Coding Agent IndexHigher is better | 66 | 61 |
| DeepSWEHigher is better | 66 | 64 |
| Terminal-Bench v2Higher is better | 83 | 84 |
| SWE-Atlas-QnAHigher is better | 49 | 37 |
| Coding $ / taskLower is better | $11.71 | $3.18 |
| Minutes / taskLower is better | 23.4 | 23.8 |
| Cost | ||
| AA $ / taskLower is better | $2.75 | $0.95 |
| Input $ / 1MLower is better | $10.00 | $3.00 |
| Output $ / 1MLower is better | $50.00 | $15.00 |
| Cache read $ / 1MLower is better | $1.00 | $0.30 |
| RuntimeLands with the usage-telemetry phase | ||
| Latency (TTFT)Not yet measured | — | — |
| Throughput (tok/s)Not yet measured | — | — |
Core fit is Core's dated product judgment; the AA index and task cost are published by Artificial Analysis; coding results name their harness. The better value in each measured row is emphasized — rows without two measurements name no winner.