Measured fields are split6612 fields name a winner

Where measured quality diverges

Zero-based bars preserve magnitude; the exact raw values remain in the chart table and full comparison below.

Claude Fable 5Kimi K3
Exact values
MetricA · Claude Fable 5B · Kimi K3
Core fit6461
AA index6057
Coding index6661
Exact side-by-side comparison
Claude Fable 5by AnthropicKimi K3by Moonshot AI
Evaluation
Core fitHigher is better6461
AA Intelligence IndexHigher is better6057
ClassFrontierFrontier
Rated effortmaxmax
Coding agentsClaude Code vs Kimi Code CLI
Coding Agent IndexHigher is better6661
DeepSWEHigher is better6664
Terminal-Bench v2Higher is better8384
SWE-Atlas-QnAHigher is better4937
Coding $ / taskLower is better$11.71$3.18
Minutes / taskLower is better23.423.8
Cost
AA $ / taskLower is better$2.75$0.95
Input $ / 1MLower is better$10.00$3.00
Output $ / 1MLower is better$50.00$15.00
Cache read $ / 1MLower is better$1.00$0.30
RuntimeLands with the usage-telemetry phase
Latency (TTFT)Not yet measured
Throughput (tok/s)Not yet measured

Core fit is Core's dated product judgment; the AA index and task cost are published by Artificial Analysis; coding results name their harness. The better value in each measured row is emphasized — rows without two measurements name no winner.