Top three
High to lowLeaders by AA index
- Rank#1Claude Fable 5by Anthropic60Artificial Analysis intelligence index
- Rank#2GPT-5.6 Solby OpenAI59Artificial Analysis intelligence index
- Rank#3Kimi K3by Moonshot AI57Artificial Analysis intelligence index
Complete field
15 models · sort any evidence columnFull model ranking
| # | Model | Class | |||||
|---|---|---|---|---|---|---|---|
| 1 | Claude Fable 5Recommendedby Anthropic | Frontier | 60 | 64 | 66 | $2.75 | $10 · $50 |
| 2 | GPT-5.6 Solby OpenAI | Strong | 59 | 56 | 67 | $1.04 | $5 · $30 |
| 3 | Kimi K3Recommendedby Moonshot AI | Frontier | 57 | 61 | 61 | $0.95 | $3 · $15 |
| 4 | Claude Opus 4.8by Anthropic | Strong | 56 | 58 | 61 | $1.80 | $5 · $25 |
| 5 | GPT-5.6 Terraby OpenAI | Efficient | 55 | 52 | 62 | $0.82 | $2.5 · $15 |
| 6 | Grok 4.5RecommendedCore defaultby Grok | Strong | 54 | 59 | 64 | $0.31 | $2 · $6 |
| 7 | Claude Sonnet 5by Anthropic | Strong | 53 | 55 | — | $1.52 | $2 · $10 |
| 8 | GLM-5.2by Z.ai | Market comparison | 51.1 | — | — | $0.32 | $1.4 · $4.4 |
| 9 | Muse Spark 1.1Recommendedby Meta | Efficient | 51 | 54 | 54 | $0.26 | $1.25 · $4.25 |
| 10 | GPT-5.6 Lunaby OpenAI | Efficient | 51 | 50 | 59 | $0.21 | $1 · $6 |
| 11 | Gemini 3.6 Flashby Google | Market comparison | 50.1 | — | — | $0.50 | $1.5 · $7.5 |
| 12 | MiniMax-M3by MiniMax | Market comparison | 44.4 | — | — | $0.12 | $0.3 · $1.2 |
| 13 | DeepSeek V4 Proby DeepSeek | Market comparison | 44.3 | — | — | $0.04 | $0.435 · $0.87 |
| 14 | Nemotron 3 Ultraby NVIDIA | Market comparison | 37.8 | — | — | $0.24 | $0.675 · $2.675 |
| 15 | gpt-oss-120bby OpenAI | Market comparison | 23.8 | — | — | $0.06 | $0.15 · $0.6 |
AA index and task cost are published by Artificial Analysis. Core fit is core product judgment across agentic knowledge work, coding breadth, reliability, and cost-aware task completion. Coding scores are harness-specific; missing runs remain — rather than being treated as zero. Routed models use Core's bundled rates; market comparisons show the public benchmark rate. Completed calls settle from provider-reported usage. Charts and benchmark lenses live in Performance.