Model profiles
15
AA coverage
15/15
Coding coverage
8/15
Evidence snapshot
Top three

Leaders by AA index

High to low
  1. Rank#1
    Claude Fable 5by Anthropic
    60Artificial Analysis intelligence index
  2. Rank#2
    GPT-5.6 Solby OpenAI
    59Artificial Analysis intelligence index
  3. Rank#3
    Kimi K3by Moonshot AI
    57Artificial Analysis intelligence index
Complete field

Full model ranking

15 models · sort any evidence column
#ModelClass
1
Claude Fable 5Recommendedby Anthropic
Frontier606466$2.75$10 · $50
2
GPT-5.6 Solby OpenAI
Strong595667$1.04$5 · $30
3
Kimi K3Recommendedby Moonshot AI
Frontier576161$0.95$3 · $15
4
Claude Opus 4.8by Anthropic
Strong565861$1.80$5 · $25
5
GPT-5.6 Terraby OpenAI
Efficient555262$0.82$2.5 · $15
6
Grok 4.5RecommendedCore defaultby Grok
Strong545964$0.31$2 · $6
7
Claude Sonnet 5by Anthropic
Strong5355$1.52$2 · $10
8
GLM-5.2by Z.ai
Market comparison51.1$0.32$1.4 · $4.4
9
Muse Spark 1.1Recommendedby Meta
Efficient515454$0.26$1.25 · $4.25
10
GPT-5.6 Lunaby OpenAI
Efficient515059$0.21$1 · $6
11Market comparison50.1$0.50$1.5 · $7.5
12
MiniMax-M3by MiniMax
Market comparison44.4$0.12$0.3 · $1.2
13
DeepSeek V4 Proby DeepSeek
Market comparison44.3$0.04$0.435 · $0.87
14Market comparison37.8$0.24$0.675 · $2.675
15
gpt-oss-120bby OpenAI
Market comparison23.8$0.06$0.15 · $0.6

AA index and task cost are published by Artificial Analysis. Core fit is core product judgment across agentic knowledge work, coding breadth, reliability, and cost-aware task completion. Coding scores are harness-specific; missing runs remain — rather than being treated as zero. Routed models use Core's bundled rates; market comparisons show the public benchmark rate. Completed calls settle from provider-reported usage. Charts and benchmark lenses live in Performance.