Capability index

16 models across 9 providers. Seed scores today; replaced by real production eval results as the data moat grows.

ModelProviderAvg capabilityBest forContextSpeedCost / call*
Claude Fable 5Anthropic
96
Code / review, Hard reasoning / analysis300K40$0.040
Claude Opus 4.8Anthropic
95
Code / review, Writing / marketing200K45$0.020
GPT-5.5OpenAI
94
Hard reasoning / analysis, Summarize400K60$0.007
Kimi K3openMoonshot
94
Research / long context, Code / review1000K45$0.012
Claude Sonnet 5Anthropic
93
Code / review, Summarize200K90$0.012
GPT-5OpenAI
93
Hard reasoning / analysis, Classify / route400K60$0.007
Claude Sonnet 4.6Anthropic
92
Code / review, Summarize200K90$0.012
Gemini 2.5 ProGoogle
90
Research / long context, Summarize1000K70$0.007
Grok 4xAI
89
Hard reasoning / analysis, Classify / route256K80$0.012
Gemini 2.5 FlashGoogle
85
Research / long context, Classify / route1000K180$0.002
GPT-5 miniOpenAI
85
Classify / route, Extract structured data400K120$0.001
Claude Haiku 4.5Anthropic
85
Classify / route, Extract structured data200K160$0.004
DeepSeek V3openDeepSeek
85
Classify / route, Code / review128K90$0.001
Qwen 3 MaxopenAlibaba
85
Classify / route, Extract structured data256K120$0.001
Mistral LargeMistral
84
Classify / route, Extract structured data128K100$0.006
Llama 4 MaverickopenMeta
82
Classify / route, Extract structured data256K130$0.001

*Representative 1,500-in / 500-out call. Prices illustrative — refresh from provider price lists.