AI Model Directory
Browse 34 AI models across 12 providers. Compare pricing, capabilities, and find the right model for your use case. Last updated: 2026-08-10.
OpenAI
API →GPT-5.6 Sol
GPT-5.6 Sol is OpenAI's flagship reasoning model, topping benchmarks for coding, mathematics, and multi-step agentic workflows. It supports native multimodal input including images and documents within its 256K context window.
GPT-5.6 Terra
GPT-5.6 Terra offers a balanced price-performance profile, delivering strong reasoning and coding at half the cost of Sol. Ideal for production workloads that need reliability without the flagship premium.
GPT-5.6 Luna
GPT-5.6 Luna is the fast, cost-effective model in the 5.6 family. It punches above its weight on everyday tasks like summarization, classification, and simple coding while keeping costs low.
GPT-5.5
GPT-5.5 is a dedicated reasoning model from OpenAI, excelling at deep analytical tasks that require extended chain-of-thought. It performs exceptionally on math, logic, and scientific problem-solving.
GPT-5.5 Pro
GPT-5.5 Pro is OpenAI's most powerful reasoning model, designed for the hardest problems in science, mathematics, and engineering. It uses extensive test-time compute to solve problems that stump other models.
GPT-5.4
GPT-5.4 was the flagship model before the 5.6 series, still delivering strong performance across a wide range of tasks. A reliable workhorse for applications that don't need the latest model.
GPT-5.4 Mini
GPT-5.4 Mini is the go-to model for everyday AI tasks at a fraction of the cost. It handles summarization, drafting, classification, and light coding with excellent speed and reliability.
GPT-5.4 Nano
GPT-5.4 Nano is OpenAI's most affordable model, designed for simple tasks that need blazing speed. Perfect for high-scale, cost-sensitive applications where milliseconds matter.
Anthropic
API →Fable 5
Fable 5 is Anthropic's next-generation frontier model, purpose-built for long-running autonomous agents. It excels at extended multi-step tasks requiring persistent context and strategic planning over hours of operation.
Opus 4.8
Opus 4.8 is Anthropic's premium model for complex agentic coding and enterprise work. It handles large-scale codebases, multi-file refactors, and sophisticated software engineering tasks with precision.
Sonnet 5
Sonnet 5 is Anthropic's high-performance workhorse for coding and agentic tasks. It delivers near-Opus quality at significantly lower cost, making it the practical choice for most production applications. Introductory pricing through August 2026.
Haiku 4.5
Haiku 4.5 is Anthropic's fastest and most cost-efficient model. It handles routine tasks like classification, extraction, and light drafting with near-instant response times at minimal cost.
Gemini 3.1 Pro
Gemini 3.1 Pro is Google's powerhouse for complex tasks, offering a massive 2 million token context window. It excels at processing entire codebases, long documents, and multi-hour video transcripts in a single prompt.
Gemini 3.5 Flash
Gemini 3.5 Flash is Google's latest frontier model for agentic coding and multi-step workflows. It delivers state-of-the-art performance on coding benchmarks while maintaining Flash-tier speed and cost efficiency.
Gemini 3.1 Flash-Lite
Gemini 3.1 Flash-Lite is Google's most cost-efficient model, designed for high-volume production workloads. It handles classification, extraction, and simple generation tasks at an extremely competitive price point.
DeepSeek
API →DeepSeek V4 Flash
DeepSeek V4 Flash offers frontier-level performance at an unprecedented price point. With cache-hit pricing as low as $0.0028 per 1M input tokens, it's one of the most cost-effective high-quality models available.
DeepSeek V4 Pro
DeepSeek V4 Pro is the flagship reasoning model from DeepSeek, delivering performance competitive with much more expensive models. It supports both thinking and non-thinking modes for flexible deployment.
Meta
API →Llama 4 Scout
Llama 4 Scout is Meta's open-source Mixture-of-Experts frontier model, offering strong reasoning and coding at a fraction of the cost of closed-source alternatives. Available via Groq with blazing fast inference.
Llama 3.3 70B
Llama 3.3 70B is Meta's refined 70B-parameter model that delivers strong all-around performance. It's a reliable workhorse for general-purpose tasks including writing, analysis, and moderate coding.
Llama 3.1 8B
Llama 3.1 8B is the most affordable model in the lineup, perfect for high-throughput, simple tasks. At just $0.05 per 1M input tokens, it's ideal when cost is the primary concern and tasks are straightforward.
Qwen (Alibaba)
API →Qwen 3.6 27B
Qwen 3.6 27B is Alibaba's latest 27B-parameter model, delivering surprisingly strong coding and reasoning for its size. Available on Groq, it offers a compelling mix of quality and speed at a moderate price point.
Qwen3 32B
Qwen3 32B offers excellent reasoning capabilities at a very competitive price. It's a strong choice for tasks requiring logical analysis and structured output, balancing quality and cost effectively.
Mistral
API →Mistral Small 4
Mistral Small 4 is a lightweight yet capable model optimized for efficiency and multilingual performance. It handles everyday tasks in multiple languages at a very competitive price, making it ideal for global deployments.
Codestral
Codestral is Mistral's specialized code generation model, fine-tuned for software development tasks. With a 256K context window, it handles large codebases and complex programming challenges at an affordable price.
xAI
API →Grok-3
Grok-3 is xAI's frontier reasoning model, delivering strong performance on analytical tasks, coding, and complex problem-solving. It brings competitive reasoning capabilities to the market with a unique personality.
Grok-3 Mini
Grok-3 Mini is a faster, lighter version of Grok-3 that retains much of its reasoning ability at a dramatically lower price. Perfect for production workloads that need speed without sacrificing too much quality.
Cohere
API →Command
Cohere Command is an enterprise-focused model optimized for agentic workflows, RAG applications, and multilingual tasks. It excels at tool use, structured output, and working with business data across languages.
Command R+
Command R+ is Cohere's previous-generation enterprise model, still offering strong performance for business applications. A reliable choice for existing deployments and applications that prioritize stability over cutting-edge.
Moonshot (Kimi)
API →Kimi K2.7 Code
Kimi K2.7 Code is Moonshot AI's strongest coding model, featuring native multimodal architecture supporting text, images, and video. It excels at programming tasks with deep thinking and long-context reasoning over 256K tokens.
Kimi K2.6
Kimi K2.6 is Moonshot AI's flagship multimodal model, combining strong language understanding with native image and video processing. It handles diverse inputs for comprehensive AI applications at a competitive price.
MiniMax
API →MiniMax M3
MiniMax M3 is the latest generation model from MiniMax with a permanent 50% discount on already competitive pricing. With up to 512K context, it's an exceptional value for long-form content and document processing.
MiniMax M2.7
MiniMax M2.7 is a fast, cost-efficient model well-suited for everyday tasks. It delivers reliable performance for classification, extraction, and content generation at a very accessible price point.
Xiaomi
API →MiMo V2.5 Pro
MiMo V2.5 Pro is Xiaomi's premium agentic model, a 310B-parameter MoE with 15B active parameters. It delivers frontier-level agentic capabilities with a 1M context window and deep multimodal understanding at a competitive credit-based price.
MiMo V2.5
MiMo V2.5 is Xiaomi's efficient agentic model offering the same 310B MoE architecture as Pro at half the credit rate. It provides excellent performance for coding, reasoning, and multimodal tasks with a generous 1M context window.
Benchmark Comparison
Independent LLM benchmarks from Artificial Analysis. Last updated: 2026-08-08.
| Model | Intelligence Index | Coding Index | Math Index | MMLU-Pro | GPQA Diamond | MATH-500 | AIME 2025 | LiveCodeBench | Humanity's Last Exam |
|---|---|---|---|---|---|---|---|---|---|
| Opus 4.8 · Anthropic | ★57.3 | ★74.3 | — | — | 92 | — | — | — | ★48.7 |
| GPT-5.4 · OpenAI | 53.1 | 71.1 | — | — | 92 | — | — | — | 43.7 |
| Gemini 3.5 Flash · Google | 52 | 70.1 | — | — | 92.2 | — | — | — | 42.7 |
| GPT-5.6 Sol · OpenAI | 50.7 | 69.7 | — | — | 89.8 | — | — | — | 39.4 |
| Gemini 3.1 Pro · Google | 47.7 | 68.8 | — | — | ★94.1 | — | — | — | 47 |
| GPT-5.6 Terra · OpenAI | 46.8 | 64.7 | — | — | 87.2 | — | — | — | 33.3 |
| MiniMax M3 · MiniMax | 45.4 | 58.6 | — | — | 92.9 | — | — | — | 39 |
| Sonnet 5 · Anthropic | 42.6 | 66.4 | — | — | 80 | — | — | — | 19 |
| DeepSeek V4 Flash · DeepSeek | 39 | 52 | — | — | 86.7 | — | — | — | 30.3 |
| MiMo V2.5 · Xiaomi | 38 | 56.8 | — | — | 84.9 | — | — | — | 27.2 |
| Kimi K2.6 · Moonshot (Kimi) | 35.4 | 61.8 | — | — | 78.8 | — | — | — | 19.6 |
| GPT-5.5 · OpenAI | 34.3 | 60.9 | — | — | 84.6 | — | — | — | 21.6 |
| GPT-5.6 Luna · OpenAI | 33.9 | 44.2 | — | — | 83.5 | — | — | — | 19.8 |
| DeepSeek V4 Pro · DeepSeek | 31.9 | 58.7 | — | — | 71.7 | — | — | — | 8.2 |
| GPT-5.4 Mini · OpenAI | 30.5 | 56.1 | — | — | 82.3 | — | — | — | 18.6 |
| MiniMax M2.7 · MiniMax | 28.9 | 52.6 | 78.3 | 82 | 77.7 | — | — | ★82.6 | 13.7 |
| MiMo V2.5 Pro · Xiaomi | 28.4 | 60.2 | — | — | 76.2 | — | — | — | 14.8 |
| Gemini 3.1 Flash-Lite · Google | 25.6 | 34.7 | — | — | 82.2 | — | — | — | 17.2 |
| Grok-3 Mini · xAI | 22.9 | — | ★84.7 | ★82.8 | 79.1 | ★99.2 | ★93.3 | 69.6 | 11 |
| Command · Cohere | 22.8 | 27.8 | 13 | 71.2 | 76.1 | 81.9 | 9.7 | 28.7 | 12 |
| Kimi K2.7 Code · Moonshot (Kimi) | 19.7 | 60.8 | 57 | 82.4 | 76.6 | 97.1 | 69.3 | 55.6 | 7.4 |
| GPT-5.4 Nano · OpenAI | 17.8 | 56.1 | — | — | 55.8 | — | — | — | 4.1 |
| Grok-3 · xAI | 15.2 | — | 58 | 79.9 | 69.3 | 87 | 33 | 42.5 | 4.1 |
| Mistral Small 4 · Mistral | 12.3 | 26.6 | — | — | 57.1 | — | — | — | 3.8 |
| Qwen3 32B · Qwen (Alibaba) | 11.4 | 15.3 | 73 | 79.8 | 66.8 | 96.1 | 80.7 | 54.6 | 7.4 |
| Llama 4 Scout · Meta | 10.3 | 8.2 | 14 | 75.2 | 58.7 | 84.4 | 28.3 | 29.9 | 3.8 |
| Llama 3.1 8B · Meta | 1.9 | — | — | 36.5 | 27 | 21.8 | 0 | 8.5 | 4.3 |
| Command R+ · Cohere | 1.7 | — | — | 33.8 | 28.4 | 16.4 | 0.7 | 4.8 | 4.8 |
★ = highest score for that benchmark. Source: Artificial Analysis. Scores are normalised to 0–100.