Wednesday, 26 August 2026 | Updating Daily AI insight, written for builders

LLM Leaderboard 2026 — AI Model Intelligence Index

In 2026, Claude Opus 4.8 tops our intelligence index at 55.7, just ahead of OpenAI’s GPT-5.5 (54.8). Google’s Gemini 3.5 Flash (50.2) and open-weight GLM-5.2 (51.1) sit close behind, with Gemini 3.1 Pro (46.5) and DeepSeek V4 Pro (44.3) rounding out the frontier. The single “best” model depends on how you weigh intelligence against price and speed.

The Intelligence column is a composite 0-100 score based on the Artificial Analysis Intelligence Index v4.1, which blends nine demanding evaluations spanning reasoning, coding, agentic tool use, and scientific knowledge. Higher is smarter — but a two-point gap at the very top rarely changes real-world output quality, so treat clusters of models as roughly equivalent rather than reading tiny differences as decisive.

Read intelligence alongside price and speed. A frontier model like Opus 4.8 costs far more per token than DeepSeek V4 Flash (40.3) or GLM-5.2, which deliver 70-90% of the intelligence at a fraction of the cost. For high-volume or latency-sensitive work, a cheaper, faster model usually wins; for the hardest reasoning and agentic tasks, the top of the table earns its premium. Sort by the metric that matches your use case.

# Model Developer Intelligence Context Input $/1M Output $/1M Open weights
1 Claude Opus 5 Anthropic 61 1M $5.00 $25.00 No
2 Claude Fable 5 Anthropic 60 1M $10.00 $50.00 No
3 GPT-5.6 Sol OpenAI 59 1.05M $5.00 $30.00 No
4 Kimi K3 Moonshot AI 57 1M $3.00 $15.00 Yes
5 Claude Opus 4.8 Anthropic 55.7 1M $5.00 $25.00 No
6 GPT-5.5 OpenAI 54.8 1.05M $5.00 $30.00 No
7 GLM 5.2 Zhipu AI 51.1 1M $1.40 $4.40 Yes
8 Gemini 3.5 Flash Google 50.2 1M $1.50 $9.00 No
9 Claude Sonnet 4.6 Anthropic 47 1M $3.00 $15.00 No
10 Gemini 3.1 Pro Google 46.5 1.05M $2.00 $12.00 No
11 DeepSeek V4-Pro DeepSeek 44.3 1M $0.44 $0.87 Yes
12 Kimi K2.7 Code Moonshot AI 42 256K $0.60 $2.50 Yes
13 DeepSeek V4-Flash DeepSeek 40.3 1M $0.14 $0.28 Yes
14 Claude Haiku 4.5 Anthropic 37 200K $1.00 $5.00 No
15 DeepSeek R1 DeepSeek 20.1 128K $0.50 $2.15 Yes
16 Mistral Large 3 Mistral AI 15.9 256K $2.00 $6.00 Yes
17 Llama 4 Maverick Meta 14.3 1M $0.20 $0.80 Yes
18 Qwen3 235B-A22B Alibaba 13 128K $0.45 $1.80 Yes
19 Qwen3 32B Alibaba 12 128K $0.08 $0.28 Yes
20 Llama 4 Scout Meta 10.0 10M $0.10 $0.30 Yes
21 Gemma 3 27B Google 7.4 128K $0.08 $0.16 Yes
22 Phi-4 Microsoft 4.9 16K $0.07 $0.14 Yes
23 Claude Sonnet 5 Anthropic 1M $2.00 $10.00 No
24 DeepSeek R1 Distill Llama 70B DeepSeek 128K $0.80 $0.80 Yes
25 Gemini 2.5 Pro Google (Google DeepMind) 1M (1,048,576 tokens) $1.25 $10.00 No
26 Gemini 3.6 Flash Google 1M $1.50 $7.50 No
27 Gemma 3 12B Google 128K $0.05 $0.15 Yes
28 Gemma 3 4B Google 128K $0.05 $0.10 Yes
29 Grok 4 xAI 256,000 tokens $3.00 $15.00 No
30 Kling 2.5 Turbo Pro Kuaishou No
31 Llama 3.1 8B Meta 128K $0.02 $0.03 Yes
32 Llama 3.3 70B Meta 128K $0.10 $0.32 Yes
33 Mistral 7B Mistral AI 32K $0.02 $0.03 Yes
34 Mistral NeMo 12B Mistral AI 128K $0.02 $0.04 Yes
35 NVIDIA Nemotron 3 Nano Omni NVIDIA 256K Yes
36 Qwen3 14B Alibaba 128K $0.12 $0.24 Yes
37 Qwen3 30B-A3B Alibaba 128K $0.12 $0.50 Yes
38 Qwen3 8B Alibaba 128K $0.04 $0.14 Yes
39 Sora 2 OpenAI No
40 Sora 2 Pro OpenAI No
41 Veo 3.1 Google No
42 Wan 2.5 Alibaba No

Click any column heading to re-sort. Intelligence = a 0–100 composite of public reasoning/knowledge benchmarks; — means not yet scored. Prices are USD per 1M tokens. Updated August 2026.

FAQ

What is the best LLM right now?

As of mid-2026, Claude Opus 4.8 is the highest-scoring model on the Artificial Analysis Intelligence Index (55.7), narrowly ahead of GPT-5.5 (54.8). Both lead on reasoning, coding, and agentic tasks, so the practical choice between them usually comes down to price, speed, and ecosystem rather than a meaningful intelligence gap.

What is the smartest open-source model?

GLM-5.2 is currently the most intelligent open-weight model, scoring 51.1 on the intelligence index — ahead of DeepSeek V4 Pro (44.3) and DeepSeek V4 Flash (40.3). That puts the best open models within roughly four points of proprietary frontier models like GPT-5.5, while remaining free to self-host or run through low-cost APIs.

Which AI model is the best value?

DeepSeek V4 Flash and GLM-5.2 offer the strongest intelligence-per-dollar. DeepSeek V4 Flash scores 40.3 at a tiny fraction of Opus 4.8’s price, delivering roughly 70% of the top score for a small percentage of the cost. For budget-sensitive, high-volume workloads, these open models are the clear value leaders.

How is the intelligence score calculated?

Our scores mirror the Artificial Analysis Intelligence Index v4.1, a 0-100 composite of nine demanding evaluations — including Humanity’s Last Exam, GPQA Diamond, Terminal-Bench, SciCode, and agentic tool-use tasks. It measures reasoning, knowledge, coding, and agentic ability in one number, so a higher score means broadly more capable across hard, real-world tasks.

Scroll to Top
Featured on There's An AI For That