Tuesday, 25 August 2026 | Updating Daily AI insight, written for builders

LLM Leaderboard 2026 — AI Model Intelligence Index

In 2026, Claude Opus 4.8 tops our intelligence index at 55.7, just ahead of OpenAI’s GPT-5.5 (54.8). Google’s Gemini 3.5 Flash (50.2) and open-weight GLM-5.2 (51.1) sit close behind, with Gemini 3.1 Pro (46.5) and DeepSeek V4 Pro (44.3) rounding out the frontier. The single “best” model depends on how you weigh intelligence against price and speed.

The Intelligence column is a composite 0-100 score based on the Artificial Analysis Intelligence Index v4.1, which blends nine demanding evaluations spanning reasoning, coding, agentic tool use, and scientific knowledge. Higher is smarter — but a two-point gap at the very top rarely changes real-world output quality, so treat clusters of models as roughly equivalent rather than reading tiny differences as decisive.

Read intelligence alongside price and speed. A frontier model like Opus 4.8 costs far more per token than DeepSeek V4 Flash (40.3) or GLM-5.2, which deliver 70-90% of the intelligence at a fraction of the cost. For high-volume or latency-sensitive work, a cheaper, faster model usually wins; for the hardest reasoning and agentic tasks, the top of the table earns its premium. Sort by the metric that matches your use case.

#Modelo DesarrolladorInteligencia Contexto Entrada: $/1 millón Salida: $/1 millón Open weights
1Claude Opus 5Anthropic611 millón$5.00$25.00No
2Claude Fable 5Anthropic601 millón$10.00$50.00No
3GPT-5.6 SolOpenAI591,05 millones$5.00$30.00No
4Kimi K3Moonshot AI571 millón$3.00$15.00
5Claude Opus 4.8Anthropic55.71 millón$5.00$25.00No
6GPT-5.5OpenAI54.81,05 millones$5.00$30.00No
7GLM 5.2Zhipu AI51.11 millón$1.40$4.40
8Gemini 3.5 FlashGoogle50.21 millón$1.50$9.00No
9Claude Sonnet 4.6Anthropic471 millón$3.00$15.00No
10Gemini 3.1 ProGoogle46.51,05 millones$2.00$12.00No
11DeepSeek V4-ProDeepSeek44.31 millón$0.44$0.87
12Kimi K2.7 CodeMoonshot AI42256K$0.60$2.50
13DeepSeek V4-FlashDeepSeek40.31 millón$0.14$0.28
14Claude Haiku 4.5Anthropic37200 000$1.00$5.00No
15DeepSeek R1DeepSeek20.1128 K$0.50$2.15
16Mistral Large 3Mistral AI15.9256K$2.00$6.00
17Llama 4 MaverickMeta14.31 millón$0.15$0.60
18Qwen3 235B-A22BAlibaba13128 K$0.45$1.80
19Qwen3 32BAlibaba12128 K$0.08$0.28
20Llama 4 ScoutMeta10.010M$0.10$0.30
21Gemma 3 27BGoogle7.4128 K$0.08$0.16
22Phi-4Microsoft4.916K$0.07$0.14
23Claude Sonnet 5Anthropic1 millón$2.00$10.00No
24DeepSeek R1 Distill Llama 70BDeepSeek128 K$0.80$0.80
25Gemini 2.5 ProGoogle (Google DeepMind)1 millón (1 048 576 tokens)$1.25$10.00No
26Gemini 3.6 FlashGoogle1 millón$1.50$7.50No
27Gemma 3 12BGoogle128 K$0.05$0.15
28Gemma 3 4BGoogle128 K$0.05$0.10
29Grok 4xAI256 000 tokens$3.00$15.00No
30Kling 2.5 Turbo ProKuaishouNo
31Llama 3.1 8BMeta128 K$0.02$0.03
32Llama 3.3 70BMeta128 K$0.10$0.32
33Mistral 7BMistral AI32K$0.02$0.03
34Mistral NeMo 12BMistral AI128 K$0.02$0.04
35NVIDIA Nemotron 3 Nano OmniNVIDIA256K
36Qwen3 14BAlibaba128 K$0.12$0.24
37Qwen3 30B-A3BAlibaba128 K$0.12$0.50
38Qwen3 8BAlibaba128 K$0.04$0.14
39Sora 2OpenAINo
40Sora 2 ProOpenAINo
41Veo 3.1GoogleNo
42Wan 2.5AlibabaNo

Click any column heading to re-sort. Intelligence = a 0–100 composite of public reasoning/knowledge benchmarks; — means not yet scored. Prices are USD per 1M tokens. Updated August 2026.

Preguntas frecuentes

What is the best LLM right now?

As of mid-2026, Claude Opus 4.8 is the highest-scoring model on the Artificial Analysis Intelligence Index (55.7), narrowly ahead of GPT-5.5 (54.8). Both lead on reasoning, coding, and agentic tasks, so the practical choice between them usually comes down to price, speed, and ecosystem rather than a meaningful intelligence gap.

What is the smartest open-source model?

GLM-5.2 is currently the most intelligent open-weight model, scoring 51.1 on the intelligence index — ahead of DeepSeek V4 Pro (44.3) and DeepSeek V4 Flash (40.3). That puts the best open models within roughly four points of proprietary frontier models like GPT-5.5, while remaining free to self-host or run through low-cost APIs.

Which AI model is the best value?

DeepSeek V4 Flash and GLM-5.2 offer the strongest intelligence-per-dollar. DeepSeek V4 Flash scores 40.3 at a tiny fraction of Opus 4.8’s price, delivering roughly 70% of the top score for a small percentage of the cost. For budget-sensitive, high-volume workloads, these open models are the clear value leaders.

How is the intelligence score calculated?

Our scores mirror the Artificial Analysis Intelligence Index v4.1, a 0-100 composite of nine demanding evaluations — including Humanity’s Last Exam, GPQA Diamond, Terminal-Bench, SciCode, and agentic tool-use tasks. It measures reasoning, knowledge, coding, and agentic ability in one number, so a higher score means broadly more capable across hard, real-world tasks.

Scroll to Top
Featured on There's An AI For That