Tuesday, 25 August 2026 | Updating Daily AI insight, written for builders

LLM Leaderboard 2026 — AI Model Intelligence Index

In 2026, Claude Opus 4.8 tops our intelligence index at 55.7, just ahead of OpenAI’s GPT-5.5 (54.8). Google’s Gemini 3.5 Flash (50.2) and open-weight GLM-5.2 (51.1) sit close behind, with Gemini 3.1 Pro (46.5) and DeepSeek V4 Pro (44.3) rounding out the frontier. The single “best” model depends on how you weigh intelligence against price and speed.

The Intelligence column is a composite 0-100 score based on the Artificial Analysis Intelligence Index v4.1, which blends nine demanding evaluations spanning reasoning, coding, agentic tool use, and scientific knowledge. Higher is smarter — but a two-point gap at the very top rarely changes real-world output quality, so treat clusters of models as roughly equivalent rather than reading tiny differences as decisive.

Read intelligence alongside price and speed. A frontier model like Opus 4.8 costs far more per token than DeepSeek V4 Flash (40.3) or GLM-5.2, which deliver 70-90% of the intelligence at a fraction of the cost. For high-volume or latency-sensitive work, a cheaper, faster model usually wins; for the hardest reasoning and agentic tasks, the top of the table earns its premium. Sort by the metric that matches your use case.

#Modell EntwicklerIntelligenz Kontext Eingabe $/1 Mio. Ausgabe $/1 Mio. Open weights
1Claude Opus 5Anthropic611 Mio.$5.00$25.00Nein
2Claude Fable 5Anthropic601 Mio.$10.00$50.00Nein
3GPT-5.6 SolOpenAI591,05 Mio.$5.00$30.00Nein
4Kimi K3Moonshot AI571 Mio.$3.00$15.00Ja
5Claude Opus 4.8Anthropic55.71 Mio.$5.00$25.00Nein
6GPT-5.5OpenAI54.81,05 Mio.$5.00$30.00Nein
7GLM 5.2Zhipu AI51.11 Mio.$1.40$4.40Ja
8Gemini 3.5 FlashGoogle50.21 Mio.$1.50$9.00Nein
9Claude Sonnet 4.6Anthropic471 Mio.$3.00$15.00Nein
10Gemini 3.1 ProGoogle46.51,05 Mio.$2.00$12.00Nein
11DeepSeek V4-ProDeepSeek44.31 Mio.$0.44$0.87Ja
12Kimi K2.7 CodeMoonshot AI42256 K$0.60$2.50Ja
13DeepSeek V4-FlashDeepSeek40.31 Mio.$0.14$0.28Ja
14Claude Haiku 4.5Anthropic37200 K$1.00$5.00Nein
15DeepSeek R1DeepSeek20.1128 K$0.50$2.15Ja
16Mistral Large 3Mistral AI15.9256 K$2.00$6.00Ja
17Llama 4 MaverickMeta14.31 Mio.$0.15$0.60Ja
18Qwen3 235B-A22BAlibaba13128 K$0.45$1.80Ja
19Qwen3 32BAlibaba12128 K$0.08$0.28Ja
20Llama 4 ScoutMeta10.010 Mio.$0.10$0.30Ja
21Gemma 3 27BGoogle7.4128 K$0.08$0.16Ja
22Phi-4Microsoft4.916K$0.07$0.14Ja
23Claude Sonnet 5Anthropic1 Mio.$2.00$10.00Nein
24DeepSeek R1 Distill Llama 70BDeepSeek128 K$0.80$0.80Ja
25Gemini 2.5 ProGoogle (Google DeepMind)1 Mio. (1.048.576 Tokens)$1.25$10.00Nein
26Gemini 3.6 FlashGoogle1 Mio.$1.50$7.50Nein
27Gemma 3 12BGoogle128 K$0.05$0.15Ja
28Gemma 3 4BGoogle128 K$0.05$0.10Ja
29Grok 4xAI256.000 Tokens$3.00$15.00Nein
30Kling 2.5 Turbo ProKuaishouNein
31Llama 3.1 8BMeta128 K$0.02$0.03Ja
32Llama 3.3 70BMeta128 K$0.10$0.32Ja
33Mistral 7BMistral AI32K$0.02$0.03Ja
34Mistral NeMo 12BMistral AI128 K$0.02$0.04Ja
35NVIDIA Nemotron 3 Nano OmniNVIDIA256 KJa
36Qwen3 14BAlibaba128 K$0.12$0.24Ja
37Qwen3 30B-A3BAlibaba128 K$0.12$0.50Ja
38Qwen3 8BAlibaba128 K$0.04$0.14Ja
39Sora 2OpenAINein
40Sora 2 ProOpenAINein
41Veo 3.1GoogleNein
42Wan 2.5AlibabaNein

Click any column heading to re-sort. Intelligence = a 0–100 composite of public reasoning/knowledge benchmarks; — means not yet scored. Prices are USD per 1M tokens. Updated August 2026.

Häufig gestellte Fragen (FAQ)

What is the best LLM right now?

As of mid-2026, Claude Opus 4.8 is the highest-scoring model on the Artificial Analysis Intelligence Index (55.7), narrowly ahead of GPT-5.5 (54.8). Both lead on reasoning, coding, and agentic tasks, so the practical choice between them usually comes down to price, speed, and ecosystem rather than a meaningful intelligence gap.

What is the smartest open-source model?

GLM-5.2 is currently the most intelligent open-weight model, scoring 51.1 on the intelligence index — ahead of DeepSeek V4 Pro (44.3) and DeepSeek V4 Flash (40.3). That puts the best open models within roughly four points of proprietary frontier models like GPT-5.5, while remaining free to self-host or run through low-cost APIs.

Which AI model is the best value?

DeepSeek V4 Flash and GLM-5.2 offer the strongest intelligence-per-dollar. DeepSeek V4 Flash scores 40.3 at a tiny fraction of Opus 4.8’s price, delivering roughly 70% of the top score for a small percentage of the cost. For budget-sensitive, high-volume workloads, these open models are the clear value leaders.

How is the intelligence score calculated?

Our scores mirror the Artificial Analysis Intelligence Index v4.1, a 0-100 composite of nine demanding evaluations — including Humanity’s Last Exam, GPQA Diamond, Terminal-Bench, SciCode, and agentic tool-use tasks. It measures reasoning, knowledge, coding, and agentic ability in one number, so a higher score means broadly more capable across hard, real-world tasks.

Scroll to Top
Featured on There's An AI For That