Tuesday, 25 August 2026 | التحديث اليومي نظرة ثاقبة للذكاء الاصطناعي، مكتوبة للبناة

LLM Leaderboard 2026 — AI Model Intelligence Index

In 2026, Claude Opus 4.8 tops our intelligence index at 55.7, just ahead of OpenAI’s GPT-5.5 (54.8). Google’s Gemini 3.5 Flash (50.2) and open-weight GLM-5.2 (51.1) sit close behind, with Gemini 3.1 Pro (46.5) and DeepSeek V4 Pro (44.3) rounding out the frontier. The single “best” model depends on how you weigh intelligence against price and speed.

The Intelligence column is a composite 0-100 score based on the Artificial Analysis Intelligence Index v4.1, which blends nine demanding evaluations spanning reasoning, coding, agentic tool use, and scientific knowledge. Higher is smarter — but a two-point gap at the very top rarely changes real-world output quality, so treat clusters of models as roughly equivalent rather than reading tiny differences as decisive.

Read intelligence alongside price and speed. A frontier model like Opus 4.8 costs far more per token than DeepSeek V4 Flash (40.3) or GLM-5.2, which deliver 70-90% of the intelligence at a fraction of the cost. For high-volume or latency-sensitive work, a cheaper, faster model usually wins; for the hardest reasoning and agentic tasks, the top of the table earns its premium. Sort by the metric that matches your use case.

#النموذج المطوّرالذكاء السياق التكلفة لكل مليون رمز (مدخلات) التكلفة لكل مليون رمز (مخرجات) Open weights
1Claude Opus 5أنثروبيك61مليون$5.00$25.00لا
2Claude Fable 5أنثروبيك60مليون$10.00$50.00لا
3GPT-5.6 SolOpenAI591.05 مليون رمز$5.00$30.00لا
4Kimi K3مون شوت آي57مليون$3.00$15.00نعم
5Claude Opus 4.8أنثروبيك55.7مليون$5.00$25.00لا
6GPT-5.5OpenAI54.81.05 مليون رمز$5.00$30.00لا
7GLM 5.2زهي بو آي51.1مليون$1.40$4.40نعم
8Gemini 3.5 Flashجوجل50.2مليون$1.50$9.00لا
9Claude Sonnet 4.6أنثروبيك47مليون$3.00$15.00لا
10Gemini 3.1 Proجوجل46.51.05 مليون رمز$2.00$12.00لا
11DeepSeek V4-ProDeepSeek44.3مليون$0.44$0.87نعم
12Kimi K2.7 Codeمون شوت آي42256 ألف رمز$0.60$2.50نعم
13DeepSeek V4-FlashDeepSeek40.3مليون$0.14$0.28نعم
14Claude Haiku 4.5أنثروبيك37200 ألف$1.00$5.00لا
15DeepSeek R1DeepSeek20.1128 ألف رمز$0.50$2.15نعم
16Mistral Large 3ميسترال إيه آي15.9256 ألف رمز$2.00$6.00نعم
17Llama 4 Maverickميتا14.3مليون$0.15$0.60نعم
18Qwen3 235B-A22Bعلي بابا13128 ألف رمز$0.45$1.80نعم
19Qwen3 32Bعلي بابا12128 ألف رمز$0.08$0.28نعم
20Llama 4 Scoutميتا10.010 ملايين$0.10$0.30نعم
21Gemma 3 27Bجوجل7.4128 ألف رمز$0.08$0.16نعم
22Phi-4مايكروسوفت4.916 ألف رمز$0.07$0.14نعم
23Claude Sonnet 5أنثروبيكمليون$2.00$10.00لا
24DeepSeek R1 Distill Llama 70BDeepSeek128 ألف رمز$0.80$0.80نعم
25Gemini 2.5 Proجوجل (جوجل ديب مايند)مليون رمز (1,048,576 رمزًا)$1.25$10.00لا
26Gemini 3.6 Flashجوجلمليون$1.50$7.50لا
27Gemma 3 12Bجوجل128 ألف رمز$0.05$0.15نعم
28Gemma 3 4Bجوجل128 ألف رمز$0.05$0.10نعم
29Grok 4إكس إيه آي256,000 رمز$3.00$15.00لا
30Kling 2.5 Turbo ProKuaishouلا
31Llama 3.1 8Bميتا128 ألف رمز$0.02$0.03نعم
32Llama 3.3 70Bميتا128 ألف رمز$0.10$0.32نعم
33Mistral 7Bميسترال إيه آي32K$0.02$0.03نعم
34Mistral NeMo 12Bميسترال إيه آي128 ألف رمز$0.02$0.04نعم
35NVIDIA Nemotron 3 Nano OmniNVIDIA256 ألف رمزنعم
36Qwen3 14Bعلي بابا128 ألف رمز$0.12$0.24نعم
37Qwen3 30B-A3Bعلي بابا128 ألف رمز$0.12$0.50نعم
38Qwen3 8Bعلي بابا128 ألف رمز$0.04$0.14نعم
39Sora 2OpenAIلا
40Sora 2 ProOpenAIلا
41Veo 3.1جوجللا
42Wan 2.5علي بابالا

Click any column heading to re-sort. Intelligence = a 0–100 composite of public reasoning/knowledge benchmarks; — means not yet scored. Prices are USD per 1M tokens. Updated August 2026.

الأسئلة الشائعة

What is the best LLM right now?

As of mid-2026, Claude Opus 4.8 is the highest-scoring model on the Artificial Analysis Intelligence Index (55.7), narrowly ahead of GPT-5.5 (54.8). Both lead on reasoning, coding, and agentic tasks, so the practical choice between them usually comes down to price, speed, and ecosystem rather than a meaningful intelligence gap.

What is the smartest open-source model?

GLM-5.2 is currently the most intelligent open-weight model, scoring 51.1 on the intelligence index — ahead of DeepSeek V4 Pro (44.3) and DeepSeek V4 Flash (40.3). That puts the best open models within roughly four points of proprietary frontier models like GPT-5.5, while remaining free to self-host or run through low-cost APIs.

Which AI model is the best value?

DeepSeek V4 Flash and GLM-5.2 offer the strongest intelligence-per-dollar. DeepSeek V4 Flash scores 40.3 at a tiny fraction of Opus 4.8’s price, delivering roughly 70% of the top score for a small percentage of the cost. For budget-sensitive, high-volume workloads, these open models are the clear value leaders.

How is the intelligence score calculated?

Our scores mirror the Artificial Analysis Intelligence Index v4.1, a 0-100 composite of nine demanding evaluations — including Humanity’s Last Exam, GPQA Diamond, Terminal-Bench, SciCode, and agentic tool-use tasks. It measures reasoning, knowledge, coding, and agentic ability in one number, so a higher score means broadly more capable across hard, real-world tasks.

انتقل إلى الأعلى
Featured on There's An AI For That