Moonshot’s Kimi K3 arrived in July 2026 as the most serious open-weight challenger yet to Anthropic’s flagship Claude Opus 4.8 — and the two models represent opposite philosophies of frontier AI. One is a 2.8-trillion-parameter open mixture-of-experts you will (from July 27) be able to download; the other is a closed, polished agentic workhorse you rent through an API. Here is how they actually compare on specs, pricing, and practical deployment, using data from our live AI models database.
Fatos Principais
- Kimi K3: 2.8T total parameters (MoE, 16 of 896 experts active), 1M context, open weights due July 27, 2026
- Claude Opus 4.8: Anthropic’s proprietary flagship, 1M context, up to 128K output tokens
- API pricing: K3 at $3/$15 per million tokens vs Opus 4.8 at $5/$25 — ~40% cheaper blended ($6.00 vs $10.00 on our 3:1 index)
- Self-hosting K3 is theoretical for most: ~1.4 TB of VRAM at 4-bit — a multi-node H200 cluster, not a workstation
- Moonshot’s published benchmarks show K3 edging past Opus 4.8 on several agentic and coding suites — treat vendor numbers as claims until independent benchmarks land
Specs, Side by Side
Head to Head
| Kimi K3 | Claude Opus 4.8 | |
|---|---|---|
| Desenvolvedor | Moonshot AI | Anthropic |
| Parâmetros | 2.8T total, MoE (16/896 experts active) | Não divulgado |
| Janela de contexto | 1 milhão de tokens | 1 milhão de tokens |
| Saída máxima | — | 128K tokens |
| Licença | Open weights (from Jul 27, 2026) | Proprietary API |
| API price (in/out per 1M) | $3 / $15 | $5 / $25 |
| Run it locally? | ~1.4 TB VRAM (4-bit) — multi-node cluster | Not available |
The Pricing Math
On our blended index (3:1 input-to-output, the ratio of typical production traffic), Kimi K3 costs $6.00 per million tokens against $10.00 for Opus 4.8 — a 40% saving at near-frontier capability. Run a real workload through it: at 100M input + 20M output tokens a month, K3 bills $600 while Opus bills $1,000. Over a year that difference funds an entire additional AI project. You can model your own volumes in our Calculadora de custos de API de IA.
The context that matters: both models sit at the premium end of our índice preço-desempenho, where efficiency-tier models deliver far more capability per dollar for routine work. The K3-vs-Opus question is really about the top 10% of tasks — long-horizon agents, complex refactors, high-stakes reasoning — where flagship quality pays for itself.
Open Weights Change the Question
From July 27, K3’s weights are downloadable — but “open” does not mean “runnable at home.” At roughly 1.4 TB of VRAM in 4-bit quantization, you need a multi-node cluster (think two 8-GPU H200 nodes) before the first token flows. For almost everyone, K3 will be consumed through APIs just like Opus. The strategic difference is real, though: open weights mean multiple competing providers, price pressure over time, fine-tuning rights, and no single-vendor kill switch — the same dynamics that made DeepSeek the value king of our index.
Análise da Convly
Choose Opus 4.8 when reliability under agentic load is the product — its 128K output ceiling and mature tooling remain the safest bet for production agents. Choose K3 when cost at frontier scale is the constraint, when you need fine-tuning rights, or when vendor independence matters strategically. And if your workload is routine, choose neither: the efficiency tier does 90% of tasks at a tenth of the price. The most expensive mistake in 2026 isn’t picking the wrong flagship — it’s using any flagship for work a $0.20 model handles.
Perguntas frequentes
O Kimi K3 é melhor que o Claude Opus 4.8?
Moonshot’s published results show K3 ahead on several agentic and coding benchmarks, but independent verification is still landing. What is certain: K3 costs ~40% less blended ($6 vs $10 per 1M tokens) and its weights go open on July 27, 2026.
Posso executar o Kimi K3 localmente?
Only with data-center hardware: ~1.4 TB of VRAM at 4-bit quantization, i.e. a multi-node GPU cluster such as two 8×H200 nodes. Check any open model’s requirements in our Calculadora de VRAM para LLMs.
Which is cheaper for production use?
Kimi K3: $3/$15 per million tokens vs $5/$25 for Opus 4.8. At a typical 3:1 input:output mix, that is $6.00 vs $10.00 blended — 40% cheaper for the same volume.
Go Deeper
- Kimi K3 Explained — the full story of Moonshot’s 2.8T open model
- Kimi K3 vs ChatGPT e Claude — the three-way user’s view
- Banco de Dados Atualizado de Modelos de IA · API Cost Calculator · Price-Performance Index

