Qwen3 8B — Spécifications
Rédigé par Mustafa Ihsan d’après la documentation officielle du fournisseur · Dernière mise à jour
| Développeur | Alibaba |
|---|---|
| Type | LLM (dense) |
| Modalité | Texte → Texte |
| Paramètres | 8B |
| Fenêtre de contexte | 128 K |
| Licence | Apache 2.0 (ouverte) |
| Poids ouverts | Oui |
| Publié | 2025 |
| Prix de l’entrée | $0.04 /1M |
| Prix de la sortie | $0.14 /1M |
| Fournisseurs d'API | Alibaba, OpenRouter, Ollama |
Exécutez-le localement
| VRAM (4 bits) | ~5 Go |
|---|---|
| GPU minimal | RTX 3060 8 Go / toute carte graphique 8 Go |
What is Qwen3 8B?
Qwen3 8B is a small, fast dense model from the Qwen3 family — Apache 2.0, a 128K context,
and about 5 GB of VRAM at 4-bit, which fits an RTX 3060 8GB or any 8 GB card. Hosted pricing
is $0.04 in / $0.14 out per million tokens.
Among models that fit an 8 GB GPU, it is one of the strongest available, and the 128K
context is what separates it from the older generation in the same bracket — Mistral 7B, its
closest historical equivalent, tops out at 32K. That difference decides whether a local
assistant can hold a real document or only a few pages. Apache 2.0 licensing means no
conditions to clear with legal, and the small footprint leaves room on the card for a
generous context or a second process. Treat it as the entry point to local inference: it will
run on hardware you almost certainly already have, and if it proves the use case, Qwen3 14B
and 32B are drop-in upgrades within the same family and licence as your hardware
allows.
Qwen3 8B pricing: API cost per 1M tokens
| Entrée (par million de jetons) | $0.0400 |
|---|---|
| Sortie (par million de jetons) | $0.140 |
| Ratio sortie/entrée | 3.5× |
| Moyenne pondérée (4:1 entrée:sortie) | $0.0600 par 1 million de jetons |
What Qwen3 8B costs per month
Dépense mensuelle réelle pour un ratio entrée-sortie de 4:1 — le rapport effectivement observé dans une charge de travail type (chat ou RAG).
| Charge de travail | Jetons/mois | Coût par mois |
|---|---|---|
| Projet secondaire | 1 million en entrée / 0,25 million en sortie | $0.08 |
| Petite équipe | 20 millions en entrée / 5 millions en sortie | $1.50 |
| Production | 200 millions en entrée / 50 millions en sortie | $15 |
Calculez vos propres coûts dans le Calculateur de coûts pour les API IA.
Cheaper alternatives to Qwen3 8B
| Modèle | Coût combiné par million de dollars | Vous économisez |
|---|---|---|
| Mistral 7B ouverte | $0.0220 | 63 % moins cher |
| Llama 3.1 8B ouverte | $0.0220 | 63 % moins cher |
| Mistral NeMo 12B ouverte | $0.0240 | 60% cheaper |
Hébergement local ou paiement via API ?
Qwen3 8B is open-weight, so you can run it yourself. It needs ~5 Go of VRAM at 4-bit (RTX 3060 8GB / any 8GB GPU). Self-hosting only beats the API once your volume is high enough to keep that hardware busy — the calculateur auto-hébergement vs API calcule le seuil de rentabilité pour votre volume de jetons.
Questions fréquemment posées
How much does Qwen3 8B cost per 1M tokens?
Qwen3 8B costs $0.0400 per 1M input tokens and $0.140 per 1M output tokens. At a typical 4:1 input-to-output mix that blends to about $0.0600 per 1M tokens.
How much does Qwen3 8B cost per month?
A small-team workload of 20M input and 5M output tokens a month costs about $1.50 on Qwen3 8B. A side project (1M in / 0.25M out) costs roughly $0.08.
What is a cheaper alternative to Qwen3 8B?
Mistral 7B is the strongest cheaper option in our database at $0.0220 per 1M blended — about 63% less than Qwen3 8B. It is also open-weight, so self-hosting is an option.
Can I run Qwen3 8B locally?
Yes. Qwen3 8B is open-weight and needs about ~5 GB of VRAM at 4-bit quantisation (RTX 3060 8GB / any 8GB GPU).
Why does Qwen3 8B charge more for output than input?
Output tokens are generated one at a time and cannot be batched the way a prompt can, so they cost the provider more to serve. Qwen3 8B charges 3.5× more for output, which is why prompt-heavy workloads are far cheaper to run than generation-heavy ones.
Les prix indiqués correspondent aux tarifs publics officiels pour l’API principale du modèle et sont mis à jour régulièrement à mesure que les fournisseurs les modifient. Les remises liées aux volumes, aux traitements par lots ou aux entrées mises en cache ne sont pas prises en compte. Comparez tous les modèles côte à côte dans le Base de données des modèles IA ou le Classement des grands modèles linguistiques (LLM).
