Friday, 7 August 2026 | Mise à jour quotidienne L'intelligence artificielle au service des constructeurs

DeepSeek V4-Pro

DeepSeek V4-Pro — Spécifications

DéveloppeurDeepSeek
TypeLLM (architecture MoE)
ModalitéTexte → Texte
Paramètres1,6 T au total / ~49 milliards actifs (MoE)
Fenêtre de contexte1 million
Sortie maximale384 K
LicenceMIT (ouverte)
Poids ouvertsOui
Publié2026-04
Prix de l’entrée0,435 $ par million
Prix de la sortie0,87 $ par million
Fournisseurs d'APIDeepSeek, OpenRouter

Exécutez-le localement

VRAM (4 bits)~800 Go
GPU minimal requisServeur multi-GPU (ex. : 8 × H100 80 Go)

Page officielle →

What is DeepSeek V4-Pro?

DeepSeek V4-Pro is the company’s open flagship: a 1.6-trillion-parameter mixture-of-experts
activating around 49B parameters per token, with a 1M-token context window and MIT-licensed
weights. The API speaks both OpenAI and Anthropic request formats, which makes it unusually
easy to drop into an existing codebase — often a base-URL and key change rather than a
rewrite.

Pricing is $0.435 in / $0.87 out per million tokens, which blends to roughly $0.52 — around
a twentieth of frontier pricing for a model that holds its own on general reasoning. That
ratio, not the parameter count, is why V4-Pro matters. The dual-format API compatibility
compounds it: the switching cost that normally protects incumbent providers largely
disappears, so V4-Pro is a genuine option for teams that would otherwise never evaluate a
Chinese lab’s model. Self-hosting is a data-centre exercise — roughly 800 GB of VRAM at
4-bit, so eight H100 80GBs or more — meaning the open licence here buys auditability and
provider portability rather than a realistic on-premises deployment for most organisations.

DeepSeek V4-Pro pricing: API cost per 1M tokens

Entrée (par million de jetons)$0.435
Sortie (par million de jetons)$0.870
Output/input ratio
Blended (4:1 in:out)$0.522 per 1M tokens

What DeepSeek V4-Pro costs per month

Real monthly spend at a 4:1 input-to-output mix — the ratio a typical chat or RAG workload actually produces.

Charge de travailJetons/moisCost / month
Projet secondaire1 million en entrée / 0,25 million en sortie$0.65
Petite équipe20 millions en entrée / 5 millions en sortie$13
Production200 millions en entrée / 50 millions en sortie$131

Run your own numbers in the Calculateur de coûts des API IA.

Cheaper alternatives to DeepSeek V4-Pro

ModèleCoût combiné par million de dollarsYou save
DeepSeek V4-Flash ouverte$0.16868 % moins cher

Self-host or pay the API?

DeepSeek V4-Pro is open-weight, so you can run it yourself. It needs ~800 Go of VRAM at 4-bit (Multi-GPU server (e.g. 8× H100 80GB)). Self-hosting only beats the API once your volume is high enough to keep that hardware busy — the calculateur auto-hébergement vs API works out the break-even point for your token volume.

Questions fréquemment posées

How much does DeepSeek V4-Pro cost per 1M tokens?

DeepSeek V4-Pro costs $0.435 per 1M input tokens and $0.870 per 1M output tokens. At a typical 4:1 input-to-output mix that blends to about $0.522 per 1M tokens.

How much does DeepSeek V4-Pro cost per month?

A small-team workload of 20M input and 5M output tokens a month costs about $13 on DeepSeek V4-Pro. A side project (1M in / 0.25M out) costs roughly $0.65.

What is a cheaper alternative to DeepSeek V4-Pro?

DeepSeek V4-Flash is the strongest cheaper option in our database at $0.168 per 1M blended — about 68% less than DeepSeek V4-Pro. It is also open-weight, so self-hosting is an option.

Can I run DeepSeek V4-Pro locally?

Yes. DeepSeek V4-Pro is open-weight and needs about ~800 GB of VRAM at 4-bit quantisation (Multi-GPU server (e.g. 8× H100 80GB)).

Why does DeepSeek V4-Pro charge more for output than input?

Output tokens are generated one at a time and cannot be batched the way a prompt can, so they cost the provider more to serve. DeepSeek V4-Pro charges 2× more for output, which is why prompt-heavy workloads are far cheaper to run than generation-heavy ones.

See every DeepSeek model priced side by side: DeepSeek API pricing.

Prices are the published list rates for the model's primary API and are reviewed as providers change them. Volume, batch and cached-input discounts are not included. Compare every model side by side in the Base de données des modèles d'IA ou le Classement des grands modèles linguistiques (LLM).

Défiler vers le haut
Featured on There's An AI For That