{"id":2100,"date":"2026-08-03T18:32:03","date_gmt":"2026-08-03T18:32:03","guid":{"rendered":"https:\/\/convly.ai\/deepseek-api-pricing\/"},"modified":"2026-08-03T18:32:03","modified_gmt":"2026-08-03T18:32:03","slug":"deepseek-api-pricing","status":"publish","type":"page","link":"https:\/\/convly.ai\/pt\/deepseek-api-pricing\/","title":{"rendered":"Pre\u00e7os da API DeepSeek (2026): Custo por 1 milh\u00e3o de tokens para cada modelo"},"content":{"rendered":"<p><strong>DeepSeek API pricing starts at $0.14 per 1M input tokens on DeepSeek V4-Flash \u2014 around<br \/>\na sixtieth of what frontier models charge.<\/strong> That gap, not the parameter counts, is why<br \/>\nDeepSeek matters. Every model in the line is also released under an open licence, so the API<br \/>\nprice is a convenience charge rather than the only way to run them.<\/p>\n<div class=\"cvp\">\n  <h2 id=\"pricing\">DeepSeek API pricing: every model, per 1M tokens<\/h2>\n  <table class=\"cm-table cvp-main\">\n    <thead><tr>\n      <th>Model<\/th><th>Input $\/1M<\/th><th>Output $\/1M<\/th>\n      <th>Blended $\/1M<\/th><th>Context<\/th>\n    <\/tr><\/thead>\n    <tbody>\n          <tr>\n        <td data-label=\"Model\"><a href=\"https:\/\/convly.ai\/pt\/model\/deepseek-v4-flash\/\">DeepSeek V4-Flash<\/a> <span class=\"cvp-tag\">open weights<\/span><\/td>\n        <td data-label=\"Input\">$0.140<\/td>\n        <td data-label=\"Output\">$0.280<\/td>\n        <td data-label=\"Blended\"><strong>$0.168<\/strong><\/td>\n        <td data-label=\"Context\">1M<\/td>\n      <\/tr>\n          <tr>\n        <td data-label=\"Model\"><a href=\"https:\/\/convly.ai\/pt\/model\/deepseek-v4-pro\/\">DeepSeek V4-Pro<\/a> <span class=\"cvp-tag\">open weights<\/span><\/td>\n        <td data-label=\"Input\">$0.435<\/td>\n        <td data-label=\"Output\">$0.870<\/td>\n        <td data-label=\"Blended\"><strong>$0.522<\/strong><\/td>\n        <td data-label=\"Context\">1M<\/td>\n      <\/tr>\n          <tr>\n        <td data-label=\"Model\"><a href=\"https:\/\/convly.ai\/pt\/model\/deepseek-r1-distill-llama-70b\/\">DeepSeek R1 Distill Llama 70B<\/a> <span class=\"cvp-tag\">open weights<\/span><\/td>\n        <td data-label=\"Input\">$0.800<\/td>\n        <td data-label=\"Output\">$0.800<\/td>\n        <td data-label=\"Blended\"><strong>$0.800<\/strong><\/td>\n        <td data-label=\"Context\">128K<\/td>\n      <\/tr>\n          <tr>\n        <td data-label=\"Model\"><a href=\"https:\/\/convly.ai\/pt\/model\/deepseek-r1\/\">DeepSeek R1<\/a> <span class=\"cvp-tag\">open weights<\/span><\/td>\n        <td data-label=\"Input\">$0.500<\/td>\n        <td data-label=\"Output\">$2.15<\/td>\n        <td data-label=\"Blended\"><strong>$0.830<\/strong><\/td>\n        <td data-label=\"Context\">128K<\/td>\n      <\/tr>\n        <\/tbody>\n  <\/table>\n  <p class=\"cvp-note\">Blended is the effective rate at a 4:1 input-to-output mix, which is what a\n     typical chat or retrieval workload actually produces. It is the number to compare across\n     vendors \u2014 a headline input price hides how much the output side costs.<\/p>\n\n    <h2>What DeepSeek costs per month<\/h2>\n  <table class=\"cm-table cvp-tiers\">\n    <thead><tr><th>Workload<\/th><th>Tokens \/ month<\/th>\n              <th>DeepSeek V4-Flash<br><span class=\"cvp-sub\">cheapest<\/span><\/th>\n        <th>DeepSeek R1<br><span class=\"cvp-sub\">most capable tier<\/span><\/th>\n          <\/tr><\/thead>\n    <tbody>\n          <tr>\n        <td data-label=\"Workload\">Side project<\/td>\n        <td data-label=\"Tokens\">1M in \/ 0.25M out<\/td>\n        <td data-label=\"DeepSeek V4-Flash\"><strong>$0.21<\/strong><\/td>\n                <td data-label=\"DeepSeek R1\"><strong>$1.04<\/strong><\/td>\n              <\/tr>\n          <tr>\n        <td data-label=\"Workload\">Small team<\/td>\n        <td data-label=\"Tokens\">20M in \/ 5M out<\/td>\n        <td data-label=\"DeepSeek V4-Flash\"><strong>$4.20<\/strong><\/td>\n                <td data-label=\"DeepSeek R1\"><strong>$21<\/strong><\/td>\n              <\/tr>\n          <tr>\n        <td data-label=\"Workload\">Production<\/td>\n        <td data-label=\"Tokens\">200M in \/ 50M out<\/td>\n        <td data-label=\"DeepSeek V4-Flash\"><strong>$42<\/strong><\/td>\n                <td data-label=\"DeepSeek R1\"><strong>$208<\/strong><\/td>\n              <\/tr>\n        <\/tbody>\n  <\/table>\n  <p class=\"cvp-note\">Model your own volumes in the\n     <a href=\"https:\/\/convly.ai\/pt\/ai-api-cost-calculator\/\">AI API cost calculator<\/a>.<\/p>\n\n    <h2>How DeepSeek compares to other providers<\/h2>\n  <p>Each provider's cheapest priced model, so the comparison is like for like on entry cost.<\/p>\n  <table class=\"cm-table cvp-cross\">\n    <thead><tr><th>Provider<\/th><th>Cheapest model<\/th><th>Blended $\/1M<\/th><\/tr><\/thead>\n    <tbody>\n          <tr>\n        <td data-label=\"Provider\">Mistral AI<\/td>\n        <td data-label=\"Model\"><a href=\"https:\/\/convly.ai\/pt\/model\/mistral-7b\/\">Mistral 7B<\/a><\/td>\n        <td data-label=\"Blended\">$0.0220<\/td>\n      <\/tr>\n          <tr>\n        <td data-label=\"Provider\">Meta<\/td>\n        <td data-label=\"Model\"><a href=\"https:\/\/convly.ai\/pt\/model\/llama-3-1-8b\/\">Llama 3.1 8B<\/a><\/td>\n        <td data-label=\"Blended\">$0.0220<\/td>\n      <\/tr>\n          <tr>\n        <td data-label=\"Provider\">Alibaba<\/td>\n        <td data-label=\"Model\"><a href=\"https:\/\/convly.ai\/pt\/model\/qwen3-8b\/\">Qwen3 8B<\/a><\/td>\n        <td data-label=\"Blended\">$0.0600<\/td>\n      <\/tr>\n          <tr>\n        <td data-label=\"Provider\">Google<\/td>\n        <td data-label=\"Model\"><a href=\"https:\/\/convly.ai\/pt\/model\/gemma-3-4b\/\">Gemma 3 4B<\/a><\/td>\n        <td data-label=\"Blended\">$0.0600<\/td>\n      <\/tr>\n          <tr>\n        <td data-label=\"Provider\">Microsoft<\/td>\n        <td data-label=\"Model\"><a href=\"https:\/\/convly.ai\/pt\/model\/phi-4\/\">Phi-4<\/a><\/td>\n        <td data-label=\"Blended\">$0.0840<\/td>\n      <\/tr>\n          <tr class=\"cvp-me\">\n        <td data-label=\"Provider\">DeepSeek <span class=\"cvp-tag\">this page<\/span><\/td>\n        <td data-label=\"Model\"><a href=\"https:\/\/convly.ai\/pt\/model\/deepseek-v4-flash\/\">DeepSeek V4-Flash<\/a><\/td>\n        <td data-label=\"Blended\">$0.168<\/td>\n      <\/tr>\n          <tr>\n        <td data-label=\"Provider\">Moonshot AI<\/td>\n        <td data-label=\"Model\"><a href=\"https:\/\/convly.ai\/pt\/model\/kimi-k2-7-code\/\">Kimi K2.7 Code<\/a><\/td>\n        <td data-label=\"Blended\">$0.980<\/td>\n      <\/tr>\n          <tr>\n        <td data-label=\"Provider\">Anthropic<\/td>\n        <td data-label=\"Model\"><a href=\"https:\/\/convly.ai\/pt\/model\/claude-haiku-4-5\/\">Claude Haiku 4.5<\/a><\/td>\n        <td data-label=\"Blended\">$1.80<\/td>\n      <\/tr>\n          <tr>\n        <td data-label=\"Provider\">Zhipu AI<\/td>\n        <td data-label=\"Model\"><a href=\"https:\/\/convly.ai\/pt\/model\/glm-5-2\/\">GLM 5.2<\/a><\/td>\n        <td data-label=\"Blended\">$2.00<\/td>\n      <\/tr>\n          <tr>\n        <td data-label=\"Provider\">xAI<\/td>\n        <td data-label=\"Model\"><a href=\"https:\/\/convly.ai\/pt\/model\/grok-4\/\">Grok 4<\/a><\/td>\n        <td data-label=\"Blended\">$5.40<\/td>\n      <\/tr>\n          <tr>\n        <td data-label=\"Provider\">OpenAI<\/td>\n        <td data-label=\"Model\"><a href=\"https:\/\/convly.ai\/pt\/model\/gpt-5-6-sol\/\">GPT-5.6 Sol<\/a><\/td>\n        <td data-label=\"Blended\">$10.00<\/td>\n      <\/tr>\n        <\/tbody>\n  <\/table>\n  <p class=\"cvp-note\">Cheapest is not the same as best value \u2014 check capability alongside price on the\n     <a href=\"https:\/\/convly.ai\/pt\/llm-leaderboard\/\">LLM leaderboard<\/a>, or browse every\n     model in the <a href=\"https:\/\/convly.ai\/pt\/models\/\">AI models database<\/a>.<\/p>\n  <\/div>\n\n<h2>Which DeepSeek model should you use?<\/h2>\n<p><strong>V4-Flash<\/strong> is the volume workhorse: 284B total parameters with roughly 13B<br \/>\nactive per token, a 1M-token context, and a blended rate near $0.17 per 1M against a frontier<br \/>\nmodel&#8217;s $10. For bulk classification, document processing, first-pass summarisation, synthetic<br \/>\ndata generation and the retrieval layer of a RAG pipeline, it is close to unmatched on<br \/>\ncapability per dollar. A 1M context at that price does not exist elsewhere.<\/p>\n<p><strong>V4-Pro<\/strong> is the open flagship \u2014 a 1.6-trillion-parameter mixture-of-experts<br \/>\nactivating around 49B parameters per token, with a 1M-token context. At roughly $0.52 blended it<br \/>\nruns about a twentieth of frontier pricing while holding its own on general reasoning.<\/p>\n<p><strong>R1<\/strong> is the landmark open reasoning model, a 671B mixture-of-experts with 37B<br \/>\nactive, delivering chain-of-thought reasoning at a fraction of proprietary cost. Its 128K<br \/>\ncontext is narrower than the V4 line&#8217;s 1M.<\/p>\n<p><strong>R1 Distill Llama 70B<\/strong> exists for one reason: hardware. Full R1 needs roughly<br \/>\n400 GB of VRAM at 4-bit; the distill needs about 40 GB, which moves it from &#8220;rent a server&#8221; to<br \/>\n&#8220;buy a workstation&#8221;. It also bills $0.80 for both input and output, so generation-heavy<br \/>\nworkloads cost no more than prompt-heavy ones \u2014 unusually simple to budget.<\/p>\n<h2>The switching cost is lower than you think<\/h2>\n<p>The detail that makes DeepSeek a genuine option rather than a curiosity is that the V4-Pro<br \/>\nAPI speaks both OpenAI and Anthropic request formats. Migrating is usually a base-URL and key<br \/>\nchange rather than a rewrite. The integration friction that normally protects incumbent<br \/>\nproviders largely disappears, which is what makes a 20x cost difference actionable instead of<br \/>\ntheoretical.<\/p>\n<p>The sensible pattern is not wholesale migration but routing: send the high-volume,<br \/>\nlow-stakes majority of traffic to V4-Flash and keep a frontier model for the requests that<br \/>\ngenuinely need it. Most applications discover that the split is heavily weighted toward the<br \/>\ncheap side.<\/p>\n<h2>Do the open weights change the maths?<\/h2>\n<p>Every DeepSeek model here is openly downloadable, which raises the obvious question. For<br \/>\nmost teams the honest answer is that the licence buys auditability and provider portability<br \/>\nrather than a realistic on-premises deployment.<\/p>\n<p>The hardware requirements are the reason. V4-Pro needs roughly 800 GB of VRAM at 4-bit \u2014 a<br \/>\ndata-centre exercise of eight H100 80GBs or more. R1 needs about 400 GB. V4-Flash needs about<br \/>\n140 GB, or two H100 80GBs. Only the R1 distill, at 40 GB, fits hardware a small team would<br \/>\nactually buy.<\/p>\n<p>Set against API pricing this low, self-hosting has to clear a high bar: it only pays once<br \/>\nyour volume is high enough to keep that hardware continuously busy, and the break-even point<br \/>\nmoves every time DeepSeek cuts prices. Work it out for your own token volume in the<br \/>\n<a href=\"https:\/\/convly.ai\/self-hosting-vs-api-calculator\/\">self-hosting vs API calculator<\/a><br \/>\nbefore committing to hardware. What the open weights do guarantee is that you can never be<br \/>\nlocked in, price-gouged, or stranded by a deprecation \u2014 which is worth something even if you<br \/>\nnever download them.<\/p>\n<div class=\"cvp\">\n  <h2 id=\"faq\">Frequently asked questions<\/h2>\n  <div class=\"cvp-faq\">\n      <details class=\"cmp-q\"><summary>How much does the DeepSeek API cost?<\/summary><p>DeepSeek API pricing runs from $0.168 per 1M blended tokens on DeepSeek V4-Flash up to $0.830 on DeepSeek R1. Blended assumes a 4:1 input-to-output mix. Input and output are billed separately, and output is always the more expensive side.<\/p><\/details>\n      <details class=\"cmp-q\"><summary>What is the cheapest DeepSeek model?<\/summary><p>DeepSeek V4-Flash, at $0.140 per 1M input tokens and $0.280 per 1M output. A small-team workload of 20M input and 5M output tokens a month costs about $4.20 on it.<\/p><\/details>\n      <details class=\"cmp-q\"><summary>How much does DeepSeek cost per month?<\/summary><p>At 20M input and 5M output tokens a month, DeepSeek V4-Flash costs about $4.20 and DeepSeek R1 about $21. A side project at 1M\/0.25M costs a small fraction of that. Use the AI API cost calculator for your own volumes.<\/p><\/details>\n      <details class=\"cmp-q\"><summary>Is DeepSeek cheaper than Mistral AI?<\/summary><p>On entry-level pricing, DeepSeek starts at $0.168 per 1M blended and Mistral AI starts at $0.0220 on Mistral 7B. Mistral AI is the cheaper entry point, though capability differs \u2014 compare both on the leaderboard before switching.<\/p><\/details>\n      <details class=\"cmp-q\"><summary>Why does DeepSeek charge more for output tokens than input?<\/summary><p>Output tokens are generated one at a time and cannot be batched the way a prompt can, so they cost more to serve. This is why prompt-heavy workloads such as retrieval and classification are far cheaper to run than generation-heavy ones, and why the blended rate matters more than the headline input price.<\/p><\/details>\n    <\/div>\n  <p class=\"cvp-note\">Prices are the published list rates for each model's primary API and are reviewed as\n     providers change them. Volume, batch and cached-input discounts are not included. Last reviewed\n     August 2026.<\/p>\n<\/div>\n\n<script type=\"application\/ld+json\">{\"@context\":\"https:\/\/schema.org\",\"@type\":\"FAQPage\",\"mainEntity\":[{\"@type\":\"Question\",\"name\":\"How much does the DeepSeek API cost?\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"DeepSeek API pricing runs from $0.168 per 1M blended tokens on DeepSeek V4-Flash up to $0.830 on DeepSeek R1. Blended assumes a 4:1 input-to-output mix. Input and output are billed separately, and output is always the more expensive side.\"}},{\"@type\":\"Question\",\"name\":\"What is the cheapest DeepSeek model?\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"DeepSeek V4-Flash, at $0.140 per 1M input tokens and $0.280 per 1M output. A small-team workload of 20M input and 5M output tokens a month costs about $4.20 on it.\"}},{\"@type\":\"Question\",\"name\":\"How much does DeepSeek cost per month?\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"At 20M input and 5M output tokens a month, DeepSeek V4-Flash costs about $4.20 and DeepSeek R1 about $21. A side project at 1M\/0.25M costs a small fraction of that. Use the AI API cost calculator for your own volumes.\"}},{\"@type\":\"Question\",\"name\":\"Is DeepSeek cheaper than Mistral AI?\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"On entry-level pricing, DeepSeek starts at $0.168 per 1M blended and Mistral AI starts at $0.0220 on Mistral 7B. Mistral AI is the cheaper entry point, though capability differs \\u2014 compare both on the leaderboard before switching.\"}},{\"@type\":\"Question\",\"name\":\"Why does DeepSeek charge more for output tokens than input?\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"Output tokens are generated one at a time and cannot be batched the way a prompt can, so they cost more to serve. This is why prompt-heavy workloads such as retrieval and classification are far cheaper to run than generation-heavy ones, and why the blended rate matters more than the headline input price.\"}}]}<\/script>\n\n","protected":false},"excerpt":{"rendered":"<p>DeepSeek API pricing starts at $0.14 per 1M input tokens on DeepSeek V4-Flash \u2014 around a sixtieth of what frontier [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":0,"parent":0,"menu_order":0,"comment_status":"closed","ping_status":"closed","template":"","meta":{"site-sidebar-layout":"default","site-content-layout":"","ast-site-content-layout":"default","site-content-style":"default","site-sidebar-style":"default","ast-global-header-display":"","ast-banner-title-visibility":"","ast-main-header-display":"","ast-hfb-above-header-display":"","ast-hfb-below-header-display":"","ast-hfb-mobile-header-display":"","site-post-title":"","ast-breadcrumbs-content":"","ast-featured-img":"","footer-sml-layout":"","ast-disable-related-posts":"","theme-transparent-header-meta":"","adv-header-id-meta":"","stick-header-meta":"","header-above-stick-meta":"","header-main-stick-meta":"","header-below-stick-meta":"","astra-migrate-meta-layouts":"default","ast-page-background-enabled":"default","ast-page-background-meta":{"desktop":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"tablet":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"mobile":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""}},"ast-content-background-meta":{"desktop":{"background-color":"var(--ast-global-color-4)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"tablet":{"background-color":"var(--ast-global-color-4)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"mobile":{"background-color":"var(--ast-global-color-4)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""}},"footnotes":""},"class_list":["post-2100","page","type-page","status-publish","hentry"],"_links":{"self":[{"href":"https:\/\/convly.ai\/pt\/wp-json\/wp\/v2\/pages\/2100","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/convly.ai\/pt\/wp-json\/wp\/v2\/pages"}],"about":[{"href":"https:\/\/convly.ai\/pt\/wp-json\/wp\/v2\/types\/page"}],"author":[{"embeddable":true,"href":"https:\/\/convly.ai\/pt\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/convly.ai\/pt\/wp-json\/wp\/v2\/comments?post=2100"}],"version-history":[{"count":0,"href":"https:\/\/convly.ai\/pt\/wp-json\/wp\/v2\/pages\/2100\/revisions"}],"wp:attachment":[{"href":"https:\/\/convly.ai\/pt\/wp-json\/wp\/v2\/media?parent=2100"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}