{"id":2099,"date":"2026-08-03T18:32:02","date_gmt":"2026-08-03T18:32:02","guid":{"rendered":"https:\/\/convly.ai\/gemini-api-pricing\/"},"modified":"2026-08-03T18:32:02","modified_gmt":"2026-08-03T18:32:02","slug":"gemini-api-pricing","status":"publish","type":"page","link":"https:\/\/convly.ai\/de\/gemini-api-pricing\/","title":{"rendered":"Gemini-API-Preise (2026): Kosten pro 1 Mio. Tokens f\u00fcr jedes Modell"},"content":{"rendered":"<p><strong>Gemini API pricing starts at $1.25 per 1M input tokens on Gemini 2.5 Pro and reaches<br \/>\n$2 on Gemini 3.1 Pro, with output between $7.50 and $12.<\/strong> The headline rates are<br \/>\ncompetitive, but Gemini&#8217;s pricing has a feature none of its rivals share to the same degree:<br \/>\nprompts above roughly 200K tokens are billed at a higher rate on several models. For a family<br \/>\nsold on its million-token context, that is the number that decides your real bill.<\/p>\n<div class=\"cvp\">\n  <h2 id=\"pricing\">Gemini API pricing: every model, per 1M tokens<\/h2>\n  <table class=\"cm-table cvp-main\">\n    <thead><tr>\n      <th>Model<\/th><th>Input $\/1M<\/th><th>Output $\/1M<\/th>\n      <th>Blended $\/1M<\/th><th>Context<\/th>\n    <\/tr><\/thead>\n    <tbody>\n          <tr>\n        <td data-label=\"Model\"><a href=\"https:\/\/convly.ai\/de\/model\/gemini-3-6-flash\/\">Gemini 3.6 Flash<\/a><\/td>\n        <td data-label=\"Input\">$1.50<\/td>\n        <td data-label=\"Output\">$7.50<\/td>\n        <td data-label=\"Blended\"><strong>$2.70<\/strong><\/td>\n        <td data-label=\"Context\">1M<\/td>\n      <\/tr>\n          <tr>\n        <td data-label=\"Model\"><a href=\"https:\/\/convly.ai\/de\/model\/gemini-2-5-pro\/\">Gemini 2.5 Pro<\/a><\/td>\n        <td data-label=\"Input\">$1.25<\/td>\n        <td data-label=\"Output\">$10.00<\/td>\n        <td data-label=\"Blended\"><strong>$3.00<\/strong><\/td>\n        <td data-label=\"Context\">1M (1,048,576 tokens)<\/td>\n      <\/tr>\n          <tr>\n        <td data-label=\"Model\"><a href=\"https:\/\/convly.ai\/de\/model\/gemini-3-5-flash\/\">Gemini 3.5 Flash<\/a><\/td>\n        <td data-label=\"Input\">$1.50<\/td>\n        <td data-label=\"Output\">$9.00<\/td>\n        <td data-label=\"Blended\"><strong>$3.00<\/strong><\/td>\n        <td data-label=\"Context\">1M<\/td>\n      <\/tr>\n          <tr>\n        <td data-label=\"Model\"><a href=\"https:\/\/convly.ai\/de\/model\/gemini-3-1-pro\/\">Gemini 3.1 Pro<\/a><\/td>\n        <td data-label=\"Input\">$2.00<\/td>\n        <td data-label=\"Output\">$12.00<\/td>\n        <td data-label=\"Blended\"><strong>$4.00<\/strong><\/td>\n        <td data-label=\"Context\">1.05M<\/td>\n      <\/tr>\n        <\/tbody>\n  <\/table>\n  <p class=\"cvp-note\">Blended is the effective rate at a 4:1 input-to-output mix, which is what a\n     typical chat or retrieval workload actually produces. It is the number to compare across\n     vendors \u2014 a headline input price hides how much the output side costs.<\/p>\n\n    <h2>What Gemini costs per month<\/h2>\n  <table class=\"cm-table cvp-tiers\">\n    <thead><tr><th>Workload<\/th><th>Tokens \/ month<\/th>\n              <th>Gemini 3.6 Flash<br><span class=\"cvp-sub\">cheapest<\/span><\/th>\n        <th>Gemini 3.1 Pro<br><span class=\"cvp-sub\">most capable tier<\/span><\/th>\n          <\/tr><\/thead>\n    <tbody>\n          <tr>\n        <td data-label=\"Workload\">Side project<\/td>\n        <td data-label=\"Tokens\">1M in \/ 0.25M out<\/td>\n        <td data-label=\"Gemini 3.6 Flash\"><strong>$3.38<\/strong><\/td>\n                <td data-label=\"Gemini 3.1 Pro\"><strong>$5.00<\/strong><\/td>\n              <\/tr>\n          <tr>\n        <td data-label=\"Workload\">Small team<\/td>\n        <td data-label=\"Tokens\">20M in \/ 5M out<\/td>\n        <td data-label=\"Gemini 3.6 Flash\"><strong>$68<\/strong><\/td>\n                <td data-label=\"Gemini 3.1 Pro\"><strong>$100<\/strong><\/td>\n              <\/tr>\n          <tr>\n        <td data-label=\"Workload\">Production<\/td>\n        <td data-label=\"Tokens\">200M in \/ 50M out<\/td>\n        <td data-label=\"Gemini 3.6 Flash\"><strong>$675<\/strong><\/td>\n                <td data-label=\"Gemini 3.1 Pro\"><strong>$1,000<\/strong><\/td>\n              <\/tr>\n        <\/tbody>\n  <\/table>\n  <p class=\"cvp-note\">Model your own volumes in the\n     <a href=\"https:\/\/convly.ai\/de\/ai-api-cost-calculator\/\">AI API cost calculator<\/a>.<\/p>\n\n    <h2>How Gemini compares to other providers<\/h2>\n  <p>Each provider's cheapest priced model, so the comparison is like for like on entry cost.<\/p>\n  <table class=\"cm-table cvp-cross\">\n    <thead><tr><th>Provider<\/th><th>Cheapest model<\/th><th>Blended $\/1M<\/th><\/tr><\/thead>\n    <tbody>\n          <tr>\n        <td data-label=\"Provider\">Mistral AI<\/td>\n        <td data-label=\"Model\"><a href=\"https:\/\/convly.ai\/de\/model\/mistral-7b\/\">Mistral 7B<\/a><\/td>\n        <td data-label=\"Blended\">$0.0220<\/td>\n      <\/tr>\n          <tr>\n        <td data-label=\"Provider\">Meta<\/td>\n        <td data-label=\"Model\"><a href=\"https:\/\/convly.ai\/de\/model\/llama-3-1-8b\/\">Llama 3.1 8B<\/a><\/td>\n        <td data-label=\"Blended\">$0.0220<\/td>\n      <\/tr>\n          <tr>\n        <td data-label=\"Provider\">Alibaba<\/td>\n        <td data-label=\"Model\"><a href=\"https:\/\/convly.ai\/de\/model\/qwen3-8b\/\">Qwen3 8B<\/a><\/td>\n        <td data-label=\"Blended\">$0.0600<\/td>\n      <\/tr>\n          <tr class=\"cvp-me\">\n        <td data-label=\"Provider\">Google <span class=\"cvp-tag\">this page<\/span><\/td>\n        <td data-label=\"Model\"><a href=\"https:\/\/convly.ai\/de\/model\/gemma-3-4b\/\">Gemma 3 4B<\/a><\/td>\n        <td data-label=\"Blended\">$0.0600<\/td>\n      <\/tr>\n          <tr>\n        <td data-label=\"Provider\">Microsoft<\/td>\n        <td data-label=\"Model\"><a href=\"https:\/\/convly.ai\/de\/model\/phi-4\/\">Phi-4<\/a><\/td>\n        <td data-label=\"Blended\">$0.0840<\/td>\n      <\/tr>\n          <tr>\n        <td data-label=\"Provider\">DeepSeek<\/td>\n        <td data-label=\"Model\"><a href=\"https:\/\/convly.ai\/de\/model\/deepseek-v4-flash\/\">DeepSeek V4-Flash<\/a><\/td>\n        <td data-label=\"Blended\">$0.168<\/td>\n      <\/tr>\n          <tr>\n        <td data-label=\"Provider\">Moonshot AI<\/td>\n        <td data-label=\"Model\"><a href=\"https:\/\/convly.ai\/de\/model\/kimi-k2-7-code\/\">Kimi K2.7 Code<\/a><\/td>\n        <td data-label=\"Blended\">$0.980<\/td>\n      <\/tr>\n          <tr>\n        <td data-label=\"Provider\">Anthropic<\/td>\n        <td data-label=\"Model\"><a href=\"https:\/\/convly.ai\/de\/model\/claude-haiku-4-5\/\">Claude Haiku 4.5<\/a><\/td>\n        <td data-label=\"Blended\">$1.80<\/td>\n      <\/tr>\n          <tr>\n        <td data-label=\"Provider\">Zhipu AI<\/td>\n        <td data-label=\"Model\"><a href=\"https:\/\/convly.ai\/de\/model\/glm-5-2\/\">GLM 5.2<\/a><\/td>\n        <td data-label=\"Blended\">$2.00<\/td>\n      <\/tr>\n          <tr>\n        <td data-label=\"Provider\">xAI<\/td>\n        <td data-label=\"Model\"><a href=\"https:\/\/convly.ai\/de\/model\/grok-4\/\">Grok 4<\/a><\/td>\n        <td data-label=\"Blended\">$5.40<\/td>\n      <\/tr>\n          <tr>\n        <td data-label=\"Provider\">OpenAI<\/td>\n        <td data-label=\"Model\"><a href=\"https:\/\/convly.ai\/de\/model\/gpt-5-6-sol\/\">GPT-5.6 Sol<\/a><\/td>\n        <td data-label=\"Blended\">$10.00<\/td>\n      <\/tr>\n        <\/tbody>\n  <\/table>\n  <p class=\"cvp-note\">Cheapest is not the same as best value \u2014 check capability alongside price on the\n     <a href=\"https:\/\/convly.ai\/de\/llm-leaderboard\/\">LLM leaderboard<\/a>, or browse every\n     model in the <a href=\"https:\/\/convly.ai\/de\/models\/\">AI models database<\/a>.<\/p>\n  <\/div>\n\n<h2>Which Gemini model should you use?<\/h2>\n<p>The interesting thing about Google&#8217;s current line is that the cheaper tier is often the<br \/>\nbetter one.<\/p>\n<p><strong>Gemini 3.5 Flash<\/strong> outperforms Gemini 3.1 Pro on coding and agentic<br \/>\nbenchmarks while running roughly four times faster and about 25% cheaper, with the same<br \/>\n1M-token context. When the faster, cheaper model also scores higher, the Pro tier stops being<br \/>\nthe default and becomes a specialist choice. For coding assistants, agent loops and anything<br \/>\nwhere a user is waiting on a response, Flash wins on every axis at once \u2014 which is rare in<br \/>\nmodel selection, where speed, cost and quality normally trade against each other.<\/p>\n<p><strong>Gemini 3.6 Flash<\/strong>, released 21 July 2026, is the newest in the line. Its<br \/>\ninput rate is unchanged at $1.50 while output drops from $9 to $7.50 \u2014 a 17% cut on the side of<br \/>\nthe bill that usually dominates. For generation-heavy workloads that is a straightforward<br \/>\nsaving with no capability trade-off implied by the pricing, which makes it the default choice<br \/>\nwithin the Flash tier for new work. On prompt-heavy retrieval workloads the two are<br \/>\ninterchangeable on cost, since input is identical. To clear up a common search: there is no<br \/>\nGemini 4 as of July 2026.<\/p>\n<p><strong>Gemini 3.1 Pro<\/strong> is worth choosing for multimodal depth rather than by<br \/>\ndefault, given that Flash beats it on coding at a lower price.<\/p>\n<p><strong>Gemini 2.5 Pro<\/strong> remains widely deployed and well supported. It reached<br \/>\ngeneral availability on 17 June 2025 and was a benchmark leader in coding, mathematics and<br \/>\nlong-context reasoning, with a native 1,048,576-token context and fully multimodal input \u2014 text,<br \/>\nimages, audio, video and PDFs. Newer 3.x models supersede it, but at $1.25 input it is still<br \/>\none of the strongest value picks among frontier reasoning models.<\/p>\n<h2>The 200K threshold is the thing to plan around<\/h2>\n<p>Gemini&#8217;s long-context billing is where budgets break. Gemini 3.1 Pro bills prompts above<br \/>\n200K tokens at roughly double the input rate. Gemini 2.5 Pro is explicitly tiered: $1.25 per<br \/>\n1M input up to 200K, rising to $2.50 above it, with output moving from $10 to $15.<\/p>\n<p>The practical consequence is that a feature costing one amount in testing can cost twice<br \/>\nthat in production, because the threshold is crossed per request and depends entirely on how<br \/>\nmuch a user pasted in. Budgeting a long-context feature on the headline figure will understate<br \/>\nthe bill substantially once real prompts start arriving.<\/p>\n<p>There are two sensible responses. Design around the threshold with aggressive retrieval and<br \/>\nreranking so prompts stay under it \u2014 which usually improves answer quality anyway, since<br \/>\nattention across very long contexts is not uniform. Or compare against a model that charges one<br \/>\nflat rate across its whole window; <a href=\"https:\/\/convly.ai\/claude-api-pricing\/\">Claude Opus<br \/>\n4.8 has no long-context premium at all<\/a>, which for a genuinely retrieval-heavy workload can<br \/>\nwork out cheaper despite a higher sticker price.<\/p>\n<h2>Google&#8217;s open models are priced separately<\/h2>\n<p>Gemma 3 is Google&#8217;s open-weight family and sits outside this table because it is a different<br \/>\nproposition: 4B, 12B and 27B models under an open licence, multimodal across text and images,<br \/>\n128K context, and cheap enough on hosted APIs ($0.05 to $0.16 per 1M) that the real decision is<br \/>\nwhether to run them yourself. Gemma 3 27B needs about 16 GB of VRAM at 4-bit, which puts it on<br \/>\na single consumer card. Size it in the<br \/>\n<a href=\"https:\/\/convly.ai\/llm-vram-calculator\/\">VRAM calculator<\/a>, or work out whether<br \/>\nself-hosting beats the API at your volume with the<br \/>\n<a href=\"https:\/\/convly.ai\/self-hosting-vs-api-calculator\/\">self-hosting vs API<br \/>\ncalculator<\/a>.<\/p>\n<div class=\"cvp\">\n  <h2 id=\"faq\">Frequently asked questions<\/h2>\n  <div class=\"cvp-faq\">\n      <details class=\"cmp-q\"><summary>How much does the Gemini API cost?<\/summary><p>Gemini API pricing runs from $2.70 per 1M blended tokens on Gemini 3.6 Flash up to $4.00 on Gemini 3.1 Pro. Blended assumes a 4:1 input-to-output mix. Input and output are billed separately, and output is always the more expensive side.<\/p><\/details>\n      <details class=\"cmp-q\"><summary>What is the cheapest Gemini model?<\/summary><p>Gemini 3.6 Flash, at $1.50 per 1M input tokens and $7.50 per 1M output. A small-team workload of 20M input and 5M output tokens a month costs about $68 on it.<\/p><\/details>\n      <details class=\"cmp-q\"><summary>How much does Gemini cost per month?<\/summary><p>At 20M input and 5M output tokens a month, Gemini 3.6 Flash costs about $68 and Gemini 3.1 Pro about $100. A side project at 1M\/0.25M costs a small fraction of that. Use the AI API cost calculator for your own volumes.<\/p><\/details>\n      <details class=\"cmp-q\"><summary>Is Gemini cheaper than Mistral AI?<\/summary><p>On entry-level pricing, Gemini starts at $2.70 per 1M blended and Mistral AI starts at $0.0220 on Mistral 7B. Mistral AI is the cheaper entry point, though capability differs \u2014 compare both on the leaderboard before switching.<\/p><\/details>\n      <details class=\"cmp-q\"><summary>Why does Gemini charge more for output tokens than input?<\/summary><p>Output tokens are generated one at a time and cannot be batched the way a prompt can, so they cost more to serve. This is why prompt-heavy workloads such as retrieval and classification are far cheaper to run than generation-heavy ones, and why the blended rate matters more than the headline input price.<\/p><\/details>\n    <\/div>\n  <p class=\"cvp-note\">Prices are the published list rates for each model's primary API and are reviewed as\n     providers change them. Volume, batch and cached-input discounts are not included. Last reviewed\n     August 2026.<\/p>\n<\/div>\n\n<script type=\"application\/ld+json\">{\"@context\":\"https:\/\/schema.org\",\"@type\":\"FAQPage\",\"mainEntity\":[{\"@type\":\"Question\",\"name\":\"How much does the Gemini API cost?\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"Gemini API pricing runs from $2.70 per 1M blended tokens on Gemini 3.6 Flash up to $4.00 on Gemini 3.1 Pro. Blended assumes a 4:1 input-to-output mix. Input and output are billed separately, and output is always the more expensive side.\"}},{\"@type\":\"Question\",\"name\":\"What is the cheapest Gemini model?\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"Gemini 3.6 Flash, at $1.50 per 1M input tokens and $7.50 per 1M output. A small-team workload of 20M input and 5M output tokens a month costs about $68 on it.\"}},{\"@type\":\"Question\",\"name\":\"How much does Gemini cost per month?\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"At 20M input and 5M output tokens a month, Gemini 3.6 Flash costs about $68 and Gemini 3.1 Pro about $100. A side project at 1M\/0.25M costs a small fraction of that. Use the AI API cost calculator for your own volumes.\"}},{\"@type\":\"Question\",\"name\":\"Is Gemini cheaper than Mistral AI?\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"On entry-level pricing, Gemini starts at $2.70 per 1M blended and Mistral AI starts at $0.0220 on Mistral 7B. Mistral AI is the cheaper entry point, though capability differs \\u2014 compare both on the leaderboard before switching.\"}},{\"@type\":\"Question\",\"name\":\"Why does Gemini charge more for output tokens than input?\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"Output tokens are generated one at a time and cannot be batched the way a prompt can, so they cost more to serve. This is why prompt-heavy workloads such as retrieval and classification are far cheaper to run than generation-heavy ones, and why the blended rate matters more than the headline input price.\"}}]}<\/script>\n\n","protected":false},"excerpt":{"rendered":"<p>Gemini API pricing starts at $1.25 per 1M input tokens on Gemini 2.5 Pro and reaches $2 on Gemini 3.1 [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":0,"parent":0,"menu_order":0,"comment_status":"closed","ping_status":"closed","template":"","meta":{"site-sidebar-layout":"default","site-content-layout":"","ast-site-content-layout":"default","site-content-style":"default","site-sidebar-style":"default","ast-global-header-display":"","ast-banner-title-visibility":"","ast-main-header-display":"","ast-hfb-above-header-display":"","ast-hfb-below-header-display":"","ast-hfb-mobile-header-display":"","site-post-title":"","ast-breadcrumbs-content":"","ast-featured-img":"","footer-sml-layout":"","ast-disable-related-posts":"","theme-transparent-header-meta":"","adv-header-id-meta":"","stick-header-meta":"","header-above-stick-meta":"","header-main-stick-meta":"","header-below-stick-meta":"","astra-migrate-meta-layouts":"default","ast-page-background-enabled":"default","ast-page-background-meta":{"desktop":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"tablet":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"mobile":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""}},"ast-content-background-meta":{"desktop":{"background-color":"var(--ast-global-color-4)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"tablet":{"background-color":"var(--ast-global-color-4)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"mobile":{"background-color":"var(--ast-global-color-4)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""}},"footnotes":""},"class_list":["post-2099","page","type-page","status-publish","hentry"],"_links":{"self":[{"href":"https:\/\/convly.ai\/de\/wp-json\/wp\/v2\/pages\/2099","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/convly.ai\/de\/wp-json\/wp\/v2\/pages"}],"about":[{"href":"https:\/\/convly.ai\/de\/wp-json\/wp\/v2\/types\/page"}],"author":[{"embeddable":true,"href":"https:\/\/convly.ai\/de\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/convly.ai\/de\/wp-json\/wp\/v2\/comments?post=2099"}],"version-history":[{"count":0,"href":"https:\/\/convly.ai\/de\/wp-json\/wp\/v2\/pages\/2099\/revisions"}],"wp:attachment":[{"href":"https:\/\/convly.ai\/de\/wp-json\/wp\/v2\/media?parent=2099"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}