Monday, 3 August 2026 | Updating Daily AI insight, written for builders

Gemini 3.6 Flash

Gemini 3.6 Flash — Specifications

DesenvolvedorGoogle
TipoFast / general
ModalidadeText, vision, audio
ParâmetrosNão divulgado
Janela de contexto1 milhão
Saída máxima64K
LicençaProprietário
Pesos abertosNão
Lançado2026-07
Preço da entradaUS$ 1,50 por 1 milhão
Preço da saída$7.50 /1M
Provedores de APIGoogle AI Studio, Vertex AI

Página oficial →

What is Gemini 3.6 Flash?

Gemini 3.6 Flash is Google’s newest Flash-tier model, released on 21 July 2026, with a
1M-token context and pricing of $1.50 in / $7.50 out per million tokens. To clear up a common
search: there is no Gemini 4 as of July 2026 — 3.6 Flash is the current latest Gemini
release, and results claiming otherwise are speculation.

Against Gemini 3.5 Flash, the input rate is unchanged at $1.50 while output drops from $9
to $7.50 — a 17% cut on the side of the bill that usually dominates. For generation-heavy
workloads, that is a straightforward saving with no capability trade-off implied by the
pricing, which makes 3.6 Flash the default choice within the Flash line for new work. Where
the difference will not register is on prompt-heavy retrieval workloads, since the input rate
is identical; there, the two models are interchangeable on cost and the decision comes down
to whatever your evaluations show. As with the rest of the Gemini family, confirm the billing
treatment of prompts above 200K tokens before designing a feature around the full context
window.

Gemini 3.6 Flash pricing: API cost per 1M tokens

Entrada (por 1 milhão de tokens)$1.50
Saída (por 1 milhão de tokens)$7.50
Output/input ratio
Blended (4:1 in:out)$2.70 per 1M tokens

What Gemini 3.6 Flash costs per month

Real monthly spend at a 4:1 input-to-output mix — the ratio a typical chat or RAG workload actually produces.

Carga de trabalhoTokens / monthCost / month
Side project1M in / 0.25M out$3.38
Small team20M in / 5M out$68
Production200M in / 50M out$675

Run your own numbers in the Calculadora de custos de API de IA.

Cheaper alternatives to Gemini 3.6 Flash

ModeloCusto combinado por $/1 milhãoYou save
Mistral 7B aberta$0.022099% cheaper
Llama 3.1 8B aberta$0.022099% cheaper
Mistral NeMo 12B aberta$0.024099% cheaper

Perguntas frequentes

How much does Gemini 3.6 Flash cost per 1M tokens?

Gemini 3.6 Flash costs $1.50 per 1M input tokens and $7.50 per 1M output tokens. At a typical 4:1 input-to-output mix that blends to about $2.70 per 1M tokens.

How much does Gemini 3.6 Flash cost per month?

A small-team workload of 20M input and 5M output tokens a month costs about $68 on Gemini 3.6 Flash. A side project (1M in / 0.25M out) costs roughly $3.38.

What is a cheaper alternative to Gemini 3.6 Flash?

Mistral 7B is the strongest cheaper option in our database at $0.0220 per 1M blended — about 99% less than Gemini 3.6 Flash. It is also open-weight, so self-hosting is an option.

Can I run Gemini 3.6 Flash locally?

No. Gemini 3.6 Flash is a closed, API-only model — the weights are not released, so it cannot be self-hosted.

Why does Gemini 3.6 Flash charge more for output than input?

Output tokens are generated one at a time and cannot be batched the way a prompt can, so they cost the provider more to serve. Gemini 3.6 Flash charges 5× more for output, which is why prompt-heavy workloads are far cheaper to run than generation-heavy ones.

Prices are the published list rates for the model's primary API and are reviewed as providers change them. Volume, batch and cached-input discounts are not included. Compare every model side by side in the Banco de dados de modelos de IA ou o quadro de classificação de LLMs.

Scroll to Top
Featured on There's An AI For That