Gemini 3.6 Flash — Specifications
| Sviluppatore | |
|---|---|
| Tipo | Fast / general |
| Modalità | Text, vision, audio |
| Parametri | Non divulgato |
| Finestra contestuale | 1 milione |
| Output massimo | 64K |
| Licenza | Proprietaria |
| Pesi aperti | No |
| Pubblicato | 2026-07 |
| Prezzo dell’input | 1,50 $ /1 milione |
| Prezzo dell’output | $7.50 /1M |
| Provider API | Google AI Studio, Vertex AI |
What is Gemini 3.6 Flash?
Gemini 3.6 Flash is Google’s newest Flash-tier model, released on 21 July 2026, with a
1M-token context and pricing of $1.50 in / $7.50 out per million tokens. To clear up a common
search: there is no Gemini 4 as of July 2026 — 3.6 Flash is the current latest Gemini
release, and results claiming otherwise are speculation.
Against Gemini 3.5 Flash, the input rate is unchanged at $1.50 while output drops from $9
to $7.50 — a 17% cut on the side of the bill that usually dominates. For generation-heavy
workloads, that is a straightforward saving with no capability trade-off implied by the
pricing, which makes 3.6 Flash the default choice within the Flash line for new work. Where
the difference will not register is on prompt-heavy retrieval workloads, since the input rate
is identical; there, the two models are interchangeable on cost and the decision comes down
to whatever your evaluations show. As with the rest of the Gemini family, confirm the billing
treatment of prompts above 200K tokens before designing a feature around the full context
window.
Gemini 3.6 Flash pricing: API cost per 1M tokens
| Input (per ogni milione di token) | $1.50 |
|---|---|
| Output (per ogni milione di token) | $7.50 |
| Output/input ratio | 5× |
| Blended (4:1 in:out) | $2.70 per 1M tokens |
What Gemini 3.6 Flash costs per month
Real monthly spend at a 4:1 input-to-output mix — the ratio a typical chat or RAG workload actually produces.
| Carico di lavoro | Tokens / month | Cost / month |
|---|---|---|
| Side project | 1M in / 0.25M out | $3.38 |
| Small team | 20M in / 5M out | $68 |
| Production | 200M in / 50M out | $675 |
Run your own numbers in the Calcolatore dei costi delle API IA.
Cheaper alternatives to Gemini 3.6 Flash
| Modello | Costo combinato per 1 milione di token | You save |
|---|---|---|
| Mistral 7B aperta | $0.0220 | 99% cheaper |
| Llama 3.1 8B aperta | $0.0220 | 99% cheaper |
| Mistral NeMo 12B aperta | $0.0240 | 99% cheaper |
Domande frequenti
How much does Gemini 3.6 Flash cost per 1M tokens?
Gemini 3.6 Flash costs $1.50 per 1M input tokens and $7.50 per 1M output tokens. At a typical 4:1 input-to-output mix that blends to about $2.70 per 1M tokens.
How much does Gemini 3.6 Flash cost per month?
A small-team workload of 20M input and 5M output tokens a month costs about $68 on Gemini 3.6 Flash. A side project (1M in / 0.25M out) costs roughly $3.38.
What is a cheaper alternative to Gemini 3.6 Flash?
Mistral 7B is the strongest cheaper option in our database at $0.0220 per 1M blended — about 99% less than Gemini 3.6 Flash. It is also open-weight, so self-hosting is an option.
Can I run Gemini 3.6 Flash locally?
No. Gemini 3.6 Flash is a closed, API-only model — the weights are not released, so it cannot be self-hosted.
Why does Gemini 3.6 Flash charge more for output than input?
Output tokens are generated one at a time and cannot be batched the way a prompt can, so they cost the provider more to serve. Gemini 3.6 Flash charges 5× more for output, which is why prompt-heavy workloads are far cheaper to run than generation-heavy ones.
Prices are the published list rates for the model's primary API and are reviewed as providers change them. Volume, batch and cached-input discounts are not included. Compare every model side by side in the Database di modelli IA o il classifica dei modelli linguistici (LLM leaderboard).

