Gemini 2.5 Pro — Specifications
| Developer | Google (Google DeepMind) |
|---|---|
| Type | LLM (frontier reasoning / thinking) |
| Modality | Text, Image, Audio, Video, PDF → Text |
| Parameters | Undisclosed |
| Context window | 1M (1,048,576 tokens) |
| Max output | 64K (65,536 tokens) |
| License | Proprietary |
| Open weights | No |
| Released | March 2025 (preview); June 17, 2025 (GA) |
| Input price | $1.25 /1M |
| Output price | $10 /1M |
| API providers | Google AI Studio (Gemini API), Vertex AI, OpenRouter |
Gemini 2.5 Pro is Google DeepMind’s flagship “thinking” model from the 2.5 generation, built to reason through problems before it answers. It debuted as an experimental preview on March 25, 2025 and reached general availability on June 17, 2025, quickly establishing itself as a benchmark leader in coding, mathematics, and long-context reasoning.
Its headline feature is a native 1,048,576-token (~1M) context window paired with fully multimodal input — text, images, audio, video, and PDFs — and text output up to 65,536 tokens. Adaptive “thinking budgets” let developers trade latency and cost for deeper reasoning. The model is served through Google AI Studio, the Gemini API, and Vertex AI, with third-party access via OpenRouter.
Pricing is tiered by prompt size: $1.25 per million input tokens up to 200K, rising to $2.50 above that, with output at $10 and $15 respectively.
Verdict: Gemini 2.5 Pro remains one of the strongest value picks among frontier reasoning models — a genuine 1M-token multimodal context at mid-tier pricing. It is closed-weight and API-only, but for long-document analysis, agentic coding, and mixed-media workloads it is hard to beat. Newer Gemini 3.x models now supersede it, yet 2.5 Pro stays widely deployed and well-supported.
Gemini 2.5 Pro pricing: API cost per 1M tokens
| Input (per 1M tokens) | $1.25 |
|---|---|
| Output (per 1M tokens) | $10.00 |
| Output/input ratio | 8× |
| Blended (4:1 in:out) | $3.00 per 1M tokens |
What Gemini 2.5 Pro costs per month
Real monthly spend at a 4:1 input-to-output mix — the ratio a typical chat or RAG workload actually produces.
| Workload | Tokens / month | Cost / month |
|---|---|---|
| Side project | 1M in / 0.25M out | $3.75 |
| Small team | 20M in / 5M out | $75 |
| Production | 200M in / 50M out | $750 |
Run your own numbers in the AI API cost calculator.
Cheaper alternatives to Gemini 2.5 Pro
| Model | Blended $/1M | You save |
|---|---|---|
| Mistral 7B open | $0.0220 | 99% cheaper |
| Llama 3.1 8B open | $0.0220 | 99% cheaper |
| Mistral NeMo 12B open | $0.0240 | 99% cheaper |
Frequently asked questions
How much does Gemini 2.5 Pro cost per 1M tokens?
Gemini 2.5 Pro costs $1.25 per 1M input tokens and $10.00 per 1M output tokens. At a typical 4:1 input-to-output mix that blends to about $3.00 per 1M tokens.
How much does Gemini 2.5 Pro cost per month?
A small-team workload of 20M input and 5M output tokens a month costs about $75 on Gemini 2.5 Pro. A side project (1M in / 0.25M out) costs roughly $3.75.
What is a cheaper alternative to Gemini 2.5 Pro?
Mistral 7B is the strongest cheaper option in our database at $0.0220 per 1M blended — about 99% less than Gemini 2.5 Pro. It is also open-weight, so self-hosting is an option.
Can I run Gemini 2.5 Pro locally?
No. Gemini 2.5 Pro is a closed, API-only model — the weights are not released, so it cannot be self-hosted.
Why does Gemini 2.5 Pro charge more for output than input?
Output tokens are generated one at a time and cannot be batched the way a prompt can, so they cost the provider more to serve. Gemini 2.5 Pro charges 8× more for output, which is why prompt-heavy workloads are far cheaper to run than generation-heavy ones.
See every Gemini model priced side by side: Gemini API pricing.
Prices are the published list rates for the model's primary API and are reviewed as providers change them. Volume, batch and cached-input discounts are not included. Compare every model side by side in the AI models database or the LLM leaderboard.

