Google Cloud Platform combines world-class data analytics, AI infrastructure (TPUs, Vertex AI), and the original managed Kubernetes. Its global fiber backbone and Preemptible VMs offer compelling price-performance for data-heavy and containerized workloads.
Gemini 2.5 Flash inference pricing
- Developer
- Quality rank
- #16
- Elo
- 1345
- Context
- 1000K
- Weights
- Closed
- Lowest output
- $2.50
Lowest output
$2.50
Median
$2.50
Highest
$2.50
3 results
| Provider | Plan | Price | Regions | Visit | |||
|---|---|---|---|---|---|---|---|
|
|
Vertex Gemini 2.5 Flash | $2.50 | $0.300 | 1000K |
$2.50
/M tokens
Input $0.300/1M tokens
Blended $0.850
Verified
|
Global | Visit → |
|
|
Gemini 2.5 Flash | $2.50 | $0.300 | 1000K |
$2.50
/M tokens
Input $0.300/1M tokens
Blended $0.850
Verified
|
Global | Visit → |
|
|
Gemini 2.5 Flash (routed) | $2.50 | $0.300 | 1000K |
$2.50
/M tokens
Input $0.300/1M tokens
Blended $0.850
Verified
|
Global | Visit → |
Providers serving this model
OpenRouter acts as a unified gateway that routes API requests across dozens of inference providers - OpenAI, Anthropic, Google, Together, Groq, and more - through a single API key. It automatically selects the best available provider for each model, with transparent pricing and the ability to fallback if one endpoint goes down.
Frequently asked questions
How much does Gemini 2.5 Flash cost per million tokens?
The lowest input price we track for Gemini 2.5 Flash is $0.300 per million tokens. Output tokens cost more; the table shows input, output, and blended pricing for every inference provider.
What is the cheapest Gemini 2.5 Flash API provider?
Sort the table by output or blended price to find the cheapest Gemini 2.5 Flash endpoint. Prices for the same model vary widely between providers, so the cheapest provider can be several times less than the most expensive.
Which providers serve the Gemini 2.5 Flash API?
Every provider with a published Gemini 2.5 Flash endpoint appears above, with input and output token pricing, context window, and throughput.
What is Gemini 2.5 Flash's context window?
Gemini 2.5 Flash supports a 1000K context window. A larger context window lets you pass more tokens (documents, code, history) in a single request.
Is Gemini 2.5 Flash open weight or closed source?
Gemini 2.5 Flash is a closed (proprietary) model available only through hosted APIs. Compare provider pricing above to find the cheapest endpoint.
What is blended LLM cost?
Blended cost weights input and output token prices by a typical 3:1 ratio so you can rank providers by one number instead of comparing two prices separately.