Google Gemini API Pricing & Models (2026 Official Guide)
Compare official Gemini 3.6 Flash, 3.5 Flash, and 3.5 Flash-Lite pricing, context limits, capabilities, free tier rules, Batch, Flex, Priority, and grounding charges.
Google Gemini stable models
Intelligence across the spectrum
Start with 3.6 Flash for the latest balance, use 3.5 Flash for sustained frontier work, or scale with 3.5 Flash-Lite.
USD per 1M tokens · Paid tier · Standard
Gemini 3.6 Flash
Google’s latest stable balance of speed and intelligence for agentic and multimodal work.
- Model ID
gemini-3.6-flash- Status
- Stable
- Input
- $1.50
- Cached input
- $0.15
- Output
- $7.50
- Batch / Flex
- $0.75 / $3.75
- Priority
- $2.70 / $13.50
- Input limit
- 1,048,576
- Output limit
- 65,536
- Latest update
- Jul 2026
- Capabilities
- Thinking, functions, code execution, file search, Search & Maps grounding, URL context, structured outputs, Computer use (Preview)
Gemini 3.5 Flash
Sustained frontier performance for agentic workflows and coding tasks.
- Model ID
gemini-3.5-flash- Status
- Stable
- Input
- $1.50
- Cached input
- $0.15
- Output
- $9.00
- Batch / Flex
- $0.75 / $4.50
- Priority
- $2.70 / $16.20
- Input limit
- 1,048,576
- Output limit
- 65,536
- Latest update
- May 2026
- Capabilities
- Thinking, functions, code execution, file search, Search & Maps grounding, URL context, structured outputs, Computer use (Preview)
Gemini 3.5 Flash-Lite
The fastest, most cost-efficient stable 3.5 model for high-throughput execution.
- Model ID
gemini-3.5-flash-lite- Status
- Stable
- Input
- $0.30
- Cached input
- $0.03
- Output
- $2.50
- Batch / Flex
- $0.15 / $1.25
- Priority
- $0.54 / $4.50
- Input limit
- 1,048,576
- Output limit
- 65,536
- Latest update
- Jul 2026
- Capabilities
- Thinking, functions, code execution, file search, Search & Maps grounding, URL context, structured outputs, Computer use (Preview)
Current Stable Gemini Models
Google’s current stable general-purpose lineup is led by Gemini 3.6 Flash, Gemini 3.5 Flash, and Gemini 3.5 Flash-Lite. Gemini 3.6 Flash is the latest balance of speed and intelligence, 3.5 Flash targets sustained frontier agentic and coding work, and 3.5 Flash-Lite prioritizes throughput and cost.
All three accept text, images, video, audio, and PDF input and return text. Each supports 1,048,576 input tokens and 65,536 output tokens. The visual comparison above uses the official Gemini Developer API model and pricing pages checked on August 10, 2026.
- Gemini 3.6 Flash — Standard input $1.50, cached input $0.15, output $7.50 per 1M tokens
- Gemini 3.5 Flash — Standard input $1.50, cached input $0.15, output $9.00 per 1M tokens
- Gemini 3.5 Flash-Lite — Standard input $0.30, cached input $0.03, output $2.50 per 1M tokens
Which Gemini Model Should You Choose?
Choose 3.6 Flash when you want Google’s newest stable general-purpose model and need strong multimodal, grounding, and agentic performance. It also has a lower Standard output price than 3.5 Flash in the current snapshot.
Choose 3.5 Flash for workloads already validated on that model or when sustained frontier coding performance matters. Choose 3.5 Flash-Lite for classification, extraction, translation, routing, and other high-volume tasks where unit cost dominates.
Standard, Batch, Flex, and Priority Pricing
Standard is the normal paid-tier rate. Batch and Flex reduce token prices for work that can accept asynchronous or flexible processing. Priority inference costs more for workloads that need prioritized capacity.
Batch and Flex are not temporary discounts. They are processing options with different delivery characteristics. Check the official documentation for availability and operational limits before routing production traffic.
- 3.6 Flash input/output — Standard $1.50/$7.50 · Batch or Flex $0.75/$3.75 · Priority $2.70/$13.50
- 3.5 Flash input/output — Standard $1.50/$9.00 · Batch or Flex $0.75/$4.50 · Priority $2.70/$16.20
- 3.5 Flash-Lite input/output — Standard $0.30/$2.50 · Batch or Flex $0.15/$1.25 · Priority $0.54/$4.50
Shared Model Capabilities
The three stable models share a broad capability surface: thinking, structured outputs, function calling, code execution, file search, URL context, Search grounding, Google Maps grounding, caching, and preview access to Computer Use.
They do not generate audio or images directly and do not use the Live API. For native image, speech, live conversation, music, or video generation, use Google’s specialized Nano Banana, TTS, Live, Lyria, Imagen, or Veo model families and check their separate units and prices.
Free Tier, Paid Tier, and Data Use
The official pricing table lists free-of-charge token usage for these stable models, subject to rate limits and availability. Paid-tier usage is billed at the rates above and is marked as not used to improve Google products, while the free tier is marked as used to improve products.
Free-tier limits are account and service constraints, not a permanent unlimited allowance. Review rate limits and billing status in Google AI Studio before estimating production capacity.
Google Search and Maps Grounding Charges
Grounding charges are separate from model token prices. The current Gemini 3.x pricing table includes 5,000 free Google Search requests per month shared across Gemini 3.x models, then $14 per 1,000 requests. Google Maps grounding similarly includes a shared monthly allowance followed by $14 per 1,000 search queries.
One customer request can trigger multiple search queries, and Google charges each individual query performed. Measure actual grounding activity rather than estimating cost only from API request count.
Estimate Costs and Verify Prices
Estimate token cost with: input tokens × input rate plus output tokens (including thinking tokens) × output rate, divided by one million. Add grounding, cache storage, or specialized media charges separately.
This page is a dated editorial snapshot, not a live billing feed. Google can add models, change preview status, or update prices. Verify the model ID, pricing tier, rate limits, and final billing terms on the official Gemini Developer API pages before deployment.
FAQ
Which stable Gemini model is cheapest?
Gemini 3.5 Flash-Lite has the lowest paid Standard rate among the three highlighted stable models: $0.30 input and $2.50 output per 1M tokens.
Does Gemini 3.6 Flash have a free tier?
The official pricing page lists free-of-charge Standard token usage, subject to rate limits and availability. Check AI Studio for the limits applied to your account.
Are Batch and Flex always half the Standard price?
They are half the Standard input and output rates for the three models in this August 10, 2026 snapshot. Do not assume the same ratio for every Gemini model or future price update.
Is Google Search grounding included in token pricing?
No. Grounding has a separate monthly allowance and per-query pricing after that allowance. A single request may produce multiple billed search queries.
Related in This Series
Related Providers
Sources
- Gemini Developer API PricingGoogle AI for Developers · Checked 2026-08-10
- Gemini API ModelsGoogle AI for Developers · Checked 2026-08-10
- Gemini 3.6 Flash ModelGoogle AI for Developers · Checked 2026-08-10
- Gemini 3.5 Flash ModelGoogle AI for Developers · Checked 2026-08-10
- Gemini 3.5 Flash-Lite ModelGoogle AI for Developers · Checked 2026-08-10