Skip to main content
Comparison10 min readPublished: 2026-08-10Updated: 2026-08-10

Google Gemini API Pricing & Models (2026 Official Guide)

Compare official Gemini 3.6 Flash, 3.5 Flash, and 3.5 Flash-Lite pricing, context limits, capabilities, free tier rules, Batch, Flex, Priority, and grounding charges.

By Heizi· Founder & Editor· Published: 2026-08-10Hands-on tested

Google Gemini stable models

Intelligence across the spectrum

Start with 3.6 Flash for the latest balance, use 3.5 Flash for sustained frontier work, or scale with 3.5 Flash-Lite.

USD per 1M tokens · Paid tier · Standard

Gemini 3.6 Flash

Google’s latest stable balance of speed and intelligence for agentic and multimodal work.

Model ID
gemini-3.6-flash
Status
Stable
Input
$1.50
Cached input
$0.15
Output
$7.50
Batch / Flex
$0.75 / $3.75
Priority
$2.70 / $13.50
Input limit
1,048,576
Output limit
65,536
Latest update
Jul 2026
Capabilities
Thinking, functions, code execution, file search, Search & Maps grounding, URL context, structured outputs, Computer use (Preview)

Gemini 3.5 Flash

Sustained frontier performance for agentic workflows and coding tasks.

Model ID
gemini-3.5-flash
Status
Stable
Input
$1.50
Cached input
$0.15
Output
$9.00
Batch / Flex
$0.75 / $4.50
Priority
$2.70 / $16.20
Input limit
1,048,576
Output limit
65,536
Latest update
May 2026
Capabilities
Thinking, functions, code execution, file search, Search & Maps grounding, URL context, structured outputs, Computer use (Preview)

Gemini 3.5 Flash-Lite

The fastest, most cost-efficient stable 3.5 model for high-throughput execution.

Model ID
gemini-3.5-flash-lite
Status
Stable
Input
$0.30
Cached input
$0.03
Output
$2.50
Batch / Flex
$0.15 / $1.25
Priority
$0.54 / $4.50
Input limit
1,048,576
Output limit
65,536
Latest update
Jul 2026
Capabilities
Thinking, functions, code execution, file search, Search & Maps grounding, URL context, structured outputs, Computer use (Preview)
Official Gemini Developer API snapshot checked August 10, 2026. Pricing may change.

Current Stable Gemini Models

Google’s current stable general-purpose lineup is led by Gemini 3.6 Flash, Gemini 3.5 Flash, and Gemini 3.5 Flash-Lite. Gemini 3.6 Flash is the latest balance of speed and intelligence, 3.5 Flash targets sustained frontier agentic and coding work, and 3.5 Flash-Lite prioritizes throughput and cost.

All three accept text, images, video, audio, and PDF input and return text. Each supports 1,048,576 input tokens and 65,536 output tokens. The visual comparison above uses the official Gemini Developer API model and pricing pages checked on August 10, 2026.

  • Gemini 3.6 Flash — Standard input $1.50, cached input $0.15, output $7.50 per 1M tokens
  • Gemini 3.5 Flash — Standard input $1.50, cached input $0.15, output $9.00 per 1M tokens
  • Gemini 3.5 Flash-Lite — Standard input $0.30, cached input $0.03, output $2.50 per 1M tokens

Which Gemini Model Should You Choose?

Choose 3.6 Flash when you want Google’s newest stable general-purpose model and need strong multimodal, grounding, and agentic performance. It also has a lower Standard output price than 3.5 Flash in the current snapshot.

Choose 3.5 Flash for workloads already validated on that model or when sustained frontier coding performance matters. Choose 3.5 Flash-Lite for classification, extraction, translation, routing, and other high-volume tasks where unit cost dominates.

Standard, Batch, Flex, and Priority Pricing

Standard is the normal paid-tier rate. Batch and Flex reduce token prices for work that can accept asynchronous or flexible processing. Priority inference costs more for workloads that need prioritized capacity.

Batch and Flex are not temporary discounts. They are processing options with different delivery characteristics. Check the official documentation for availability and operational limits before routing production traffic.

  • 3.6 Flash input/output — Standard $1.50/$7.50 · Batch or Flex $0.75/$3.75 · Priority $2.70/$13.50
  • 3.5 Flash input/output — Standard $1.50/$9.00 · Batch or Flex $0.75/$4.50 · Priority $2.70/$16.20
  • 3.5 Flash-Lite input/output — Standard $0.30/$2.50 · Batch or Flex $0.15/$1.25 · Priority $0.54/$4.50

Shared Model Capabilities

The three stable models share a broad capability surface: thinking, structured outputs, function calling, code execution, file search, URL context, Search grounding, Google Maps grounding, caching, and preview access to Computer Use.

They do not generate audio or images directly and do not use the Live API. For native image, speech, live conversation, music, or video generation, use Google’s specialized Nano Banana, TTS, Live, Lyria, Imagen, or Veo model families and check their separate units and prices.

Free Tier, Paid Tier, and Data Use

The official pricing table lists free-of-charge token usage for these stable models, subject to rate limits and availability. Paid-tier usage is billed at the rates above and is marked as not used to improve Google products, while the free tier is marked as used to improve products.

Free-tier limits are account and service constraints, not a permanent unlimited allowance. Review rate limits and billing status in Google AI Studio before estimating production capacity.

Google Search and Maps Grounding Charges

Grounding charges are separate from model token prices. The current Gemini 3.x pricing table includes 5,000 free Google Search requests per month shared across Gemini 3.x models, then $14 per 1,000 requests. Google Maps grounding similarly includes a shared monthly allowance followed by $14 per 1,000 search queries.

One customer request can trigger multiple search queries, and Google charges each individual query performed. Measure actual grounding activity rather than estimating cost only from API request count.

Estimate Costs and Verify Prices

Estimate token cost with: input tokens × input rate plus output tokens (including thinking tokens) × output rate, divided by one million. Add grounding, cache storage, or specialized media charges separately.

This page is a dated editorial snapshot, not a live billing feed. Google can add models, change preview status, or update prices. Verify the model ID, pricing tier, rate limits, and final billing terms on the official Gemini Developer API pages before deployment.

FAQ

Which stable Gemini model is cheapest?

Gemini 3.5 Flash-Lite has the lowest paid Standard rate among the three highlighted stable models: $0.30 input and $2.50 output per 1M tokens.

Does Gemini 3.6 Flash have a free tier?

The official pricing page lists free-of-charge Standard token usage, subject to rate limits and availability. Check AI Studio for the limits applied to your account.

Are Batch and Flex always half the Standard price?

They are half the Standard input and output rates for the three models in this August 10, 2026 snapshot. Do not assume the same ratio for every Gemini model or future price update.

Is Google Search grounding included in token pricing?

No. Grounding has a separate monthly allowance and per-query pricing after that allowance. A single request may produce multiple billed search queries.

Related Providers

Sources