Google logo

Google · Language Model

Gemini 3 Flash

Frontier intelligence built for speed and low cost.

4.4Released December 2025Fast · ~4× cheaper than Pro

Overview

Gemini 3 Flash extends the Gemini 3 family with a model tuned for speed and price — Google's pitch is "frontier intelligence at a fraction of the cost." Released in December 2025 across the Gemini app, AI Mode in Search, AI Studio and Vertex AI, it delivers a large share of Gemini 3 Pro's capability while running roughly 4× cheaper per token.

Flash is the right default for high-volume and interactive workloads where latency and cost matter: chat, summarisation, routing, lightweight coding and agent sub-steps. Pro still wins on the hardest reasoning and the most demanding multimodal tasks, but for everyday work Flash offers an excellent quality-to-cost ratio.

Key capabilities

  • Roughly 4× cheaper per token than Gemini 3 Pro
  • Fast responses for interactive and high-volume use
  • Inherits Gemini 3's multimodal strengths
  • Wide availability across Google surfaces

At a glance

Cost vs Pro

~4× cheaper

Speed

Optimised for low latency

Pricing: Low-cost Flash tier — built for volume.

Pros & cons

What we like

  • Excellent speed-to-cost ratio
  • Keeps much of Gemini 3's intelligence
  • Great default for everyday tasks

Trade-offs

  • Trails Pro on the hardest reasoning
  • Less headroom for demanding multimodal work

The verdict

4.4Aggregate review

Gemini 3 Flash is the value pick of the Gemini 3 line — fast, cheap and smart enough for the bulk of real work. Run it as your default and reach for Gemini 3 Pro only when a task genuinely needs the extra reasoning or context.

Best for: Fast, cost-efficient everyday and high-volume workloads.

Frequently asked questions

How much cheaper is Gemini 3 Flash than Pro?

Roughly four times cheaper per token, while retaining a large share of Gemini 3 Pro's capability — a strong value trade for everyday tasks.

When should I upgrade from Flash to Pro?

Step up to Gemini 3 Pro for the hardest reasoning, the longest context work and the most demanding multimodal tasks.

Further reading

Related models