On July 21, 2026, Google announced three new models in the Gemini family. It isn't a generic "new model" launch — each one has a different purpose, a different price and a different audience: Gemini 3.6 Flash for general agent work, Gemini 3.5 Flash-Lite for maximum speed at minimum cost, and Gemini 3.5 Flash Cyber for government cybersecurity. It's the multi-model portfolio strategy taken to the extreme.

Gemini 3.6 Flash — the workhorse

Gemini 3.6 Flash arrives with 17% fewer output tokens than its predecessor 3.5 Flash on general benchmarks, and up to 65% fewer on coding tasks — as measured by the DeepSWE benchmark. On coding, the model reaches 49% versus 37% for the previous model. On machine-learning research, it rises from 49.7% to 63.9%.

Pricing stays at $1.50 per million input tokens and $7.50 per million output tokens. The most significant technical novelty is native "computer use" built directly into the API, letting the model control graphical interfaces without intermediate layers. It also includes strengthened safety capabilities in the CBRN domain (chemical, biological, radiological, nuclear).

Gemini 3.5 Flash-Lite — extreme speed, minimum price

Flash-Lite is the most aggressive model of the launch in terms of market positioning: 350 tokens per second, the fastest in the entire Gemini family. At $0.30 per million input tokens and $2.50 per million output tokens, it competes directly with the lowest-cost models on the market.

Benchmarks show substantial improvements over previous versions: Terminal-Bench 2.1 rises from 31% to 54%, and GDM-MRCR v2 goes from 60.1% to 72.2%. The clearest sign of its maturity is that Google has already deployed it in Google Search in production — real-world validation at a scale no benchmark can match.

"350 tokens per second on the cheapest model in the line isn't just a technical figure — it's Google telling the market that speed no longer has to cost a premium."

Gemini 3.5 Flash Cyber — the restricted model

Flash Cyber is the most unusual of the three. It is a model fine-tuned specifically to detect and repair cybersecurity vulnerabilities, designed to run inside CodeMender — Google DeepMind's security agent — with multiple instances running in parallel to speed up analysis.

Google decided to restrict access to governments and trusted partners through a pilot program. The justification is explicit: an AI capable of finding and exploiting security vulnerabilities is "dual use" by nature — it can protect or attack depending on who uses it. It's the first time Google has launched a Gemini model with deliberately limited access from day one.

What this means for the market

The strategy of multiple specialized models isn't accidental. It lets Google capture three market segments whose needs don't really overlap: developers and teams that prioritize efficiency on complex agent tasks, high-volume applications that need minimal latency at minimal cost, and government cybersecurity defenders with controlled access.

The move is also a response to competitive pressure. With OpenAI, Anthropic and Meta shipping new models in ever-shorter cycles, Google shows it can respond quickly — and with real differentiation between variants, not just incremental updates to a single model.

Worth noting: Gemini 3.5 Pro still has no confirmed release date and is being tested with selected partners. Meanwhile, Google has already started pre-training Gemini 4 — the next generation of its frontier model. The pace of releases suggests each version's life cycle is getting shorter.

  • Gemini 3.6 Flash available in the Gemini API and AI Studio
  • Gemini 3.5 Flash-Lite available in the API and rolling out in Search
  • Gemini 3.5 Flash Cyber: governments and partners only, through a pilot program
  • Gemini 3.5 Pro still in testing with partners, no confirmed date
  • Google has already started pre-training Gemini 4

Frequently asked questions

What's the difference between Gemini 3.6 Flash and 3.5 Flash-Lite?

Gemini 3.6 Flash is the main workhorse model, more efficient on complex agent and coding tasks. 3.5 Flash-Lite prioritizes extreme speed (350 tokens/sec) and minimum price ($0.30 per million input tokens), ideal for applications that need low latency at scale.

Who can use Gemini 3.5 Flash Cyber?

It is only available to governments and trusted partners through a limited pilot program. Google restricts access because an AI specialized in finding security vulnerabilities is dual use by nature.

When will Gemini 3.5 Pro arrive?

It doesn't have a confirmed release date yet. It is being tested with selected partners. Google has already started pre-training Gemini 4, its next frontier model.