All posts
Production Published 14 min

Gemini 3.6 Flash shipped, 3.5 Pro did not: the routing checklist I run after July 21

On July 21, 2026 Google released Gemini 3.6 Flash ($1.50/$7.50 per million), 3.5 Flash-Lite ($0.30/$2.50, ~350 tok/s), and a limited-access 3.5 Flash Cyber via CodeMender. Gemini 3.5 Pro stays in partner testing. Here is how I re-benchmark cost per completed task against Sonnet 5 and Fable without betting the roadmap on Pro.

Jigar JoshiJigar JoshiAgentic AI Architect and Consultant
In this post (6 sections)

Introduction

The market wanted Pro. Google shipped efficiency. That is not a failure for agent builders if you measure the right thing. I already treat Flash as the default Gemini lane in production routers from the Gemini 3.5 Flash vs Sonnet 4.6 playbook. July 21 is a model-ID and price refresh inside that lane, plus a clear statement that Pro is still not a date you can schedule against.

What shipped (and what did not)

July 21 Gemini releases at a glance
ModelRoleList price (per 1M tokens)Availability
Gemini 3.6 FlashWorkhorse agent/coding/knowledge$1.50 in / $7.50 outGA: API, AI Studio, Enterprise Agent Platform, Gemini app
Gemini 3.5 Flash-LiteHigh-throughput / low latency$0.30 in / $2.50 outGA; also rolling into Search
Gemini 3.5 Flash CyberVulnerability find/fixNot a public list priceLimited: governments and trusted partners via CodeMender
Gemini 3.5 ProFrontier flagship (expected)n/aPartner testing only; "as soon as it's ready"

Google also says computer use is available as a built-in client-side tool on the Flash surfaces, and that Gemini 4 pre-training has started. Useful context. Not a reason to pause this week's evals.

How I update the routing sheet

  1. 01
    Add 3.6 Flash and 3.5 Flash-Lite as named rows
    Do not silently alias "gemini-flash" in config. Pin gemini-3.6-flash and gemini-3.5-flash-lite (or your provider's exact IDs) with a documented fallback.
  2. 02
    Re-run cost per completed task, not tokens alone
    3.6 Flash claims fewer output tokens and fewer tool loops on some agentic benches. That only matters if your workflows finish cheaper and more reliably. Measure completed tasks on your suite from eval datasets beyond the happy path.
  3. 03
    Compare against Sonnet 5 and Fable fallbacks
    Sonnet 5 still sits at intro pricing through August 31 on Anthropic's side. Fable remains a gated frontier ID with safeguard routing. Use the same three-column sheet: Gemini lane, Claude mid-tier, Claude frontier. See Sonnet 5 migration and Fable redeployment.
  4. 04
    Put Flash-Lite on classification, extraction, and fan-out reads
    350 tok/s and $0.30/$2.50 is where high-volume subagents belong. Keep 3.6 Flash for multi-step coding and knowledge work that used 3.5 Flash before.
  5. 05
    Leave 3.5 Pro off committed roadmaps
    Partner testing is not GA. Bloomberg-style delay reporting plus Google's own "when ready" language means Pro is upside, not a dependency. Same discipline as any revocable frontier model.
  6. 06
    Treat Flash Cyber as a governance product, not a model swap
    Limited access through CodeMender is closer to Anthropic's Glasswing pattern than to a public API tier. Do not plan general coding agents on Cyber.

Suggested default lanes (until your evals say otherwise)

  • High-volume extract / classify / translate: Gemini 3.5 Flash-Lite.
  • Default agentic coding and multimodal knowledge work on Google: Gemini 3.6 Flash.
  • Claude mid-tier default where tool schemas and Agent SDK matter: Sonnet 5.
  • Long-horizon frontier maker work: Fable 5 only with fallbacks and spend caps.
  • Security vuln find/fix at scale: wait for partner programs (Flash Cyber / Glasswing-class), do not DIY with a public chat model.

Common mistakes this week

  • Rewriting the roadmap around Pro because a leak said July 17.
  • Flipping all Gemini traffic to 3.6 Flash without measuring tool-call count regressions.
  • Putting Flash-Lite on tasks that need deep multi-file reasoning just because it is cheap.
  • Mixing Flash Cyber headlines into app-sec agent demos for customers who cannot access it.

Conclusion

July 21 improved the Gemini efficiency lane and clarified that Pro is still not a date. Update IDs, re-run cost per completed task, keep Pro off the critical path. The routing layer you already have matters more than the missing flagship.

Sources: Google blog "Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber" (July 21, 2026) at Google blog; TechCrunch coverage of the same-day launch noting Pro remains unreleased.

The weekly take

Agentic AI patterns, delivered Thursdays

What I am shipping, watching, and pruning out of client stacks each week. One email. No fluff.

Shipping an agentic AI project this quarter?
Book a 30-min consult
Frequently asked

Questions readers ask about this post

Share this post
LinkedIn Facebook