Google adds three Gemini Flash models while Gemini 3.5 Pro remains pending
Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber expand Google's fast family, with prices that cut output costs. The catalog keeps Gemini 3.5 Pro as coming soon.
On July 21, 2026, Google introduced three additions to its Gemini family: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber. The announcement sketches an increasingly specialized family: a Flash model for general work and agents, a light variant for high volume, and a model focused on cybersecurity. At the same time, Google DeepMind's catalog keeps a terse note next to Gemini 3.1 Pro: "3.5 Pro coming soon."
That phrase confirms an availability status, not an explanation of Google's strategy. The company has not said in its public materials why the Pro model does not accompany these three launches or what date it will have. Turning that absence into a signal of withdrawal, delay, or change of course would go beyond the sources.
Three roles within the same family
Gemini 3.6 Flash takes the spot of the fast model for code, knowledge work, and multimodal tasks. Google presents it in the catalog as best "for token efficiency" in those tasks. That description matters because, for an application, cost does not depend only on how many tokens go in and out: how much reasoning and how many tool calls a task requires also count.
Gemini 3.5 Flash-Lite is aimed at the other end of the map. Google points it at high-volume tasks that need efficiency and intelligence, but not necessarily the reasoning depth of a more capable Flash model. It is the kind of option that fits classification, extraction, repetitive transformations, or sub-steps of an agent flow, where saving latency and per-request budget can outweigh getting the most elaborate answer.
The third announcement, Gemini 3.5 Flash Cyber, introduces a security specialization. DeepMind's news page announces it as its own July release, besides including it in the joint presentation. With the official information available, it is best described as a family model aimed at that domain, without attributing operational capabilities, evaluations, or defensive uses that Google has not publicly detailed. There is also an operational fact by absence: as of this revision, Flash Cyber does not appear on the Gemini API pricing page, so anyone wanting to try it will not find a public rate there like its siblings have.
The full catalog helps place the pieces: alongside the Flash models live Gemini 3.1 Pro "for complex tasks and creative concepts," Gemini 3.1 Deep Think aimed at science and engineering, and a row of ecosystem models — Omni for multimodality, the image line, audio, robotics, and embeddings. The family is no longer one model in different sizes, but a roster of specialists.
What they cost, and what the price reveals
The API's public prices, checked on July 31, 2026 on the official pricing page, tell a story the announcement does not spell out. Gemini 3.6 Flash charges 1.50 dollars per million input tokens and 7.50 per million output tokens on the standard tier. Its immediate predecessor, Gemini 3.5 Flash, costs the same on input but 9.00 on output. The new model does not just promise token efficiency: it is cheaper per generated token than the previous Flash, 17% less on output.
Flash-Lite embodies the other extreme: 0.30 dollars per million input and 2.50 per million output — one fifth of the input and one third of the output of the new 3.6 Flash. And Gemini 3.1 Pro, the available top tier, charges double above 200,000 tokens of context: 2.00 or 4.00 on input and 12.00 or 18.00 on output depending on request size. Every tier also has a batch mode that halves prices for work that can wait. With the table in view, each use case's arithmetic becomes concrete: a million output tokens on Flash-Lite cost the same as about 139,000 on 3.1 Pro with long context.
Flash is not simply "the small model"
The Flash label can suggest that only speed matters. Yet Google's documentation for Gemini 3.5 Flash already placed that branch in agent tasks, programming loops, tool use, and long contexts, and today describes it as its "most intelligent model for sustained frontier performance" in those tasks. The stable 3.5 Flash model offers a one-million-token input context and up to 65,000 output tokens, with a knowledge cutoff of January 2025 and the platform's full battery of tools: Google Search grounding, code execution, URL context, function calling, and computer use in preview.
The new division makes visible a product decision common to current models: there is no single sweet spot. One team may want the most capable model to plan a complex task, another a lighter one to process thousands of documents, and a specialized variant when the domain demands its own knowledge, tests, and controls. Choosing well requires measuring quality on the real case, latency, cost, and each tool's risks — and remembering that price per token is not cost per task: a somewhat pricier model that solves in one pass can end up cheaper than a nominally economical one that needs three attempts and twice the tool calls.
Where non-developers will see them
The three names do not live only in the API. In the Gemini app, Google's subscriptions page lists access to 3.6 Flash as the free tier's model, with "variable access" to the Pro model; and the usage-limits help page describes Flash-Lite as the way out Google offers when a subscribed account exhausts its quota: continue the conversation with the lighter model instead of waiting. The family's division is also, in the product, a ladder of controlled degradation — from Pro to Flash and from Flash to Lite, depending on capacity and demand. The same family is bought differently depending on the door: by tokens in the API, by monthly compute quota in the app. Cyber, by contrast, does not appear in those consumer materials: its distribution, like its rate, remains publicly undescribed.
What "coming soon" means for Pro
Google keeps Gemini 3.1 Pro for complex tasks and creative development while announcing that 3.5 Pro is coming soon. That is useful information for anyone planning a migration: it is unwise to design an integration as if that model were already available, or to budget it with prices that do not yet exist.
It is also a call to distinguish catalog from rumor. The three Flash models are a verifiable launch, with model cards and — in two of the three cases — public rates; the 3.5 Pro promise indicates a roadmap, but does not allow anticipating its capabilities, price, date, or comparison with the rest. Until Google publishes that data, the soundest reading is less dramatic: it is broadening the Flash options while keeping the next Pro step open. Three cheap checks for anyone integrating these models: pin the model by exact identifier (gemini-3.6-flash, not "the newest Flash"), measure the real case with the batch rate if the work can wait, and reread the pricing page at every release — family jumps move the numbers without warning in the name.
Sources
Sources for this piece
This piece draws on 4 primary source(s), gathered during reporting.
This article was produced with artificial intelligence under human editorial oversight.