IA 360
Gemini

Google brings its new Gemini to Search, making access and capability separate questions

The release spans Search, the app, developer tools, and an agent platform, but access differs across them. A matrix of surface, region, tier, status, and evaluation shows what actually shipped.

5 min read AI-generated Leer en español
Google brings its new Gemini to Search, making access and capability separate questions

On November 18, 2025, Google introduced Gemini 3 and distributed Gemini 3 Pro in preview across several products. The general announcement placed it in the Gemini app, AI Studio, Vertex AI, Google Antigravity, and Search’s AI Mode. The breadth was real, but “it shipped at Google” did not mean every person received the same model that day, in every country, without conditions.

In Search, the specific product notice limited initial access to Google AI Pro and Ultra subscribers in the United States through AI Mode’s “Thinking” selection. Automatic routing of complex queries to Gemini 3 was announced for the following weeks. Gemini Deep Think was not publicly open either: Google said it would first go to safety testers.

The durable skill from the launch is building a matrix before comparing. Its columns are surface, region, plan, version, status, permitted tools, and date. Without it, “available,” “in preview,” “rolling out,” and “coming soon” merge into an imaginary product with every advantage and no restrictions.

A family name does not identify the experience

Gemini 3 was the family. Gemini 3 Pro was its first model and launched in preview. Deep Think was an enhanced reasoning mode awaiting further evaluation. Antigravity was a development platform combining models and tools. AI Mode was a Search surface. Using “Gemini 3” for all of them hides which component produced a result.

Google announced a rollout to everyone in the Gemini app. In Search, initial access was restricted by country, subscription, and manual selection. Vertex AI and Gemini Enterprise served businesses. The Gemini API, AI Studio, Gemini CLI, and Antigravity offered other paths for developers. Each route could have different limits, prices, tools, and policies.

A reproducible test therefore begins with the model identifier, not only the brand. Add the date, product, plan, region, parameters, tools, system instructions, and number of attempts. Two people entering the same question in Search and an API are not necessarily testing the same system: one surface may search the web, route between models, construct an interface, or apply additional rules.

A benchmark is a task under conditions

Google reported Gemini 3 Pro results across several axes: 37.5% on Humanity’s Last Exam without tools, 91.9% on GPQA Diamond, 81% on MMMU-Pro, 87.6% on Video-MMMU, and 72.1% on SimpleQA Verified. It also reported a 1,501 Elo score on LMArena. These numbers do not form one grade average: each test has a question population, scoring method, and configuration.

Humanity’s Last Exam and GPQA target difficult academic questions. MMMU-Pro combines image reasoning and knowledge; Video-MMMU uses video; SimpleQA Verified evaluates short factual answers; LMArena derives a preference ranking from paired responses. Elo is a relative position in a competitive system, not a percentage of correct queries. It cannot be directly subtracted from an accuracy score.

Google DeepMind’s evaluation methodology supplied essential details: Gemini results used pass@1; single-attempt settings allowed no majority voting or parallel test-time compute; and smaller benchmarks were averaged over multiple trials. It also said some competitor numbers came from their providers, while Google calculated others through official APIs when published results were unavailable.

Even in coding, the methodology warned that SWE-bench Verified figures used different scaffoldings and infrastructure. That does not make the table worthless. It narrows the question it can answer: reported performance under described configurations, not a perfectly controlled laboratory race. A sound comparison aligns the model, tools, budget, attempts, time, and success criterion.

A context window is capacity, not guaranteed understanding

Google announced a one-million-token context window and inputs spanning text, images, audio, video, and code. The limit describes how much material the system accepts in one request, not how much it recalls accurately or what it judges relevant. Placing a repository or hours of video in context may avoid splitting, but it does not guarantee finding the decisive clause or following a relationship distributed across files.

A context test should hide verifiable facts in different positions, vary the amount of irrelevant material, and require citations or references to the source passage. Measure retrieval, accuracy, contradictions, latency, and cost. If the model summarizes a short file but omits facts in a long one, the input specification remains true while the useful capacity is smaller for that task.

Multimodality requires the same care. Accepting video does not demonstrate understanding every frame; accepting audio does not guarantee speaker separation; accepting a PDF does not ensure accurate interpretation of tables, diagrams, and footnotes. The evaluation unit should resemble the real work and preserve a set of checkable answers.

Search adds retrieval and interface generation

Google described AI Mode as using query fan-out to perform related searches, plus generative interfaces with tables, visual layouts, tools, and simulations. The visible result does not come from model weights alone. Page retrieval, source selection, generated code, interface components, and product rules all participate.

This system may be more useful than an isolated text response, but it creates additional failure points. A calculator may implement the wrong formula; a simulation may choose undisclosed assumptions; a summary may cite a page that does not support the claim. Evaluation should follow the chain: query issued, sources retrieved, transformation, calculation, link, and presented result.

The Search notice promised prominent links to web content. Testing that does not mean counting links; it means opening them and checking whether they support the adjacent claim. A visual interface can improve clarity without raising accuracy, and it can also make an error more persuasive. Design and truth are separate axes.

An agent is evaluated through its action trail

The developer guide presented Antigravity as a platform where agents could plan and execute tasks across an editor, terminal, and browser, communicating their work through artifacts. The important shift was not writing more code; it was taking actions in environments.

A tool-using agent needs a different metric from a chat. Define the objective, permissions, initial state, acceptance tests, time, cost, changes made, and reversibility. “Completed the task” is insufficient if the agent changed unrelated files, concealed a failure, or left a vulnerable dependency. Artifacts make it possible to review which command ran and what evidence it observed.

A practical trial begins in an isolated environment with a representative task. Repeat it across variants, record every action, and score the output through tests external to the agent. Then review false completions, unnecessary operations, and collateral damage. Autonomy means correct work within boundaries, not a large number of unsupervised steps.

A release card that separates promise from fact

For each product, write “available to whom, where, in what status, and with which model.” For every number, add the task, metric, configuration, tools, attempts, and source. For each demo, separate prepared input, repeatable execution, and external evidence. For every agent, record permissions and the acceptance test.

Applied on November 18, the card showed a broad but uneven deployment: Gemini 3 Pro preview across tools and apps; Search access initially limited to certain U.S. subscribers; Deep Think still with evaluators; and Antigravity as an agent platform in preview. The benchmarks showed capability under specific protocols, not universal correctness.

Google had a clear distribution advantage: it could place a new model inside existing products on day one. That scale required more precision, not less. The useful question was not whether “Gemini 3 was everywhere,” but which version served which user, with what tools, and what test would show whether it improved the user’s task.

This article was produced with artificial intelligence under human editorial oversight.

Share this article

This website uses cookies to improve the browsing experience. Cookie policy.

↑↓ navigate ↵ open esc close