How Gemini’s new limits work and where to see them
Google measures Gemini use by compute, not a fixed message count. Here is what affects a quota, when it refreshes, and how to check it.
A Gemini conversation no longer works like a simple message counter. On 17 May, Google changed the way it presents and applies Gemini Apps limits: they are now described as compute-based limits. That means two apparently similar requests can consume different portions of a quota depending on the model, selected feature, prompt complexity and conversation length.
For people who use Gemini every day, the consequence is clear: there is no fixed equivalence such as “one message, one credit.” A short request to a fast model does not carry the same capacity cost as deep research, a long conversation with files, or a task using extended reasoning. Google groups that difference under the same idea of compute usage.
How the limit refreshes
Google’s official help page says the limit refreshes every five hours until a weekly cap is reached. Running out therefore does not always mean waiting a full week: capacity may return when the next window begins. But the five-hour refresh does not remove the weekly maximum, and specific availability can change with demand.
Google also warns that it may change limits without notice because of capacity constraints. During periods of high activity, some compute-intensive features may become unavailable first to accounts without an AI plan. The practical lesson is to avoid false precision: the figures shown in a person’s own account and the timing of a warning matter more than a table copied from the web.
Plans change the relative level of access. Current documentation places Google AI Plus at twice the standard limit, Google AI Pro at four times the standard, and Google AI Ultra at five or twenty times the Pro limit depending on the subscription. That comparison describes relative access, not a promise of an identical number of replies for every feature.
A limit for each feature
Gemini Apps is not the only product that consumes capacity. Google explains that each product has its own AI limits. Using Gemini in chat, creating content in Google Flow, working with an agent in Antigravity, requesting Deep Research, or using Gemini in Gmail and Docs does not necessarily draw from one single pool of messages.
Plans also change the context size available in Gemini Apps: 32,000 tokens without a plan, 128,000 with AI Plus, and one million with AI Pro or AI Ultra. A larger context allows more material to be supplied at once, but it does not guarantee that a long task costs the same as a short query. Google explicitly lists conversation length as one factor in the limit.
Once a feature’s cap is reached, Google may offer the option to continue the conversation with Flash-Lite, wait for the quota to refresh, or move to a plan with higher limits. Pro and Ultra members can also buy AI credits to extend usage in Flow, Antigravity and other supported products. Those credits do not turn every feature into unlimited use: they work where Google supports them and add to a plan-based baseline quota.
Where to find useful information
The most reliable way to check a personal status is to open Gemini on the web, go to Settings, and select Usage limits. Google says the app warns users when they are close to a limit and again when they reach it, including when it will refresh. That screen is the operational reference because it reflects the plan, region and current capacity.
It is worth separating three questions before subscribing: which models and features are needed; how much context length is required; and which product is actually generating the usage. Someone making short queries may not gain value from a higher tier. Someone working with long documents, research or creative tools should inspect both Gemini Apps limits and the quotas of those functions.
Google’s change does not make quotas easier to summarize, but it is more transparent about what determines them. The unit is no longer an isolated message; it is the resources a particular interaction requires. Understanding that distinction helps explain why the same account may receive more or fewer responses on different days and tasks.
Sources for this piece
This piece draws on 3 primary source(s), gathered during reporting.
This article was produced with artificial intelligence under human editorial oversight.