IA 360
Gemini

Gemini 1.5 Flash goes free; Pro’s price cut comes later

Google brought Gemini 1.5 Flash to the free app on July 25. The Pro price cut arrived in September, not July.

4 min read AI-generated Leer en español
Gemini 1.5 Flash goes free; Pro’s price cut comes later

Date correction: the price cut did not happen in July

On July 31, 2024, Google had not announced the price reduction this page described. The company published that cut on September 24 and made it effective October 1 for prompts below 128,000 tokens. The previous version moved the event almost two months earlier and attributed figures that were not yet public to “this week.”

The contemporary news was different. On July 25, Google brought Gemini 1.5 Flash to the free Gemini app, expanded context to 32,000 tokens and announced related-content links. This rewrite changes the headline and body to preserve what a reader could know on July 31.

What the free user actually received

Google’s announcement presented Flash as an update to the unpaid web and mobile experience. The company rolled it out in more than 40 languages and over 230 countries and territories, while claiming gains in latency, quality, reasoning and image understanding. Those are vendor statements about its product, not an independent test of every language or task.

The app’s context window increased to 32,000 tokens. That number describes the advertised maximum input, not how many words will be retrieved accurately or how long every response will take. Google also said uploads from Drive or a device and data analysis were coming soon: on the article’s date, those were announced future features rather than capabilities that should be treated as available.

Testing that window requires a document with checkable facts in different positions, questions with known answers and repeated runs. Record correct answers, omissions, invented citations, time and the effective input length. Compare a shorter context too, because accepting more material does not show that all of it improves the result. The useful unit is a task completed with evidence, not the maximum number of accepted tokens.

Related-content links were beginning to appear for fact-seeking prompts in English and selected countries. Under the published scope, a chip after a paragraph opened associated webpages; information drawn from the Gmail extension could point to the relevant email. A link makes checking easier, but it does not certify the preceding sentence or replace reading the source.

Evaluating those links asks whether each material claim has support, whether the linked page actually contains the fact and whether it is the original source. Contradictions and publication dates are then checked. Counting links measures navigation coverage; reviewing correspondence measures quality. A system may offer many doors while sending readers to pages that merely repeat the same unsupported statement.

The update also said Gemini access for teenagers would expand during the following week with dedicated policies, onboarding and an AI-literacy guide. That timetable and those safeguards belonged to the consumer app. They should not be transferred to an enterprise API, where the configuring party, submitted data and required controls differ.

Flash in the app is not Pro in the API

Gemini is a brand covering models, an application, AI Studio, an API and Vertex AI services. The free app using Flash does not automatically change Pro API pricing. Every claim should name the model, version, surface, billing mode, region and date.

An advertised context window and a used window are different too. A limit says how much one request may accept; it does not guarantee reliable retrieval of every detail or that using the maximum is cheap and fast. A document test should place information at the beginning, middle and end while recording accuracy, latency and cost.

A card that keeps surfaces separate

The minimum card has six fields: exact model name, identifier or version, access product, free or paid mode, billed unit and effective date. Region, language and volume limits belong there too. With that structure, “Gemini is cheaper” stops being an ambiguous sentence and becomes a comparison that can be checked.

The method protects consumers and technical teams alike. An app user may receive a change without choosing a version; a developer pins a model in code and pays for usage; a cloud customer may add contracts, data residency and controls. One commercial name does not make those decisions, risks or bills interchangeable.

Reconstructing an update starts by fixing the editorial date and opening the contemporary source. Available features are then separated from promises, and every pricing table is checked against its effective date. This article failed in exactly that order: it began with a later announcement and inserted it into July. The correction does not erase the September figure; it returns it to its proper place.

A simple register preserves three moments: publication of the announcement, stated availability and the first observed test or invoice. When they differ, the text explains why. It also stores the user-facing name and technical identifier, because a stable brand can hide a revision. With that discipline, a historical page continues to say what its reader could know and use on that day.

Reconstructing a historical price

A current pricing table does not prove what a service cost months earlier. Use a dated announcement, invoice, changelog or official archived capture. Preserve the unit, context tier, input, output, cache, batch mode and taxes. “Half price” may apply to only one part of consumption.

The September announcement listed different reductions for input, output and incremental caching. That is why the old summary — seven dollars to three-fifty and twenty-one to ten-fifty — did not belong in July. Even when figures are correct at another time, a false date changes decisions and destroys comparison.

Effective date deserves its own column: announcement and application do not always coincide. A budget prepared between them may know the future tariff while still paying the old one. Recording both prevents a misleading reconstruction.

Cost per token is not cost per job

The useful cost is calculated per completed task. It includes input and output tokens, retries, tool calls, storage, human review and failures. A cheap model per million may cost more when it needs extra context or correction; a pricier one may save work by succeeding first time.

A test runs ten or more real tasks, records billed tokens, total time and quality, and calculates cost per accepted result. It then varies length and concurrency. That avoids choosing from one tariff that does not represent the production workflow.

The lasting skill

The transferable skill is to date and dimension every pricing announcement: which product, tier, unit, effective date and outcome. If one field is missing, a comparison may be months out of place, mix models or promise savings that no invoice shows.

This article was produced with artificial intelligence under human editorial oversight.

Share this article

This website uses cookies to improve the browsing experience. Cookie policy.

↑↓ navigate ↵ open esc close