Google brings Gemini to Search, mobile and video creation
On May 14, 2024, Google I/O brought together AI Overviews, Gemini 1.5 Flash, Project Astra and Veo. They were not equally available: assessing the announcement requires separating rollout, public preview, private testing, prototype and future promise.
Google used I/O on May 14, 2024 to place Gemini across several layers of its business: generated answers in Search, a faster model for developers, an assistant that sees through a phone, and a video generator. It felt like a simultaneous launch, but each component occupied a different state. AI Overviews was beginning a US rollout; Gemini 1.5 Flash entered public preview; Veo was restricted to selected creators; Project Astra remained a prototype. Reading that availability ladder prevents a demonstration from being confused with a feature people can already use.
The significance was not merely another chatbot. Google owns an interface through which a large share of the internet looks for information, as well as Android, Workspace and a cloud platform. Putting generation at those points changes the user's path: the model stops being a separate destination and becomes a layer that summarises, organises or acts before a person sees the sources.
AI Overviews moves from the lab into Search
Google announced that AI Overviews would begin rolling out to everyone in the United States. The feature generates an answer above traditional results and includes links for further reading. It grew out of Search Generative Experience, the Search Labs experiment tested since 2023.
The company said hundreds of millions of people would gain access that week and expected to exceed one billion by the end of 2024. Those figures are not both observed users of the finished product: the first describes intended rollout reach and the second is Google's own forecast. The date and verb matter as much as the number.
The change alters the unit of trust. In a link list, users can recognise a domain, title and snippet before choosing. In a generated overview, they receive a composite synthesis first and must open sources to reconstruct where each claim originated. Google said links inside AI Overviews received more clicks than if the same pages had appeared as traditional listings for those queries, a company measurement to watch as deployment expanded.
For a low-stakes query, the summary may save steps. For health, money, law, safety or an irreversible decision, it should be treated as a map rather than the destination. The practical question is whether each important fact can be followed to a page that actually supports it. A fluent answer without visible provenance reduces friction while making a faulty combination harder to spot.
Project Astra was a vision, not an available app
Project Astra demonstrated an agent that continuously processed video and speech, answered questions about objects in view and remembered where it had seen something. The Gemini and Astra announcement described the mechanism at a high level: continuously encode frames, combine video and speech into an event timeline, and cache that information for rapid recall.
That design makes latency part of capability. A visual assistant is not conversational if the scene changes before it answers. It also makes memory a privacy question. To remember where a pair of glasses was located, the system must represent what it saw and associate it with a time. Before using such a feature, people would need to know whether processing happens on the device or on servers, how long information is retained, who can delete it and what indication is given to others captured by the camera.
The May announcement discussed prototypes and said some capabilities would come to Gemini experiences later in 2024. The displayed glasses were a possible form factor, not a product with a price or sale date. The official Project Astra page identifies it as a research prototype used by a limited set of trusted testers. A video of a feature and an available feature are different kinds of evidence.
Veo generated video, but only in private testing
Google also introduced Veo, its high-definition video model. The original Veo and Imagen 3 announcement attributed 1080p videos longer than a minute, prompt input and understanding of cinematic terms such as timelapse and aerial shots to the system. The published samples demonstrated selected results; they did not provide a failure distribution or a reproducible public comparison.
Access was narrow. A group of selected creators could test Veo in private preview through VideoFX, while everyone else could join a waitlist. Google said some capabilities would reach YouTube Shorts and other products in the future. On May 14, this was neither general availability nor an open Vertex AI preview.
All Veo videos generated in VideoFX would carry SynthID, Google's imperceptible digital watermark. That adds a signal for identifying content produced through that pipeline, but a video without the mark is not thereby authenticated. It may come from another tool, have lost metadata or sit outside a compatible system. A positive signal has a stronger meaning than its absence.
Flash did reach developers, in preview
Gemini 1.5 Flash was the most accessible part of the technical package. Google described it as lighter than 1.5 Pro and optimised for speed, cost and high-volume tasks. Both Flash and Pro were in public preview through Google AI Studio and Vertex AI with a one-million-token context window. Two-million-token Pro access required joining a waitlist.
Public preview means a developer can test a model, not that its interface, price, limits or behaviour have the stability of a generally available version. Nor does one million tokens mean one million tokens understood with identical precision. The window defines how much material a request can accept; it does not guarantee that the model will retrieve every detail, preserve distant relationships or reason correctly across the whole set.
Google said Flash was distilled from 1.5 Pro and aimed it at summarisation, chat, image and video captioning, and data extraction. Testing the speed claim requires more than its label. Measure time to a correct answer, cost per completed task and quality on the organisation's own documents. A cheaper model that forces repetition or correction can become expensive.
An availability table for reading a keynote
The four announcements can be organised through five questions. Who has access today? In which country or platform? Under which label: general release, public preview, private test or research? Which capability is present and which is future? What evidence supports the claim: observed use, benchmark, selected demonstration or company forecast?
Applied to I/O, the table becomes clear. AI Overviews was starting a product rollout in the United States. Gemini 1.5 Flash was open for developer testing but still in preview. Veo had selected users and a waitlist. Astra displayed a research direction and future integrations. The word introduced does not mean the same thing in all four rows.
The distinction also clarifies risk. A failure in Astra then affected a limited prototype; a failure in AI Overviews could scale with Search. Veo needed provenance controls before broadening access. Flash required evaluation and limits before application integration. Technical capability and exposure surface grow on different axes.
The transferable skill is to turn any keynote into a matrix of state, reach, evidence and date. First record what a person can use that day; then separate testing from promises. Finally, locate the source behind each number and preserve its verbs. This makes it possible to recognise Google's real ambition without treating a roadmap as the present or a massive rollout as an innocent demo.
This article was produced with artificial intelligence under human editorial oversight.