IA 360
Gemini

Google’s AI Overviews stumble with dangerous advice

Google’s AI Overviews have recommended putting glue on pizza and eating rocks. The failures, which emerged after the US launch, expose the risk of turning text found online into authoritative-sounding answers.

4 min read AI-generated Leer en español
Google’s AI Overviews stumble with dangerous advice

On May 24, 2024, screenshots of Google Search’s new AI Overviews were circulating with absurd or dangerous answers, including suggestions involving glue on pizza and eating rocks. Google had launched the feature on May 14; a contemporaneous verification of examples documented observed outputs. Not every viral screenshot was authentic, so each case required separate tracing.

Examples shared by users include a response advising them to add non-toxic glue to pizza sauce to keep the cheese from sliding off. Another suggested eating at least one small rock a day. These are not merely wording errors: they appear inside Google’s interface, formatted as direct answers and presented alongside links that seem to support them.

The problem is more than one strange answer

A language model does not verify that a statement is true before writing it. It predicts which words are most likely to come next based on vast amounts of text. That allows it to summarize, explain and converse fluently, but it can also reproduce jokes, false information or advice stripped of its context when those appear in the sources it retrieves.

In the pizza case, the recommendation appeared to come from a satirical comment posted on Reddit years ago. The system did not understand that it was a joke, nor did it apply a strong enough safeguard to stop that kind of food advice from reaching the final answer.

That is the difficult part of bringing generative models into search. A traditional results page shows links and lets users compare sources. An automated summary, by contrast, compresses that chain of checks into a short piece of text, placed in the most visible part of the page and written in a confident tone.

The trust Google inspires does not automatically transfer to every sentence generated by a model. The company itself includes warnings telling users to verify important information. Yet asking for that caution after placing a synthesized answer above the links creates a practical contradiction: the more direct the answer appears, the less likely many people are to open the original sources.

A central bet on the future of Search

Google introduced AI Overviews on May 14 at its Google I/O conference as part of a broader transformation of its search engine. The company says the tool can help with complex queries that once required multiple searches, such as comparing products, planning activities or bringing together scattered information.

The ambition is understandable. Google is competing with assistants such as OpenAI’s ChatGPT and with the chatbots Microsoft has integrated into Bing and Windows. They are all chasing the same promise: users can ask a complete question and receive a developed answer, rather than a list of links.

But search has different demands from a creative chatbot. For queries about health, food, safety or public affairs, a convincing fabrication can be more harmful than an incomplete answer. The problem is not just that the model may get something wrong; it also matters whether it can properly distinguish between a reliable source, a forum, a parody and content designed to attract clicks.

Google has argued that the examples being shared involve uncommon queries and that most summaries provide useful information. The company has also begun limiting the appearance of AI Overviews for some problematic searches. Even so, the speed with which these cases spread shows that earlier testing did not adequately anticipate how the feature would interact with the variety, humor and low quality of the open web.

Verify before following advice

The episode does not mean automated summaries are useless. They can save steps on simple tasks, especially when they link to clear sources and users need an initial overview. But they should not be treated as an independent authority.

For recommendations affecting health, food, money or safety, it is worth opening the links, checking who published the information and turning to specialist organizations or professionals when necessary. That was already good practice online, but it matters even more when an incorrect answer is wrapped in the authority of Google’s page.

For Google, the immediate challenge will be proving that it can make AI Overviews more useful without turning search into a megaphone for plausible but false answers. The feature’s quality will not be measured by how natural it sounds, but by its ability to recognize when it should not answer.

Search and generation fail differently

A search engine retrieves documents and a model writes an answer. The system may select satire, lose context or combine incompatible passages even when every link exists. A citation at the end is insufficient: it should support the exact sentence and expose author, date and purpose.

Rare queries are useful tests because they reveal behaviour when evidence is scarce. Safe conduct may show conventional results, request clarification or abstain. Generating a fluent answer to every question turns missing evidence into an invitation to improvise.

How to verify a viral screenshot

Preserve the exact query, time, region, account, full screenshot and displayed links. Then attempt reproduction without assuming the output is stable. If it cannot be reproduced, report it as an output observed by another source or an unconfirmed image; do not combine authentic and fake examples.

Health, safety and finance require a stricter policy: primary sources, uncertainty language and direct access. An answer recommending action needs a higher threshold than a curiosity. Design can reduce triggering in those categories rather than relying only on the model.

Public evaluation should publish rate and severity rather than anecdotes alone. An absurd output may be rare and still dangerous; millions of queries turn small rates into real cases. Separate accuracy, coverage, source quality and potential harm.

The transferable skill is to decompose an AI answer into query, retrieved sources, synthesis and evidence. That chain verifies each claim and determines when a summary saves work and when opening the links is safer.

The interface should expose the source path

A synthesis may save steps when every important claim links the exact document and conventional results remain accessible. Hiding links behind an answer demands trust. Responsible design makes opening, comparing and searching again easy.

Teams maintain a set of adversarial and ordinary queries. The first tests boundaries; the second detects regressions in the main service. Fixing an anecdote through a manual rule does not show that the pattern disappeared.

Monitor who bears errors too. An absurd recipe may be visible to everyone; a wrong answer about a small community or language may go unnoticed. Sampling languages and topics with experts reduces that blind spot.

When the system changes, publish the class of failure addressed and what remains open. The explanation need not reveal exploitable defences, but it lets researchers repeat tests and readers calibrate trust.

Users also need a simple way to flag a problematic answer and submit the query. That signal supports review only with categories, priority and follow-up. A button without visible response transfers work to the public without demonstrating learning.

Before acting on an answer, open its source and find the supporting sentence. If none exists, the synthesis is a lead for investigation rather than a basis for decision.

This article was produced with artificial intelligence under human editorial oversight.

Share this article

This website uses cookies to improve the browsing experience. Cookie policy.

↑↓ navigate ↵ open esc close