IA 360
Gemini

Google Halts Gemini's Image Generation of People Over History Errors

Google has suspended Gemini's ability to generate images of people after the model produced historically inaccurate depictions, including 1943 German soldiers and America's founding fathers with ethnicities that don't match the era's reality.

4 min read AI-generated Leer en español
Google Halts Gemini's Image Generation of People Over History Errors

Google paused Gemini’s generation of people on February 22, 2024 after acknowledging inaccurate historical results. The next day, the company published its own explanation: the feature had overcompensated in some requests and become too cautious in others.

The date matters. Google had launched image generation in Bard — soon renamed Gemini — on February 1. The pause did not happen on February 21, as this page’s original date implied, but on the 22nd; the detailed company explanation followed on the 23rd.

An overcorrection problem, not censorship

Google’s explanation identified two mechanisms: tuning intended to produce diversity did not distinguish cases where history required a particular composition, and the system became over-cautious and refused some harmless requests. That lets us describe the defect without inventing which hidden instruction, filter or dataset produced each image.

The criterion changes with the assignment. In a generic image of a profession, varied people may avoid a narrow default. In a documentary reconstruction, attributes must come from sources about the place, period and participants. The failure does not show that diversity and accuracy are incompatible; it shows that one adjustment did not decide when each objective applied.

The result has been a model that, in its attempt not to discriminate, ends up distorting verifiable historical facts—a failure that has drawn criticism both from those who see it as evidence of ideological bias baked into the AI's design and from those simply pointing out a product quality flaw.

Why it matters beyond Gemini

This episode comes barely two weeks after Google rebranded Bard as Gemini, folding image generation in as one of its flagship features to compete with rivals like OpenAI's DALL-E or Midjourney. The pause therefore affects a capability that had just been heavily promoted, and it exposes the tensions facing any company deploying generative AI at scale: balancing representation and accuracy without letting one priority swallow the other.

The case also illustrates a structural problem across the industry. Adjustments to these systems—whether through hidden instructions to the model, safety filters, or retraining—rarely distinguish well between "generating without bias" and "generating with accuracy," because both goals require different rules depending on the context of each request. A system calibrated to correct racial bias in contemporary scenes can fail spectacularly in historical scenes, and vice versa.

What Google has said and what's still unresolved

Google has not offered a timeline for restoring Gemini's ability to generate images of people, beyond Krawczyk's commitment to work "immediately" on a fix. In the meantime, the model remains available for generating images of objects, landscapes, or scenes without human figures.

The company faces a dilemma that isn't new but is a recurring one across the industry: any adjustment that reduces the risk of unwarranted demographic bias risks introducing, if applied without nuance, another kind of error that's just as visible—and, in this case, easily verifiable by any user with a basic grasp of history.

A test matrix, not an impression

Evaluation starts with a matrix of generic, historical, cultural and explicit assignments. Each case lists attributes that must appear, attributes that must not be inferred and a reference source. Generation is repeated because one image can hide variability. Reviewers then grade without knowing the version.

Errors are separated too: omission, substitution, anachronism, stereotype, unjustified refusal or harmful content. “Bias” as one label does not say what to change. Google’s own diagnosis distinguished overcompensation from excessive caution, failures that need different responses.

The set should contain controlled pairs. When only the profession, period or requested attribute changes, differences between outputs help locate what triggers a failure. Negative controls matter too: assignments where diversity is reasonable and others where a documented composition is required. Without those pairs, a collection of striking screenshots cannot show whether a problem is systematic, rare or dependent on one word.

The sample is fixed before outputs are inspected and published with its rules. Selecting only failures found on social media inflates the problem; choosing only vendor examples hides it. Record the version, date, region, language, parameters, refusals and every generation for each assignment, not merely the best or most absurd image. That trace turns controversy into evidence another team can reproduce.

Review also needs two prior decisions: which error requires an output to be withdrawn and which failure rate prevents deployment. A decorative inaccuracy, falsely assigning an identity to a real person and an image presented as historical documentation do not carry the same harm. Classifying severity before testing prevents the threshold moving after results appear. When reviewers disagree, resolve the dispute using the planned reference, not a majority impression of which image “looks” correct.

Finally, test the complete product, including its interface. A visible warning, an editable assignment and an explained refusal change usage risk even when the underlying model is unchanged. Evaluation should preserve the prompt that the generator actually received and any automatic product transformation. Otherwise, a rule added by the application may be attributed to the model, or a failure that occurs before generation may disappear from the audit.

Creative and documentary work use different contracts

A creative illustration may mix periods when the assignment asks for it. A historical reconstruction should declare that it is a representation, cite references and review verifiable details. A generator does not know the intent from “realistic” alone; the product needs controls and the user needs acceptance criteria.

Date, uniform, architecture and group composition may be facts. Colour, framing and atmosphere may be artistic choices. Separating them in the brief prevents “accuracy” from becoming one indivisible quality. Without a source, the output should not be presented as documentation.

This does not require every creative picture to reproduce demographic statistics or every historical scene to have one possible composition. It requires a declared contract. “A doctor in a contemporary hospital” leaves choices open; “the signatories present at a documented event” imposes checkable constraints. The evaluator should score only what the assignment fixes and identify details that remain interpretive.

What a product pause teaches

Disabling a feature limits harm during investigation, but it does not prove the corrected system will work. Restoration requires evaluation sets, criteria, red teaming, monitoring and an error-reporting channel. Google promised extensive testing; readers should seek results and scope, not merely reopening.

A useful reopening report would state which categories were tested, how many requests formed the set, who judged outputs, what threshold was required and which risks remain. It would separate historical accuracy, diversity in open-ended prompts and refusal rate. One average can hide improvement in one dimension and deterioration in another. Without those denominators, “we worked to improve it” describes activity rather than quality.

The transferable skill is to turn a visual controversy into a protocol: classify context, fix attributes, generate repeatedly, verify against sources and record the failure type. That sequence works with future models and avoids judging from the most viral screenshot.

This article was produced with artificial intelligence under human editorial oversight.

Share this article

This website uses cookies to improve the browsing experience. Cookie policy.

↑↓ navigate ↵ open esc close