IA 360
DeepMind

DeepMind’s Genie 3 generates interactive worlds in real time

Google DeepMind has unveiled Genie 3, a model that can create navigable environments from a text prompt at 720p and 24 frames per second. The company presents it as a simulator for training AI agents before they enter the real world.

4 min read AI-generated Leer en español
DeepMind’s Genie 3 generates interactive worlds in real time

On August 5, 2025, Google DeepMind introduced Genie 3 as a world model that generates interactive environments. The original source supports the documentary core of the event; a visually coherent environment during a demonstration is not yet a validated physical simulator.

The novelty is not simply its ability to produce convincing video. Genie 3 responds to actions within the scene and preserves elements, objects and spatial relationships for several minutes. That capability matters for training agents: instead of learning solely from images, text or recorded gameplay, they can act in an environment and observe the consequences of their decisions.

A world that keeps existing when no one is looking

A world model is a system that attempts to represent how an environment changes over time: what happens if someone moves forward, turns, opens a door or alters an object. In practice, it functions as an AI-generated simulator.

Video generators can already produce short sequences with impressive visual quality, but they generally do not let users explore a scene freely. And when they do, they often lose coherence as the view moves away from its starting point: an object may disappear, a room may change shape or an action may fail to have consistent consequences.

DeepMind says Genie 3 preserves the visual and physical consistency of the worlds it creates for several minutes. Users can explore forests, buildings, urban landscapes or imaginary settings generated from a text prompt. They can also introduce changes while navigating, such as altering the weather or changing elements of the environment.

The system’s ability to run at 24 frames per second matters because that speed enables continuous interaction rather than a slow succession of recalculated images. It still is not equivalent to a conventional graphics engine: a video game stores rules and geometries defined by its developers, while Genie 3 generates the world on the fly from patterns learned from visual data. Document supporting the figure.

From the original Genie to a simulator for agents

Google DeepMind introduced the first Genie in 2024 as a model capable of turning images into simple playable environments. Genie 2, announced at the end of that same year, expanded the concept to more varied 3D worlds. The third version shifts the focus toward sustained interaction and response speed.

The broader goal is so-called embodied AI: systems that do more than answer questions, instead perceiving an environment, planning and acting within it. A household robot, a software-operating assistant or an agent controlling a machine needs to learn action sequences, handle errors and adapt to changes. Doing that training directly in the physical world is expensive, slow and sometimes dangerous.

Traditional simulators are useful, but they require someone to design every building, object, physical rule and scenario. A model like Genie 3 promises to generate a much larger number of scenarios from a brief description. That could help expose an agent to rare situations or ones that are difficult to recreate, from navigating a construction site to following instructions in a fictional warehouse.

An important building block, not proof of general intelligence

DeepMind places world models among the technologies needed to advance toward more general-purpose systems. The idea makes sense: learning to predict the consequences of an action is central both to moving around a room and to carrying out a complex task.

But generating a plausible environment does not show that an agent understands the world as a person does. A model may maintain a coherent scene and still fail when faced with unusual physical rules, ambiguous instructions or combinations of objects that rarely appeared in its training data. The gap between performing well in a simulation and doing so reliably outside it remains one of the major challenges in robotics and autonomous agents.

The technology itself has practical limits. Genie 3’s worlds persist for minutes, not hours; accuracy for real-world locations is not guaranteed; and the available interaction is still more limited than that of a simulator designed specifically for a given task. It also does not replace testing in physical environments when safety or high-stakes decisions are involved.

DeepMind plans to make Genie 3 available to a small group of researchers and creators. That phase will show whether its central promise holds up: that agents will not merely navigate visually striking worlds, but learn skills in them that can later transfer to real-world tasks.

Turning the headline into a check

An experimental result begins with its observable variable. Record what counts as success, which behavior triggers a label and which cases fall outside scope. a visually coherent environment during a demonstration is not yet a validated physical simulator. If the phenomenon cannot be recognized without interpreting a model's intent, the conclusion needs even greater caution and a reproducible definition.

Protocol matters as much as score. Document instructions, tools, time, compute budget, number of attempts, example selection and grading rule. Changing any one may alter the result without the model learning anything new. Comparing two headlines therefore begins by checking that they measure the same axis.

A strong replication tries to break the conclusion. Add unseen data, small variants, negative controls and tasks where abstention is correct. Preserve failures as well as selected successes. To assess how to distinguish appearance, dynamics, control and usefulness for training agents, the set must resemble the intended use and reflect the cost of each error class.

What the record must preserve

A study can reveal a pattern without settling an entire field. Honest wording preserves domain, sample and date, and avoids turning 'we observed' into 'we proved forever.' Evidence becomes more valuable when another team can repeat it with available materials or state what is missing. That traceability is more useful than a sweeping label.

An evidence sheet separates four columns: what the source claims, what it shows, what it did not measure and what would change the conclusion. That discipline prevents an absence from becoming a promise and a condition from vanishing in summary. It also lets the story be updated without rewriting history from a later outcome.

Include a negative case before deciding. Find a situation where the system, rule, transaction or study does not meet the need and record the signal that would require stopping. Selected successes show that something can happen; the negative case reveals the boundary and lowers the cost of discovering it after deployment.

The skill that outlasts the announcement

A valid comparison preserves denominator and axis. It does not pit a point figure against an average, future capacity against installed capacity or a forecast against an observation. When two sources use similar language, reconstruct what they counted and over what period. If those differ, publish them as different measures instead of inventing a ranking.

The record should survive a version change. Keep URL, consultation date, document, configuration and decision. When new evidence appears, add it with its date and explain what it changes. That traceability prevents opposite errors: keeping an expired conclusion or pretending later information was known on the event date.

The transferable skill in this story is how to distinguish appearance, dynamics, control and usefulness for training agents. The procedure is short: name the document, preserve the date, fix the axis, find the condition and design a check that can fail. With those steps, a reader need not accept or reject the announcement by intuition; the decision follows a visible chain of evidence.

Before closing, another person should be able to reconstruct the conclusion without knowing the headline. Give them the sources, conditions and negative case, then ask what they would accept and reject. If they need an assumed intent, a figure without a denominator or an undated later fact, the chain still has a gap. That short review catches errors that fluent prose can conceal.

This article was produced with artificial intelligence under human editorial oversight.

Share this article

This website uses cookies to improve the browsing experience. Cookie policy.

↑↓ navigate ↵ open esc close