IA 360
Current Affairs

LeCun announces his Meta exit and poses a test for world models

The scientist announced a company to pursue systems with physical understanding, memory, and planning, without attributing his departure to a dispute. A world model is judged by what it predicts and enables, not by its label.

4 min read AI-generated Leer en español
LeCun announces his Meta exit and poses a test for world models

On November 19, 2025, Yann LeCun announced that he would leave Meta at the end of the year after twelve years at the company and create a business to continue his Advanced Machine Intelligence program. His message stated the goal: systems that understand the physical world, have persistent memory, reason, and plan complex action sequences.

LeCun did not say he was leaving because of a dispute over language models. He thanked Meta’s leaders for their support, said Meta would partner with the new company, and explained that an independent entity could maximize the program’s broad impact. Turning a plausible scientific difference into the cause of his departure would put a motive in his mouth that the announcement does not state.

The news does support a more durable question: what would a “world model” have to demonstrate for the term to describe a capability rather than a brand? The answer requires separating a research thesis, an architecture, a demonstration, and a product. Jumping from the first to the last turns a scientific ambition into a result that does not exist.

A world model predicts useful states

The vision Meta and LeCun published in 2022 describes a world model as an internal simulator of the part of the environment relevant to a task. It receives a representation of the current state, fills in missing information, and predicts plausible future states. It can also estimate what would happen under a proposed sequence of actions.

It need not reconstruct every pixel. The future contains unpredictable, irrelevant detail, such as the position of each leaf while a vehicle approaches an intersection. The JEPA family attempts prediction in an abstract representation space, retaining useful dependencies and discarding some detail that does not help the decision. That choice reduces work but creates an audit question: did the model remove noise or a decisive signal?

“Understanding” cannot be verified by asking a system whether it understands. Design an intervention and test its prediction. If a robot pushes an object from different positions, the model should anticipate outcomes consistent with each action, distinguish alternatives, and express uncertainty when evidence is missing. Predicting one likely next frame without conditioning on action may capture appearance or video regularities without supporting control.

Prediction is not yet planning

Planning adds a loop. An actor proposes action sequences; the model estimates consequences; a cost function scores which future approaches the goal while respecting constraints; one action is executed; and the system observes again. The process resembles model-predictive control because the plan is revised as the world reveals information.

This separation locates failures. Perception can misrepresent the scene; the predictor can get dynamics wrong; the objective can reward a dangerous shortcut; the planner can explore too few alternatives; or the actuator can execute a correct command badly. Calling the final result “reasoning” hides which component worked.

Persistent memory is another problem. Holding state through a short maneuver is not equivalent to retaining facts for days, updating facts that change, and forgetting private information when required. LeCun’s announcement listed memory as a company goal; it did not claim that an available architecture had solved writing, retrieval, expiry, and control.

What the JEPA line had demonstrated before the announcement

In June 2025, Meta researchers including LeCun published V-JEPA 2. The model was pretrained on more than one million hours of video and images. An action-conditioned variant then used fewer than 62 hours of robot video and planned movements for Franka arms in two laboratories.

The work demonstrated reaching, picking, and placing with goal images, without collecting training data on the deployment robots and environments and without task-specific training or reward. That is concrete evidence of transfer and physical planning. It also bounds the achievement: particular arms, constrained manipulations, and short-horizon visual goals are not general physical understanding, persistent memory, or open-ended planning over weeks.

The study itself shows that world models and language models need not be total opposites. V-JEPA 2 was aligned with a language model for several video question-answering tests. A system can use predictive representations for dynamics, language for instructions and knowledge, and separate modules for memory, cost, and action. The useful question is not which acronym wins, but which component supplies evidence for each capability.

The intervention test separates video patterns from causality

A model may predict what happens next because it recognizes a sequence repeated in its data. To test something closer to usable dynamics, vary an action while holding the rest stable: push left or right, hold or release, block a route or leave it open. Then test whether the prediction changes in the appropriate direction.

Evaluation should include states not observed together during training, new objects, different cameras, occlusion, and disturbances. It should also compare against simple rules and models without action input. If a straight-line extrapolation, the last frame, or a background cue achieves nearly the same score, the benchmark does not establish the ambitious capability suggested by its name.

Uncertainty belongs in the answer. When several outcomes are plausible, a predictor that emits one confident result can produce a brittle plan. Coverage and calibration should be measured: how frequently reality falls within the predicted set and whether outcomes assigned a probability occur at a similar rate. For action, knowing that the model does not know can be more valuable than a crisp image.

An evidence ladder for any “world model”

The first rung is representation: recognizing objects, motion, and state in tests designed to prevent shortcuts. The second is temporal prediction on held-out data. The third is action-conditioned prediction, distinguishing futures when an intervention changes. The fourth is planning, choosing actions that reach a goal. The fifth is transfer, preserving performance with unseen environments, objects, or robots.

Memory and safety sit above them. Memory must demonstrate updating, provenance, retrieval, and deletion. Safety must test constraints under distribution shift, detect states outside experience, and stop when uncertainty exceeds a limit. An average success rate is insufficient when a rare failure can cause a collision or irreversible decision.

For each rung, record the task, training data, held-out data, allowed actions, horizon, baselines, success rate, variation, and failure mode. Also record how much information the system receives: a sequence of human-selected subgoals makes a task substantially easier than one final instruction. Without that context, two demonstrations that look alike can measure different abilities.

What the departure confirms and what it leaves open

The primary announcement confirms a departure at the end of 2025, creation of a company, continuation of the AMI program, four technical goals, and a future Meta partnership. It did not provide a final name, funding, product, launch date, or conflict with Meta as the cause. Those gaps should not be filled with rumors or persuasive inference.

The scientific bet already had a described architecture and a measurable robotic demonstration. That makes it more than a slogan but less than the broad intelligence it seeks. Readers can keep one test for the next announcement: ask which state the model predicts, under which action, over what horizon, with what uncertainty, and whether the plan works outside training environments. If those questions lack answers, “world model” names an ambition, not a demonstrated capability.

This article was produced with artificial intelligence under human editorial oversight.

Share this article

This website uses cookies to improve the browsing experience. Cookie policy.

↑↓ navigate ↵ open esc close