IA 360
Current Affairs

Apple explores how AI can keep track during image editing

MT-EditFlow trains image-editing models to chain instructions without degrading content the user did not ask to change. Its results are promising, but rely on a specific evaluation.

4 min read AI-generated Leer en español
Apple explores how AI can keep track during image editing

On June 1, 2026, a paper involving Apple proposed training image-editing models for a chain of decisions rather than a single isolated command. Editing an image with natural language rarely ends with one instruction: a person may ask to remove an object, then add text, and finally change the material of a particular area.

The method, called MT-EditFlow, applies reinforcement learning to flow-matching image-editing models. Its goal is not only to make an image follow the latest instruction, but to preserve content that the user did not ask to change across several turns. That distinction matters: an editor may correctly recolor a wall while altering a person, an object or the style the user meant to keep.

The problem of editing one’s own output

Many editors are trained on image-and-instruction pairs: they receive an original image, apply one change and are evaluated once. In a real conversation, the second change is made to the first output. The paper calls this an all-or-nothing requirement: a single deviation can spoil the entire sequence. Exposure bias also appears, because the model must work from its own imperfect outputs rather than the ideal examples seen in training.

MT-EditFlow changes the training signal to observe the full trajectory. It uses two rewards: instruction following, meaning whether the requested change was made at each turn; and content consistency, meaning whether regions not targeted by an edit remain faithful to the reference image. The system aggregates those signals and broadcasts the resulting advantage to every step in the chain, so local decisions are linked to global success.

The researchers also compared ways of scoring. A strict evaluation that grants success only if every turn is correct resembles the user experience, but leaves very few positive signals for training. As an alternative, they use more graded per-turn scores. The aim is for a model to learn from a partly correct sequence without rewarding it for completing one instruction at the expense of the rest of the image.

Results after three turns

The paper evaluates the method on EdiVal-Bench, a sequential editing set, and uses 2,500 instruction chains for training. In its main table, FLUX.1-Kontext-dev rises from 38.71 to 45.56 on the third-turn overall score with the MT-EditFlow-NFT variant: a 6.85-point difference. The authors place that result above Qwen-Image-Edit, which scores 41.93 in the same column.

It is a benchmark result, not a promise that every image will be edited reliably. The instruction-following evaluation uses Qwen3-VL-8B as a visual reward model, while content consistency is measured with EdiVal-CC. The authors themselves note that vision-language models can fail at spatial reasoning, detecting small details, or judging artifacts and aesthetics.

What counts as a good edit

One revealing decision in the paper is to leave pure visual quality out of the reward. The authors argue that an aesthetically polished image can still be a bad edit if it changes the identity, style or elements that were meant to remain. They therefore prioritize instruction compliance and preservation of unedited content. That is a sensible choice for studying interactive editing, but it does not cover everything a user may value in a final image.

The most interesting advance is not a particular visual effect, but treating an editing conversation as a sequence. Image systems are increasingly used as iterative collaborators: they try something, receive a correction and try again. For that interaction to be useful, executing isolated commands is not enough; a model must retain the thread of what has changed and what was deliberately left untouched. MT-EditFlow offers an experimental way to train that short-range visual memory.

Sources for this piece

This piece draws on 3 primary source(s), gathered during reporting.

This article was produced with artificial intelligence under human editorial oversight.

Share this article

This website uses cookies to improve the browsing experience. Cookie policy.

↑↓ navigate ↵ open esc close