IA 360
Current Affairs

Meta unveils Movie Gen, its model for creating video and sound

Meta has unveiled Movie Gen, a family of models that generates videos with synchronized audio and lets users edit footage through text prompts. The company presents it as a research advance and has not released it as a public product.

5 min read AI-generated Leer en español
Meta unveils Movie Gen, its model for creating video and sound

On October 4, 2024, Meta introduced Movie Gen, a research model family for generating and editing video, personalising it from an image and producing synchronised audio. Meta’s paper describes architecture, data, evaluations and limits; it did not announce a public product. That distinction determines what readers can test and what remains a laboratory demonstration.

The announcement matters because it brings together two capabilities that have so far usually arrived separately: moving-image generation and audio generation. Rather than producing a silent video and then requiring another tool to add music, effects or ambient sound, Movie Gen aims to build both layers in coordination.

Up to 16 seconds of high-definition video

The family’s main component is Movie Gen Video, a model with 30 billion parameters. Parameters are the internal values a model adjusts during training to learn patterns in images, text and sound. Meta says it can generate clips of up to 16 seconds at 1080p resolution and a frame rate of 16 frames per second. Source

The company showcases text-generated scenes featuring people, animals and complex environments, as well as sequences in which characters perform actions described in the prompt. As with other video generators, the challenge is not just producing an attractive image: it is maintaining subjects’ identities, object consistency and plausible physics throughout the sequence.

Movie Gen is entering a race that already includes OpenAI with Sora, Runway with Gen-3 Alpha, Luma AI and Chinese companies such as Kling. Video generation has become one of the most visible fronts in generative AI because it requires combining language understanding, visual composition, motion and temporal continuity.

Sound is no longer an external add-on

Meta has also developed Movie Gen Audio, a model with 13 billion parameters designed to generate soundtracks, effects and ambient sounds from text, video or both. It can, for example, pair the sound of an engine with a moving vehicle, create a musical backing track for a scene or add ambient sounds that fit the setting shown.

The key detail is synchronization. In conventional video, sound is usually added during post-production and requires human decisions: when music starts, which effect should accompany an action or how it should be mixed with the ambient sound. Meta’s proposal attempts to automate the relationship between what happens on screen and what viewers hear.

Even so, generating plausible sound is not the same as fully solving dubbing or character dialogue. Human speech requires lip-sync, performance, continuity across shots and control over language—all areas that remain particularly challenging for generative tools.

Edit a video with a single sentence

Movie Gen does not only work from a blank screen. Meta has shown features for editing existing footage through written instructions. The system can add an object to a scene, replace the background, change a person’s appearance or modify specific elements without manually reconstructing every shot.

It also includes a video personalization feature: given an image of a person, the model can place them in a generated scene while preserving their features. This capability is appealing for personalized content, advertising and audiovisual creativity, but it raises familiar concerns around consent and impersonation. The easier it becomes to produce a realistic sequence featuring someone’s likeness, the more important it will be for platforms and creators to clearly indicate when an image has been generated or altered by AI.

A research advance, not an available application

Meta has not made Movie Gen generally available to the public. The company presents Movie Gen as research and has not announced its availability as a public product.

That caution has a practical reason. Generated video could be used to preview campaigns, create educational material, prototype scenes or produce social media content with fewer resources. But the same techniques lower the cost of making deceptive videos, imitating real people or reusing works and styles without authorization.

The decisive comparison will not be limited to the quality of demonstration clips. Movie Gen will have to show that it offers enough control for real-world production: consistency across scenes, precise editing, rights to the materials used and effective mechanisms for identifying synthetic content. Meta has shown ambitious technology; what remains to be seen is under what conditions it decides to turn it into an accessible tool.

A generated video is a sequence, not a long image

Temporal quality requires identity, geometry, lighting and causality to persist across frames. A beautiful still can hide changing hands, appearing objects or discontinuous motion. Testing selects actions with beginnings and endings, tracks small elements and records when coherence breaks.

Audio adds another synchronisation problem. Separate speech, effects, ambience and music, then check whether sound matches the visible event and preserves perspective. “Generates sound” does not say whether a system understands an action or merely associates common patterns. Unlikely cases and deliberate silence reveal the difference more clearly.

Editing, generation and personalisation are different tasks

For editing, the central criterion is changing what was requested while preserving everything else. Keep the original video, a mask of the target region and a list of invariants. Score instruction compliance together with preservation. A spectacular transformation that rewrites the whole scene is not precise editing.

Personalisation needs consent, provenance and deletion controls as well as resemblance. Ask who may upload an image, how it is stored, whether output is marked and which remedy the depicted person has. Better identity fidelity increases creative ability and impersonation risk at the same time.

A vendor benchmark is a hypothesis

Human comparisons depend on example selection, interface, order and question. Repetition requires all attempts, exact models and criteria to be published. Failures should be evaluated, not just chosen pairs. When the model is unavailable, readers can audit method and published material but cannot verify general superiority.

Technical provenance and a visible label serve different functions. A credential may record origin and edits; a label informs within a platform. Both can disappear after cropping, recompression or republication. An audit follows the file through that route and does not present one signal as an absolute guarantee.

The transferable skill is to evaluate generative video across separate axes: prompt fidelity, temporal coherence, synchronisation, editing preservation, consent and provenance. That record can compare the next model even when brand and demonstration change.

A minimum record accompanies every example

Preserve prompt, seed when available, duration, resolution, aspect ratio, version and every attempt. A montage containing the best of many runs does not share conditions with a random sample. Generation time and cost also determine whether a capability supports idea exploration or production at scale.

Human evaluation needs separate questions. “Which do you prefer?” mixes sharpness, motion, aesthetics and compliance. Asking for one decision per axis and permitting ties produces a diagnosis that can guide improvement. Reviewers should see alternating order to reduce position effects.

Finally, document data and rights as far as the paper permits. If the set cannot be inspected, do not invent its composition. State the limit and avoid inferring from an output that every training source was authorised.

This article was produced with artificial intelligence under human editorial oversight.

Share this article

This website uses cookies to improve the browsing experience. Cookie policy.

↑↓ navigate ↵ open esc close