IA 360
Current Affairs

OpenAI unveils Sora, a model that creates videos up to one minute

Sora generates videos up to one minute, while OpenAI acknowledges physics and continuity failures. Checking provenance, source, context and content is stronger than hunting for one visual flaw.

4 min read AI-generated Leer en español
OpenAI unveils Sora, a model that creates videos up to one minute

OpenAI introduced Sora on February 15, 2024 as a model capable of generating videos up to one minute long from text instructions. Demonstrations showed camera movement, multiple characters and detailed visual styles, but the system was not released publicly. The company gave access to specialists tasked with finding risks and to creative professionals providing feedback.

The longer duration matters because video must preserve objects, actions and scenes over time. It also weakens a popular habit: deciding whether footage is authentic by hunting for strange hands, flicker or impossible motion. Those flaws can disappear, while genuine footage can contain compression, blur and editing. The durable skill is verifying audiovisual material as a chain of evidence.

What Sora demonstrated — and what it did not

In its original Sora presentation, OpenAI said the model could produce video up to a minute long while maintaining visual quality and adherence to a prompt. It could also animate an image, extend a video forwards or backwards and fill missing frames. “Up to” is a declared maximum capability; it does not mean every request reaches 60 seconds or every second remains coherent.

The same page listed failures. Sora could model physics incorrectly, confuse left and right, make people or objects appear spontaneously, deform rigid items and lose the relationship between an action and its consequence. The example of a person biting a cookie without leaving a mark captures the problem: the model creates a plausible visual sequence without necessarily preserving the world’s causal state.

OpenAI published selected examples and a qualitative evaluation, not a success rate across a complete prompt set. Its technical report on video models also said it omitted model and implementation details. The scenes provide evidence of possibilities and limitations, but not their frequency. A demonstration proves possibility, not a reliability distribution.

How visual data becomes a sequence

Sora is a diffusion transformer. During generation it starts from a noisy representation and learns to predict cleaner visual units. First, a network compresses images and videos into a lower-dimensional latent space. That representation is divided into spacetime patches that play a role similar to tokens in a language model.

Patches make it possible to train on videos and images with different durations, resolutions and aspect ratios. The report said the same model could produce widescreen, vertical and other dimensions. OpenAI also used expanded descriptions of training videos to improve correspondence between an instruction and the output.

This architecture helps explain continuity, but it is not a physics engine. Learning regularities from many videos can produce convincing shadows, trajectories and perspective. It does not ensure that mass, rigidity, cause or identity exist as stable variables. “World simulator” was OpenAI’s interpretation of a research direction, not a certificate of physical accuracy.

Provenance is not truth

OpenAI said it was testing a classifier for Sora-generated video and planned to include C2PA metadata if the model entered a product. C2PA provides something different from a detector: it attaches claims about creation and edits, binds them to a file and signs them digitally. A verifier can check that a statement came from a signer and was not altered without detection.

The C2PA 2.0 specification, published in January 2024, is explicit about scope: it validates that assertions are attached to an asset, properly formed and free from tampering; it does not judge whether declared provenance is “good” or “bad.” A valid credential might say a tool generated the file. It does not prove the depicted scene happened.

Missing credentials do not establish that footage is real either. Metadata may never have been created, a platform can strip it during conversion or someone can screen-record the result into a new file. There are at least four states: valid credential, invalid or broken credential, credential declaring generation or editing, and no credential. Only the first confirms a signed chain; none turns content into truth by itself.

The four-layer protocol

The first layer is provenance. Obtain the file closest to the original, inspect its credential with a compatible tool and record signer, date, declared actions and validation state. A social-media capture is weaker than a file supplied by the person who claims to have recorded it. Preserve an untouched copy and conduct analysis on duplicates.

The second is source. Find the earliest publisher, not merely an account that reposted it. Check whether the account, domain or person has a verifiable relationship to the location and offers additional material. Request the original file, other angles or the sequence before and after. An identified source can still lie, but its responsibility, history and access can be examined.

The third is context. Split the claim into location, date, participants and action. Compare each element with independent material: official notices, schedules, maps, archived weather or other recordings, depending on the event. Authentic footage can be assigned the wrong date or described as another incident; finding an original does not validate the accompanying caption.

The fourth is content. Inspect shadows, reflections, continuity, anatomy, audio and physics, but treat them as clues that prompt further checks. Extract key frames and search for earlier versions. A visual error can reveal generation or editing; the absence of one does not establish authenticity. Forensic observations are strongest when they converge with provenance, source and context.

A decision scale before sharing

The conclusion should express uncertainty. “Verified” requires a coherent capture or publication chain and corroboration of the event. “Probably authentic” or “probably synthetic” signals strong but incomplete evidence. “Unverified” is correct when the original is missing or corroboration cannot be found. Identifying the exact tool is unnecessary when deciding that a claim lacks sufficient support.

Effort should follow potential harm. A fantastic clip shared as entertainment may need only a clear label. Video accusing a person, affecting an election or depicting an emergency requires pausing distribution, preserving evidence and seeking independent confirmation. Pressure to share quickly does not lower the standard; it raises the cost of error.

Sora showed that synthetic-video duration and coherence could advance together, even as its own examples revealed causal limits. That balance will change with later models. The useful method does not rely on a permanent tell: checking credentials, origin, context and content will still work when pixels look perfect. Faced with striking footage, the strongest question is not “does this look like AI?” but “what evidence chain supports that this happened?”

The record that makes verification reproducible

A useful verification leaves a record another person can inspect. It contains the exact claim attached to the video, discovery URL and time, publishing account, a file copy and its hash, credential result, key frames, corroborating sources and unsuccessful searches. It also separates observation from inference: “the shadow changes direction” is observable; “therefore Sora made it” is a hypothesis.

Finally, record missing evidence and what finding would change the conclusion. If only a recompressed version exists, state that limitation. Two accounts copying one origin are not two corroborations. Independence matters more than quantity: another angle or a document from the location adds more than one hundred reposts. Publishing the record with the conclusion allows correction if the original appears and prevents “unverified” from becoming “false” through repetition.

This article was produced with artificial intelligence under human editorial oversight.

Share this article

This website uses cookies to improve the browsing experience. Cookie policy.

↑↓ navigate ↵ open esc close