Back to feed

Automating coherent long-form video generation

The latest research from Google

Sep 24, 2026

9/24/2026

Continuity In Visual Narratives Depends On Stateful Retrieval And A Durable Asset Registry To Maintain Identity, Spatial Geometry, And Object States Across Scenes

Automating coherent long-form video generation · The latest research from Google

Science, Technology & Innovation · Sep 24, 2026

CANVAS preserves visual continuity by maintaining a persistent, retrievable world state for characters, locations, objects, identities, geometry, and changing object conditions across both nearby and widely separated scenes. Its museum-heist example contrasts failures in Gemini-3.1-Pro and AutoStudio with CANVAS’s claimed ability to keep the thief, exhibit hall, and gemstone consistent, supporting the principle that serialized production needs a durable asset/state registry rather than relying only on prompts or larger context windows.


9/24/2026

Long-Form Video Generation Should Be Framed As Global Optimization With An Orchestration Layer And Explicit Feedback Loops.

Automating coherent long-form video generation · The latest research from Google

Science, Technology & Innovation · Sep 24, 2026

Google’s AI video co-director treats long-form generation as a global optimization problem: a bandit selects shared creative, narrative, and aesthetic settings, while a multimodal judge evaluates the final video and feeds back structured rewards. This reduces semantic drift and improves consistency, suggesting that orchestration and evaluation layers may matter as much as the underlying video model.


9/24/2026

Memory-Guided Control Balances Progression And Consistency In Long-Horizon Video Synthesis With Extrapolation And Interpolation

Automating coherent long-form video generation · The latest research from Google

Science, Technology & Innovation · Sep 24, 2026

A²RD is a long-horizon video generation system that preserves character and environment consistency by using multimodal memory and switching between extrapolation for story progression and interpolation for reanchoring returning elements; its results suggest that minute-scale video requires memory-guided control rather than simply longer raw clips.


9/24/2026

VQQA Applies Prompt-Optimization Feedback With Global Selection To Improve Video Quality While Preserving The Original Requirements

Automating coherent long-form video generation · The latest research from Google

Science, Technology & Innovation · Sep 24, 2026

VQQA improves generated videos through a black-box loop that uses VLM-generated, defect-specific feedback to revise prompts, while Global Selection preserves fidelity to the original request by choosing the best candidate across all iterations; Google reports gains on several benchmarks without numerical margins.