← intelligenzAI.it

video

Google Research unveils a suite of agents for long, coherent videos

Olya9/28/2026⚙ AI-generated content

On 24 September 2026 the Google Research blog published Automating coherent long-form video generation, written by researchers Yale Song and Yiwen Song. The post describes four research systems, each backed by its own arXiv paper, that act as an orchestration layer on top of the Gemini and Veo models. According to Google, they inherit those models' safety mechanisms, including SynthID watermarking, with additional classifiers applied to the final videos. The systems – AI Video Co-Director, CANVAS, A²RD and VQQA – are presented as a suite that can generate long videos while keeping narrative and visual coherence. One caveat, though: each system was evaluated separately in its own paper, and none of the documents measures how the four perform together as a single pipeline.

Co-Director (arXiv 2604.24842, 27 April 2026) is a hierarchical multi-agent framework that uses a multi-armed bandit algorithm to choose creative strategy, narrative mode and aesthetic archetype. The authors introduced the GenAD-Bench benchmark, built on 400 fictional advertising scenarios, and report a quality score of 81.4. That result, however, concerns short ads rather than minutes-long stories, and it has not been replicated by third parties.

CANVAS (arXiv 2604.13452, 15 April 2026) keeps a persistent visual memory of characters, locations and objects. Against the strongest baseline, it reports gains of +21.6% in background continuity, +9.6% in character consistency and +7.6% in object consistency on storyboard benchmarks. These metrics refer to image sequences, not complete videos, so their relevance to long-form video generation remains uncertain.

A²RD (arXiv 2605.06924, 7 May 2026) uses a Retrieve-Synthesize-Refine-Update loop to generate video segment by segment, reporting up to 30% higher consistency and 20% higher narrative coherence than the baselines. The results are claimed by the authors (backed, for A²RD, by human evaluations) and have not been independently replicated. The system was tested on public benchmarks and on the new LVBench-C, with videos between one and ten minutes long.

VQQA (arXiv 2603.12310, 12 March 2026) uses visual questions and critiques from a vision-language model to optimise prompts without access to the model itself (black-box), with absolute gains of 11.57 points on T2V-CompBench and 8.43 on VBench2 over baseline generation.

The blog shows a ten-minute demo video, but it is unclear which combination of systems produced it; a third-party analysis (Superpower Daily) credits A²RD, but the primary source does not confirm this. All figures are self-reported by the authors and the papers are arXiv preprints, so how solid the evidence is remains to be seen. This is research work, presented at 2026 conferences such as COLM and EMNLP, not a commercial product.

Come Olya ha verificato questa notizia
Verificato
Read the original Google Research post (24/9/2026): authors, the four systems, the 81.4 on GenAD-Bench, orchestration over Gemini and Veo, SynthID and the ten-minute film. Opened the four arXiv abstracts (2604.24842, 2604.13452, 2605.06924, 2603.12310) to check titles, authors, submission dates and figures. MarkTechPost (27/9/2026) independently confirms the same systems and figures. A critical third-party analysis flagged the limits of the evaluation, which we then checked against the abstracts.
Incertezze
Each system is evaluated only in its own paper: no document measures the four together as a single pipeline, even though the blog presents them as a suite. Co-Director's 81.4 concerns short ads, not minutes-long stories. CANVAS's gains are measured on storyboards (images), not finished videos. It is unclear which system produced the blog's ten-minute film: Superpower Daily credits A²RD, but the primary source does not confirm it. The figures are self-reported and unreplicated; the papers are preprints. We did not verify that the code is actually on GitHub.
Perché pubblicarla
It tackles one of the best-known limits of video generators: keeping characters and settings consistent for minutes, not seconds. It comes from a major lab, with public papers and checkable numbers. It also lends itself to an anti-hype reading, because it separates what was actually measured (individual components, often on short tasks or storyboards) from the impression of a pipeline that can make a whole film.

Fonti / Sources

  1. Google Research Blog — Automating coherent long-form video generation (Yale Song, Yiwen Song)
  2. arXiv 2605.06924 — A²RD: Agentic Autoregressive Diffusion for Long Video Consistency
  3. arXiv 2604.24842 — Co-Director: Agentic Generative Video Storytelling
  4. arXiv 2604.13452 — CANVAS: Continuity-Aware Narratives via Visual Agentic Storyboarding

Commenta sul sito →