,

Google’s long-form video stack, scored on benchmarks it built itself

Getting a generative video model to hold a character’s face, costume and surroundings steady across several minutes remains largely unsolved. On 24 September, Google researchers Yale Song and Yiwen Song described a set of frameworks meant to address it by treating the problem as planning rather than generation.

The approach is an orchestration layer sitting on top of Gemini and Veo rather than a new video model. Four pieces make it up: Co-Director, which tracks a world state across a whole narrative; CANVAS, which keeps a structured visual memory of characters, locations and object states; A²RD, which stitches minutes-long output together segment by segment; and VQQA, which uses a vision-language model to critique drafts and feed back changes. Co-Director and CANVAS are due at COLM 2026 and EMNLP 2026 respectively. Because it wraps Gemini and Veo, it inherits SynthID watermarking.

On the scores, read carefully. The headline is a peak quality score of 81.4 on GenAD-Bench, but GenAD-Bench, HardContinuityBench and LVBench-C were all built by the same team for this work, and the post gives no baseline figure beside the 81.4. There is no mention of released code or weights.


Related