If AI can generate ten beautiful shots, does that mean it can deliver a film?

In the first shot, the character wears a black coat. In the third, their face suddenly changes. When we return to the room, the door has moved, too. Each clip is clear on its own, but they’re unusable when stitched together. That’s the costly part of producing long-form video.

Google has scheduled a Co-Director demo at the COLM conference, running October 6–9. Combined with research announced on September 24, it focuses on planning multi-shot stories as a whole. The related CANVAS framework continuously tracks the state of characters, locations, and objects, so a scene that reappears after several shots can still pick up where it left off.

The progress here is in cross-shot consistency, and a conference demo doesn’t mean commercial revenue has been generated. A single generation may be cheaper, but if fixing one item of clothing means redoing half the film, delivery costs could still be high.

What I’m watching for: whether characters remain consistent when they leave and return, whether changes to props make sense for the story, and whether revising an earlier scene breaks a later one. If these checks can be passed reliably, production teams will find it easier to shift budgets from manual fixes to actual production.

For RENDER, FET, and NEAR, which are focused on AI compute and agents, it’s also important to track real-world adoption and delivery performance. There’s no evidence here that they’re involved in Google’s project.

The second image is a reference photo of a Google building, not from the conference.

$RENDER $FET $NEAR

Tap my profile picture to view my live trading account and trade calls