Quote · All-In with Chamath, Jason, Sacks & Friedberg
Open Source Wins, AGI Is Here, and Scorsese's AI Toolkit with CEOs of Cerebras & Black Forest Labs
Where this was said
Martin Scorsese, Robots, and the Future of Hollywood IP
At 49:58 · chapter starts 47:31
Rombach steps back from the film discussion to articulate the deeper architectural insight: the model that makes a movie and the model that drives a robot are not two different technologies — they are the same model. Pre-training on video at scale gives a model implicit understanding of physics, causality, and spatial relationships — the foundations of intelligent physical action. From that base, only a few hours of task-specific fine-tuning data are needed to deploy the model on a specific robot in a factory setting. The long-term goal is to make even that fine-tuning unnecessary, enabling in-context robot instruction — just tell a robot in natural language what to do, the same way you prompt a language model. Rombach acknowledges this is still a research problem, but the trajectory is clear: world models, action models, and generative video models are all converging into the same underlying architecture.
Robin Rombach sat with Martin Scorsese multiple times to demonstrate Black Forest Labs' generative tools. What captivated Scorsese wasn't automation — it was the ability to take a visual scene living in his imagination and externalize it for his team to iterate on. Language is lossy. Images are not.
Martin Scorsese worked directly with Robin Rombach to use Black Forest Labs' generative models to visualize pre-production scene concepts for a potential new film.
A Bitcoin movie starring Gal Gadot was filmed entirely on a sound stage, with all scenery generated by AI in post. The result: a $30M production that would have cost $150M with traditional set builds. It never would have been greenlit at $150M — generative AI didn't just cut costs, it made the film possible.
A Bitcoin movie starring Gal Gadot was made for $30M using generative AI for all scenery — it would have cost $150M with traditional set builds and might never have been greenlit.
The most underappreciated insight in AI right now: a single multimodal generative model can produce a movie and act as the perception and action brain for a physical robot. Pre-training on video gives implicit understanding of real-world physics — which transfers directly into robotic action prediction.
Black Forest Labs' visual understanding models require only a few hours of fine-tuning data to adapt to a specific robotic task, dramatically reducing deployment friction.