Black Forest Labs' visual understanding models require only a few hours of fine-tuning data to adapt to a specific robotic task, dramatically reducing deployment friction.
Snapshot · All-In with Chamath, Jason, Sacks & Friedberg
Black Forest Labs' visual understanding models require only a few hours of fine-tuning data to adapt to a specific robotic task, dramatically reducing deployment friction.
Where this was said
At 57:27 · chapter starts 47:31
Rombach steps back from the film discussion to articulate the deeper architectural insight: the model that makes a movie and the model that drives a robot are not two different technologies — they are the same model. Pre-training on video at scale gives a model implicit understanding of physics, causality, and spatial relationships — the foundations of intelligent physical action. From that base, only a few hours of task-specific fine-tuning data are needed to deploy the model on a specific robot in a factory setting. The long-term goal is to make even that fine-tuning unnecessary, enabling in-context robot instruction — just tell a robot in natural language what to do, the same way you prompt a language model. Rombach acknowledges this is still a research problem, but the trajectory is clear: world models, action models, and generative video models are all converging into the same underlying architecture.
Robin Rombach sat with Martin Scorsese multiple times to demonstrate Black Forest Labs' generative tools. What captivated Scorsese wasn't automation — it was the ability to take a visual scene living in his imagination and externalize it for his team to iterate on. Language is lossy. Images are not.
Martin Scorsese worked directly with Robin Rombach to use Black Forest Labs' generative models to visualize pre-production scene concepts for a potential new film.
A Bitcoin movie starring Gal Gadot was filmed entirely on a sound stage, with all scenery generated by AI in post. The result: a $30M production that would have cost $150M with traditional set builds. It never would have been greenlit at $150M — generative AI didn't just cut costs, it made the film possible.
A Bitcoin movie starring Gal Gadot was made for $30M using generative AI for all scenery — it would have cost $150M with traditional set builds and might never have been greenlit.
The most underappreciated insight in AI right now: a single multimodal generative model can produce a movie and act as the perception and action brain for a physical robot. Pre-training on video gives implicit understanding of real-world physics — which transfers directly into robotic action prediction.
Ad-based monetization works well for game apps where users spend extended time in-session, as seen with Grid and Wordle.
Tool-focused apps like PuffCount are poor candidates for ad monetization because users don't stay in-session long enough.
A hard paywall is a screen that blocks all app features unless the user pays or starts a free trial — it cannot be dismissed.
Mobile apps are primarily monetized through either ads (best for games) or in-app purchases/subscriptions (best for tools).
According to the episode, YouTube outperforms every other social platform for building trust and driving SaaS conversions.
Vasco stated that the majority of his app's user base came directly from his YouTube channel.
SEO Bot features a 'Boost My Domain Rating' button that routes users directly to Listing Bot, an example of in-product cross-selling.
The founder's entire product portfolio is AI-related, making it easier to package products attractively for directories.
The founder attached their SaaS demo to the trending debate about whether AI coding is actually good enough to build a full SaaS product.
We use essential and analytics cookies to run Vuci. To understand how the site is used: Privacy Policy.
Install Vuci on your phone
Add it to your home screen for a faster, app-like experience.
Install Vuci on your phone
Tap the Share button, then “Add to Home Screen”.
A new version is available
Reload to get the latest Vuci.