The most underappreciated insight in AI right now: a single multimodal generative model can produce a movie and act as the perception and action brain for a physical robot. Pre-training on video gives implicit understanding of real-world physics — which transfers directly into robotic action prediction.