Dr. Fei-Fei Li co-founded WorldLabs at the beginning of 2024 to build spatial intelligence foundation models capable of generating 3D and 4D worlds.
Snapshot · Huberman Lab
Dr. Fei-Fei Li co-founded WorldLabs at the beginning of 2024 to build spatial intelligence foundation models capable of generating 3D and 4D worlds.
Where this was said
At 1:51:30 · chapter starts 1:50:10
Huberman invites Fei-Fei Li to describe WorldLabs, and she frames it as the culmination of her life's work. Language AI has transformed information access, but humans evolved in a spatial, three-dimensional physical world — and a corresponding AI capability has been missing. WorldLabs is building foundation models for spatial intelligence that can generate 3D and 4D environments from text prompts, images, or sketches [1] — Fei-Fei Li "Language AI is powerful, but humans evolved in a spatial, physical world. WorldLabs is building foundational models for spatial and 3D inte…" 1:51:20 . The use cases are broad: entertainment companies can prototype environments, architects can visualize spaces, robotics teams can train in synthetic worlds, and healthcare providers can build interactive simulation environments. She describes the company as model-first and PhD-heavy, currently transitioning into product development. She also reflects on her identity as a builder and immigrant who finds deep meaning in rolling up her sleeves with a young, brilliant team to build something from scratch.
Language AI is powerful, but humans evolved in a spatial, physical world. WorldLabs is building foundational models for spatial and 3D intelligence — letting people generate entire environments from a sentence or sketch. The applications span filmmaking, robotics training, architecture, and healthcare.
There are more than 90,000 Flock surveillance cameras currently in use around the United States.
A 2023 report estimated that 10 million Americans own Ring cameras, roughly 1 in 5 households having a video-enabled doorbell.
The ImageNet dataset collected 15 million images to drive machine learning, becoming a cornerstone of the modern AI revolution.
A Stanford graduate student benchmarked human performance on the ImageNet 1,000-category challenge at roughly 4% error rate, a figure AI surpassed by 2016.
From the 2012 ImageNet breakthrough, it took only about 3–4 more years for AI algorithms to surpass human performance in naming 1,000 object categories.
Flock's surveillance network scans more than 20 billion license plates per month across the United States.
OpenAI's Sora, released in January 2024, demonstrated AI's ability to generate realistic video from text prompts, marking a key milestone in video generation.
AlphaGo's Move 37 against Lee Sedol was a move that human Go masters had never considered, illustrating a unique form of AI creativity within constrained mathematical rules.
The internet is not a random data source — it is the largest-ever multimodal archive of human behavior including text, images, video, and audio accumulated over decades.
We use essential and analytics cookies to run Vuci. To understand how the site is used: Privacy Policy.
Install Vuci on your phone
Add it to your home screen for a faster, app-like experience.
Install Vuci on your phone
Tap the Share button, then “Add to Home Screen”.
A new version is available
Reload to get the latest Vuci.