Where this was said
Comparing human vs AI sample efficiency
At 5:22 · chapter starts 3:11
Three objections addressed: (1) evolution pre-training humans — debunked via genome size; (2) multimodal sensory data — debunked via blind/deaf intelligence; and (3) scaling up models — addressed next. [1] — Dwarkesh Patel "The common objection — that billions of years of evolution pre-trained humans, making data comparisons unfair — doesn't hold up. The human …" 04:43 [2] — Dwarkesh Patel "If multimodal sensory data were the secret ingredient behind human intelligence, blind and deaf people would lack general intelligence. The…" 05:48
Frontier AI models are trained on tens to hundreds of trillions of tokens, versus roughly 200 million tokens a human sees from birth to adulthood — nearly a million-fold difference.
Frontier AI models are trained on tens to hundreds of trillions of tokens. A human sees about 200 million from birth to adulthood. That's close to a million-fold difference — and it reveals just how data-hungry these systems really are.
Humans can learn to teleoperate any humanoid or robot arm within hours, but AI systems require millions of hours of demonstrations and still can't perform complex open-ended tasks.
A teenager can learn to drive a car with about 20 hours of practice, yet self-driving car models from Waymo and Tesla require 3–4 orders of magnitude more data.
A teenager learns to drive in 20 hours. Even accounting for 16 years of growing up and building physical intuition, that's still 3–4 orders of magnitude less data than Waymo and Tesla use to train self-driving cars. This gap is the sample efficiency problem in concrete terms.
The common objection — that billions of years of evolution pre-trained humans, making data comparisons unfair — doesn't hold up. The human genome is only 3 GB, and 1–2% is protein-coding. That's nowhere near enough to store pre-trained neural network weights. Evolution found the right hyperparameters; it didn't train the weights.
The human genome is only 3 gigabytes and only 1–2% is protein-coding, which Dwarkesh argues is not enough space to store pre-trained neural network weights from evolution.
If multimodal sensory data were the secret ingredient behind human intelligence, blind and deaf people would lack general intelligence. They don't. This suggests billions of sensory tokens aren't the key — and may mean the human-AI data gap is even larger than estimated.