Where this was said
What is really driving AI progress?
At 2:25 · chapter starts 0:00
Dwarkesh argues that AI's main progress driver is data volume and quality, not algorithmic breakthroughs. RL is framed as synthetic data generation. Expert data labeling is described as a billion-dollar industry. [1] — Dwarkesh Patel "AI capabilities look like a galaxy of stars, but at the center is an invisible black hole of data. The main way AIs have gotten better is n…"
AI capabilities look like a galaxy of stars, but at the center is an invisible black hole of data. The main way AIs have gotten better is not through better architectures or training tricks — it's by adding more and better data and scaling compute to generate it.
The industry producing expert labels and RL training environments is already earning billions per year in revenue, soon to be tens of billions.
AI models aren't learning the way humans do. They're more like Frankenstein's monster — stitched together from a billion carefully constructed example graphs. That's the uncomfortable truth behind what looks like fluid, general intelligence.
With GRPO, AI models generate hundreds to thousands of rollouts per task to solve the credit assignment problem, far more than the one or two times a human student might practice a problem.
Open models trail frontier models by only 4 months. The reason is that data — the actual driver of progress — can be distilled from public APIs. Hyperparameters and training tricks can't be. If architecture were the real edge, the gap would be far larger.
Epoch AI reported that open models lag state-of-the-art frontier models by only 4 months, which Dwarkesh attributes to data being the real driver — easily distilled from public APIs.