Dwarkesh Podcast

Snapshot · Dwarkesh Podcast

Eric Jang – Building AlphaGo from scratch

Explore episode May 15, 2026

Where this was said

RL is even more information inefficient than you thought

At 2:14:18 · chapter starts 2:08:56

Dwarkesh presents an information-theoretic argument: naive RL learns near-zero bits per sample at low pass rates, while supervised learning learns negative log(pass rate) bits. The gap is enormous.

Similar snapshots