Dwarkesh Podcast

Quote · Dwarkesh Podcast

Eric Jang – Building AlphaGo from scratch

Explore episode May 15, 2026

Where this was said

Off-policy training

At 2:03:20 · chapter starts 2:01:09

Eric explains the DAGGER-inspired view of replay buffers: off-policy data is helpful if it covers reachable states, harmful if it covers states the current policy would never visit.

Similar quotes