Dwarkesh Podcast

Quote · Dwarkesh Podcast

Eric Jang – Building AlphaGo from scratch

Explore episode May 15, 2026

Where this was said

Self-play

At 1:18:15 · chapter starts 1:00:33

Eric explains how MCTS produces a more confident action distribution than the raw policy, and how AlphaGo trains the policy to imitate that improved distribution — the core self-play loop.

Similar quotes