Dwarkesh Podcast

Snapshot · Dwarkesh Podcast

The next big breakthrough will be AIs learning on the job

Explore episode Jun 26, 2026

Where this was said

Getting the learning back to the weights

At 9:20 · chapter starts 8:41

Dwarkesh explores on-policy self-distillation (OPSD) as the leading technique for continual learning, contrasting it with naive SFT and explaining why RL's sparse parameter updates are a feature, not a bug.

Technology
On-Policy Self-Distillation (OPSD): Compressing Sessions Into Weights

The next big breakthrough will be AIs learning on the job · Jun 26, 2026 Technology

OPSD trains the base model to match what a veteran 'teacher' model — loaded with a full session's context — would predict. It needs no outer-loop reward, provides denser per-token supervision than RL, and avoids overwriting existing knowledge. It's the most credible path to continual learning today.

Technology
Dreaming: AI Builds Its Own Training Simulations

The next big breakthrough will be AIs learning on the job · Jun 26, 2026 Technology

Dreaming means the model spends compute generating its own simulated RL environments, then trains against them — rehearsing skills relevant to the actual user's real-world context. Like EfficientZero playing dozens of internal games per real step, this could make AI far more sample-efficient without requiring external data.

Similar snapshots