Dwarkesh Podcast

Snapshot · Dwarkesh Podcast

The next big breakthrough will be AIs learning on the job

Explore episode Jun 26, 2026

Where this was said

Getting the learning back to the weights

At 13:15 · chapter starts 8:41

Dwarkesh explores on-policy self-distillation (OPSD) as the leading technique for continual learning, contrasting it with naive SFT and explaining why RL's sparse parameter updates are a feature, not a bug.

Technology
On-Policy Self-Distillation (OPSD): Compressing Sessions Into Weights

The next big breakthrough will be AIs learning on the job · Jun 26, 2026 Technology

OPSD trains the base model to match what a veteran 'teacher' model — loaded with a full session's context — would predict. It needs no outer-loop reward, provides denser per-token supervision than RL, and avoids overwriting existing knowledge. It's the most credible path to continual learning today.

Technology
Dreaming: AI Builds Its Own Training Simulations

The next big breakthrough will be AIs learning on the job · Jun 26, 2026 Technology

Dreaming means the model spends compute generating its own simulated RL environments, then trains against them — rehearsing skills relevant to the actual user's real-world context. Like EfficientZero playing dozens of internal games per real step, this could make AI far more sample-efficient without requiring external data.

Similar snapshots