Dwarkesh Podcast

Snapshot · Dwarkesh Podcast

The next big breakthrough will be AIs learning on the job

Explore episode Jun 26, 2026

Where this was said

Dreaming

At 15:30 · chapter starts 15:23

Dwarkesh introduces 'dreaming' — AI building its own RL environments to rehearse skills — using EfficientZero as an analogy, and positions it as a potential fourth scaling axis alongside pre-training, RL, and inference compute.

Technology
What 2027 Looks Like: AI Learning on the Job

The next big breakthrough will be AIs learning on the job · Jun 26, 2026 Technology

RLVR produces an agent competent enough to deploy. That agent gets a week-long context window and does real work. At the end of the week, a thumbs up triggers OPSD or dreaming to distill everything it learned back into the base weights. Repeat. The AI's skills start expanding beyond its original training domains — with each cycle.

Similar snapshots