Quote · Dwarkesh Podcast
The next big breakthrough will be AIs learning on the job
Where this was said
Dreaming
At 15:52 · chapter starts 15:23
Dwarkesh introduces 'dreaming' — AI building its own RL environments to rehearse skills — using EfficientZero as an analogy, and positions it as a potential fourth scaling axis alongside pre-training, RL, and inference compute [1] — Dwarkesh Patel "Dreaming means the model spends compute generating its own simulated RL environments, then trains against them — rehearsing skills relevant…" 14:00 .
Dreaming — where a model generates and trains against its own simulated environments — could become a fourth scaling axis alongside pre-training, RL, and inference-time compute.
RLVR produces an agent competent enough to deploy. That agent gets a week-long context window and does real work. At the end of the week, a thumbs up triggers OPSD or dreaming to distill everything it learned back into the base weights. Repeat. The AI's skills start expanding beyond its original training domains — with each cycle.
Dwarkesh envisions that by 2027–28, effective context lengths will expand enough for AIs to co-work with a human for a full week of wall-clock time.