Quote · Dwarkesh Podcast
The next big breakthrough will be AIs learning on the job
Where this was said
What 2027 looks like
At 17:48 · chapter starts 17:23
Dwarkesh sketches a concrete 2027–28 scenario where RLVR-primed agents get week-long context windows, do real work, and distill session learnings back into weights — enabling AI to improve primarily from deployment experience [1] — Dwarkesh Patel "RLVR produces an agent competent enough to deploy. That agent gets a week-long context window and does real work. At the end of the week, a…" 16:30 .
Once continual learning arrives, AI improvement will stop being driven by pre-training runs. Instead, every user interaction globally feeds back into the model. It's a network effect applied to intelligence — and it's fundamentally different from anything that exists today.