Quote · Dwarkesh Podcast
The next big breakthrough will be AIs learning on the job
Where this was said
Will RLVR alone generalize?
At 7:07 · chapter starts 6:10
Dwarkesh questions whether RLVR generalizes from containerized tasks to complex real-world domains like politics or business, citing Dario Amodei's context-length degradation comment as a key data point [1] — Dwarkesh Patel "Dario Amodei noted that training at short context lengths and serving at long ones can cause performance degradation. Dwarkesh reads this a…" 06:55 .
Dario Amodei noted that training at short context lengths and serving at long ones can cause performance degradation. Dwarkesh reads this as evidence that RLVR generalization is not infinite — a crack in the foundation of the labs' AGI bet.
Labs burn 30–50% of compute on inference — and currently gain zero improvement to the model from it. Meanwhile, deployment is where the most valuable learning signals actually exist: real organizational context, real mistakes, real feedback.
Labs spend 30 to 50% of their total compute on inference, yet this compute currently plays no productive role in improving the model.