Quote · Dwarkesh Podcast
8 Predictions for the Era of Continual Learning
Where this was said
Prediction 8: Inference Economics Will Strongly Favor Large Enterprises
At 8:25 · chapter starts 7:32
The final prediction dives into the technical economics of inference at scale. Patel explains that serving a set of weights efficiently requires many concurrent sequences to be decoded against it simultaneously — a concept he explored in depth with Ryder Pope in a prior episode [1] — Dwarkesh Patel "Serving personalized AI weights efficiently requires batching thousands of sequences simultaneously — the optimal batch size for a sparse m…" 07:32 . Back-of-the-envelope estimates suggest the optimal batch size for a sparse model like DeepSeek V3 exceeds 2,400 concurrent sequences; falling short of this means leaving compute on the table. A large enterprise with thousands of employees and agents running diverse tasks can hit that batch size, efficiently utilizing its personalized weight fork. An individual user, by contrast, runs at batch size 1 — potentially suffering more than two orders of magnitude worse efficiency. The economic conclusion is sharp: personalized AI weights are a corporate-scale technology. The costs and efficiencies of the continual learning era will flow disproportionately to large organizations, not individuals. Patel closes by acknowledging that the most important consequences of continual learning are probably the ones hardest to anticipate — but the eight above seem clear enough from here.
Serving personalized AI weights efficiently requires batching thousands of sequences simultaneously — the optimal batch size for a sparse model like DeepSeek V3 exceeds 2,400. Individual users running batch size 1 face more than 100x worse compute efficiency, meaning the economics of personalized AI strongly favor large organizations.
An individual user running batch size 1 may suffer more than 2 orders of magnitude worse compute efficiency compared to a large organization efficiently serving personalized weights.