Back-of-the-envelope math suggests the optimal inference batch size for a sparse model like DeepSeek V3 is more than 2,400 concurrent sequences to avoid underutilizing compute.
Snapshot · Dwarkesh Podcast
Back-of-the-envelope math suggests the optimal inference batch size for a sparse model like DeepSeek V3 is more than 2,400 concurrent sequences to avoid underutilizing compute.
Where this was said
At 7:30 · chapter starts 6:45
Prediction seven explores the incentive dynamics that continual learning creates for how AI labs structure their pricing and access policies. If deployment experience is the primary driver of model improvement, then every session a user runs generates valuable training data. Labs will therefore have a strong financial incentive to subsidize — or even give away — access to users who consent to training on their sessions [1] — Dwarkesh Patel "If real-world usage becomes the primary driver of model improvement, AI labs have every incentive to subsidize users who let them train on …" 06:45 . Patel observes this is already visible in the generous deals offered to new users of coding products. The deeper analogy is Google Search: Google gives it away for free because the real product is the behavioral data generated at scale. Conversely, labs may respond to enterprises that refuse training consent by restricting them to inferior models. Both the carrot and the stick point in the same direction: maximizing the flow of real-world experience back into the model.
If real-world usage becomes the primary driver of model improvement, AI labs have every incentive to subsidize users who let them train on sessions — exactly like Google giving away free search. Enterprises that refuse may find themselves locked out of the best models.
AI labs may subsidize or give preferential access to enterprises that allow their sessions to be used for training, similar to how Google gives away search.
AI lab revenues are increasing far faster than their compute costs, evidence of large economies of scale already present in AI training.
Serving personalized AI weights efficiently requires batching thousands of sequences simultaneously — the optimal batch size for a sparse model like DeepSeek V3 exceeds 2,400. Individual users running batch size 1 face more than 100x worse compute efficiency, meaning the economics of personalized AI strongly favor large organizations.
Ad-based monetization works well for game apps where users spend extended time in-session, as seen with Grid and Wordle.
Tool-focused apps like PuffCount are poor candidates for ad monetization because users don't stay in-session long enough.
A hard paywall is a screen that blocks all app features unless the user pays or starts a free trial — it cannot be dismissed.
Mobile apps are primarily monetized through either ads (best for games) or in-app purchases/subscriptions (best for tools).
According to the episode, YouTube outperforms every other social platform for building trust and driving SaaS conversions.
Vasco stated that the majority of his app's user base came directly from his YouTube channel.
SEO Bot features a 'Boost My Domain Rating' button that routes users directly to Listing Bot, an example of in-product cross-selling.
The founder's entire product portfolio is AI-related, making it easier to package products attractively for directories.
The founder attached their SaaS demo to the trending debate about whether AI coding is actually good enough to build a full SaaS product.
We use essential and analytics cookies to run Vuci. To understand how the site is used: Privacy Policy.
Install Vuci on your phone
Add it to your home screen for a faster, app-like experience.
Install Vuci on your phone
Tap the Share button, then “Add to Home Screen”.
A new version is available
Reload to get the latest Vuci.