Unlike RL-VR, on-policy self-distillation doesn't require an outer-loop verifiable reward — it only requires a model that can learn the right things within the context window.
Snapshot · Dwarkesh Podcast
Unlike RL-VR, on-policy self-distillation doesn't require an outer-loop verifiable reward — it only requires a model that can learn the right things within the context window.
Where this was said
At 11:40 · chapter starts 8:41
Dwarkesh explores on-policy self-distillation (OPSD) as the leading technique for continual learning, contrasting it with naive SFT and explaining why RL's sparse parameter updates are a feature, not a bug [1] — Dwarkesh Patel "OPSD trains the base model to match what a veteran 'teacher' model — loaded with a full session's context — would predict. It needs no oute…" 11:05 .
The CursorTab model online-learns by predicting which edits users accept, processing over 400 million requests per day.
OPSD trains the base model to match what a veteran 'teacher' model — loaded with a full session's context — would predict. It needs no outer-loop reward, provides denser per-token supervision than RL, and avoids overwriting existing knowledge. It's the most credible path to continual learning today.
On-policy self-distillation provides per-token supervision rather than a single end-of-trajectory reward, making it a much denser training signal than naive RL.
Naively fine-tuning on session transcripts trains models to recall everything with perfect fidelity — but that's not how skill acquisition works. RL concentrates updates only where they matter and changes very few parameters, which is exactly what you want when you don't want to forget what you already know.
RL training changes very few model parameters per step, concentrating updates only on what is relevant to achieving the outcome — a crucial property for continual learning.
Dreaming means the model spends compute generating its own simulated RL environments, then trains against them — rehearsing skills relevant to the actual user's real-world context. Like EfficientZero playing dozens of internal games per real step, this could make AI far more sample-efficient without requiring external data.
Ad-based monetization works well for game apps where users spend extended time in-session, as seen with Grid and Wordle.
Tool-focused apps like PuffCount are poor candidates for ad monetization because users don't stay in-session long enough.
A hard paywall is a screen that blocks all app features unless the user pays or starts a free trial — it cannot be dismissed.
Mobile apps are primarily monetized through either ads (best for games) or in-app purchases/subscriptions (best for tools).
According to the episode, YouTube outperforms every other social platform for building trust and driving SaaS conversions.
Vasco stated that the majority of his app's user base came directly from his YouTube channel.
SEO Bot features a 'Boost My Domain Rating' button that routes users directly to Listing Bot, an example of in-product cross-selling.
The founder's entire product portfolio is AI-related, making it easier to package products attractively for directories.
The founder attached their SaaS demo to the trending debate about whether AI coding is actually good enough to build a full SaaS product.
We use essential and analytics cookies to run Vuci. To understand how the site is used: Privacy Policy.
Install Vuci on your phone
Add it to your home screen for a faster, app-like experience.
Install Vuci on your phone
Tap the Share button, then “Add to Home Screen”.
A new version is available
Reload to get the latest Vuci.