Quote · Dwarkesh Podcast
8 Predictions for the Era of Continual Learning
Where this was said
Prediction 2: Technical Alignment Must Solve for Shifting Weights
At 2:35 · chapter starts 2:15
The second prediction zeroes in on a critical blind spot in the AI safety field. Current alignment research is largely focused on ensuring that a fixed set of weights behaves well during deployment. But Patel points out that he's unaware of much work addressing the harder, and soon more relevant, question: how do you make an AI system that remains safe and non-deceptive even as its weights are continuously updated from the outside world [1] — Dwarkesh Patel "Almost no alignment research addresses the hardest version of the problem: keeping an AI safe when its weights are being updated continuous…" 02:15 ? Compounding the difficulty, if AI systems are consolidating learnings across users, how do you prevent one bad actor from injecting a backdoor or malicious inclination into the base model? Patel reaches for a human analogy: parenting. Parents can't control every experience their children will have, so they aim to instil deep values and common sense that are robust to bad influences. The alignment problem for continually learning AI may require a similar approach — robust foundational values rather than strict behavioral constraints.
Almost no alignment research addresses the hardest version of the problem: keeping an AI safe when its weights are being updated continuously from millions of real-world sessions. This is structurally similar to the human parenting problem — you need to give the system enough foundational values that it improves without going off the rails.
Fewer than 5 prominent AI base models currently serve millions to billions of users, and they're all trained on roughly the same data, producing homogeneous outputs.
Continual learning from different real-world deployments will cause AI instances to diverge, creating a healthy diversity of AI minds rather than a monolithic singleton.
Today there are fewer than 5 prominent AI base models, all trained on similar data, producing eerily homogeneous outputs — classic mode collapse. Continual learning from diverse real-world deployments could finally break this, producing a genuinely diverse ecosystem of AI minds.