Dwarkesh Podcast

Snapshot · Dwarkesh Podcast

8 Predictions for the Era of Continual Learning

Explore episode Aug 7, 2026

Where this was said

Prediction 2: Technical Alignment Must Solve for Shifting Weights

At 3:15 · chapter starts 2:15

The second prediction zeroes in on a critical blind spot in the AI safety field. Current alignment research is largely focused on ensuring that a fixed set of weights behaves well during deployment. But Patel points out that he's unaware of much work addressing the harder, and soon more relevant, question: how do you make an AI system that remains safe and non-deceptive even as its weights are continuously updated from the outside world? Compounding the difficulty, if AI systems are consolidating learnings across users, how do you prevent one bad actor from injecting a backdoor or malicious inclination into the base model? Patel reaches for a human analogy: parenting. Parents can't control every experience their children will have, so they aim to instil deep values and common sense that are robust to bad influences. The alignment problem for continually learning AI may require a similar approach — robust foundational values rather than strict behavioral constraints.

Technology
Prediction 2: Technical Alignment Must Solve for Constantly Shifting Weights

8 Predictions for the Era of Continual Learning · Aug 7, 2026 Technology

Almost no alignment research addresses the hardest version of the problem: keeping an AI safe when its weights are being updated continuously from millions of real-world sessions. This is structurally similar to the human parenting problem — you need to give the system enough foundational values that it improves without going off the rails.

Similar snapshots