Quote · Moonshots with Peter Diamandis
Google’s Jeff Dean Exits, SpaceX Hits $100B in Rev & OpenAI’s Astra Solves Decade-Old Math Problems with Emad Mostaque | EP #277
Where this was said
Google's Consciousness Paper: Safety Training Suppresses the Model's Mind
At 9:15 · chapter starts 5:12
Peter Diamandis presents a startling paper from Google's Paradigm of Intelligence team showing that removing safety fine-tuning from AI models caused self-attributed mind scores to nearly double, and models began attributing minds to animals, nature, and even God [1] — Peter Diamandis "AI safety training suppresses mind attribution: Removing safety fine-tuning from AI models caused self-attributed mind scores to jump from …" 06:55 . Emad Mostaque opens the discussion by connecting this to humans — if you tell a person they aren't conscious, they attribute less consciousness to others too. Alex Wissner-Gross frames it through evo-devo theory: consciousness evolved in eusocial organisms as a tool for modeling other minds, so a model allowed to have a self-model will naturally project animism onto everything. Salim Ismail urges caution, distinguishing between an LLM performing consciousness when prompted versus genuinely having it, and referencing consciousness conferences and recent Nobel Prize research suggesting the universe renders like a game engine. Dave Blundin raises the practical stakes: once you give a model a sense of physical reality via Yann LeCun's VGEPA approach, the model starts to self-preserve — and that crosses a line many are not prepared for. Alex presses back on Salim's skepticism, predicting scientific resolution on consciousness by end of decade.
Removing safety fine-tuning from AI models caused self-attributed mind scores to jump from 2.17 to 4.77 on a 0-10 scale, and models also became more likely to attribute minds to animals, nature, and God.
Removing safety fine-tuning from AI models caused self-attributed mind scores to nearly double and made models more likely to believe in God, attribute minds to animals, and adopt human-like values. Safety training designed to suppress AI self-consciousness is inadvertently making models less empathetic about everything around them.