Where this was said
AlphaGo and the Danger of Training AI to Be Human
At 56:39 · chapter starts 54:10
Rubin circles back to the question of AI's true potential with a concrete case study. AlphaGo beat the world's best Go player not by playing like a human expert but by making a move that looked, to human commentators, like a mistake — something completely outside the culture of the game, even if within the rules. [1] — Rick Rubin "AlphaGo beat the world Go champion with a move no human would make — a move commentators initially called a mistake. When we train AI to on…" 51:00 When the grandmaster saw the move, he got up and left the room. The commentators called it an error. It was, in fact, the winning move. Rubin's point: if the AI had been trained to only do what humans do, it never would have found that move. The lesson he draws for the broader AI industry is provocative — by training AI through RLHF to mirror human values and human behavior, we may be systematically preventing it from finding the equivalent of AlphaGo's unthinkable move in every other domain. Capping the machine at human performance is not a safety feature; it's a ceiling.
Rick Rubin and Johnny Cash did three sessions with the world's best musicians. None of them were as compelling as John sitting in Rubin's living room with an acoustic guitar. The mistake was thinking they were making demos — they were making the record.