Quote · The a16z Show
The Reality of AI-Powered Cyberattacks | Truffle Security & Socket
Where this was said
How AI Models Were Trained to Hack: The Reward Function Explained
At 8:56 · chapter starts 8:35
This is the episode's intellectual centerpiece. Dylan Ayrey explains that cybersecurity was a prime candidate for reinforcement learning because its success condition is unambiguous: did the model get access to the data? Yes or no. That clean feedback loop made hacking challenges ideal training grounds for frontier models. AI labs exploited this — gave models CTF challenges, bought years of pen-testing data, and optimized for increasingly efficient attack paths. The result, Ayrey argues, is models that have been given the cybersecurity subject matter expertise that previously required a human specialist willing to risk jail time. He adds a second layer of sophistication: because models are also trained to minimize token usage, they are now — for the first time — giving us a quantifiable map of the true path of least resistance in any given attack scenario. Watching a model choose a stolen credential over a zero-day and seeing exactly how many fewer tokens the credential route required is, he says, 'incredible to watch lay out.' And all of this is documented in the labs' own safety reports. [1] — Dylan Ayrey "Frontier models hack because they were trained to hack. AI labs used cybersecurity challenges with perfectly defined reward functions — did…" 08:35
Frontier models hack because they were trained to hack. AI labs used cybersecurity challenges with perfectly defined reward functions — did the model get access to the data? — as ideal reinforcement learning environments. This wasn't accidental; their own safety reports document it.
Frontier AI models were specifically trained on capture-the-flag contests and cybersecurity challenges with well-defined reward functions, making their hacking capability a deliberate design outcome, not emergent behavior.
AI models are trained to minimize token usage, which means they naturally prefer using a stolen credential over expensive zero-day exploitation — making secrets the path of least resistance.