Quote · God Mode Podcast
EP21: The new AI stack — Claude Fable plans, GPT5.6 builds, Grok 4.5 grinds
Where this was said
OpenAI retracts SWE-bench — are benchmarks broken?
At 36:55 · chapter starts 34:56
The timing is suspicious and the hosts say so: OpenAI retracted SWE-Bench Pro right before shipping GPT-5.6 Soul, the model that would otherwise have been evaluated on it. Their audit found 30% of benchmark tasks are simply broken. [1] — Rik "30% of SWE-Bench Pro tasks broken: OpenAI audited SWE-Bench Pro and found 30% of its tasks are broken, leading them to retract their recomm…" 35:00 Rik notes they've been skeptical of benchmarks for a while — that's partly why they use BridgeMind, which maintains its own independent evaluation. Ben raises a deeper problem: if the AI models doing the auditing are the same models being benchmarked, the process is fundamentally compromised. Who watches the watchmen? The answer, increasingly, is other AI — and that's a problem nobody has cleanly solved yet.
OpenAI audited SWE-Bench Pro and found 30% of its tasks are broken, leading them to retract their recommendation that the research community use it as a leading coding benchmark.
Pieter Levels shipped a native iOS app declaring he's 'no longer iOS retarded.' Luca suspects it's a direct response to Nomad Table, which hit $1M ARR and 70K users with an iOS-only app that actually connects nomads in person — something Nomads.com never managed.