God Mode Podcast

Podbit · God Mode Podcast

EP21: The new AI stack — Claude Fable plans, GPT5.6 builds, Grok 4.5 grinds

Explore episode Jul 9, 2026

Where this was said

OpenAI retracts SWE-bench — are benchmarks broken?

At 37:36 · chapter starts 34:56

The timing is suspicious and the hosts say so: OpenAI retracted SWE-Bench Pro right before shipping GPT-5.6 Soul, the model that would otherwise have been evaluated on it. Their audit found 30% of benchmark tasks are simply broken. Rik notes they've been skeptical of benchmarks for a while — that's partly why they use BridgeMind, which maintains its own independent evaluation. Ben raises a deeper problem: if the AI models doing the auditing are the same models being benchmarked, the process is fundamentally compromised. Who watches the watchmen? The answer, increasingly, is other AI — and that's a problem nobody has cleanly solved yet.

Similar podbits