OpenAI audited SWE-Bench Pro and found 30% of tasks are broken, then retracted it as a leading coding eval — right before shipping GPT-5.6 Soul. Ben raised the obvious conflict: was it GPT-5.6 that found the flaws in the benchmark it was being tested on?