OpenAI audited SWE-Bench Pro and found 30% of its tasks are broken, leading them to retract their recommendation that the research community use it as a leading coding benchmark.
Snapshot · God Mode Podcast
OpenAI audited SWE-Bench Pro and found 30% of its tasks are broken, leading them to retract their recommendation that the research community use it as a leading coding benchmark.
Where this was said
At 35:00 · chapter starts 34:56
The timing is suspicious and the hosts say so: OpenAI retracted SWE-Bench Pro right before shipping GPT-5.6 Soul, the model that would otherwise have been evaluated on it. Their audit found 30% of benchmark tasks are simply broken. [1] — Rik "30% of SWE-Bench Pro tasks broken: OpenAI audited SWE-Bench Pro and found 30% of its tasks are broken, leading them to retract their recomm…" 35:00 Rik notes they've been skeptical of benchmarks for a while — that's partly why they use BridgeMind, which maintains its own independent evaluation. Ben raises a deeper problem: if the AI models doing the auditing are the same models being benchmarked, the process is fundamentally compromised. Who watches the watchmen? The answer, increasingly, is other AI — and that's a problem nobody has cleanly solved yet.
Pieter Levels shipped a native iOS app declaring he's 'no longer iOS retarded.' Luca suspects it's a direct response to Nomad Table, which hit $1M ARR and 70K users with an iOS-only app that actually connects nomads in person — something Nomads.com never managed.
Bhanu and his team built approximately 50 free tools to attract search traffic, each linked back to SiteGPT.
With AI coding tools like Cursor, Bhanu can now create a new free marketing tool in less than 5 minutes by referencing existing tools.
Bhanu filters Ahrefs keyword results to show only those with a keyword difficulty below 10, making them realistic ranking targets for any decent website.
Bhanu sets a minimum search volume of 1,000 monthly searches when selecting keywords to target with free tools.
PropGPT averaged 20 downloads per day right after launching on the App Store through influencer marketing.
Eyal and Yali shut down all marketing and spent 4 months completely rebuilding PropGPT from scratch.
PropGPT has accumulated over 40,000 total downloads since launch.
PropGPT's large language model (AI) operating costs are just $20 per month, and the cost is continually falling.
Ad-based monetization works well for game apps where users spend extended time in-session, as seen with Grid and Wordle.
We use essential and analytics cookies to run Vuci. To understand how the site is used: Privacy Policy.
Install Vuci on your phone
Add it to your home screen for a faster, app-like experience.
Install Vuci on your phone
Tap the Share button, then “Add to Home Screen”.
A new version is available
Reload to get the latest Vuci.