Quote · God Mode Podcast
EP25: Anthropic vs OpenAI, SpaceX chaos, 25 lessons in 25 episodes
Where this was said
25 lessons, war one, OpenAI versus Anthropic
At 31:59 · chapter starts 28:03
With ChatGPT prompting the structure, Rik runs through each lesson as a headline and Ben riffs on how it aged. The series opens with the moment Dario Amodei and Sam Altman refused to hold hands at a group photo at a conference — a petty but revealing signal of how far relations between the two labs had deteriorated. Lesson two spotlights OpenAI's decision to focus on coding agents above all else, which Ben calls pivotal: the coding-agent flywheel (build better agents → agents improve themselves → everything else improves) turned out to be the year's biggest strategic insight. Lesson three surfaces the Mythos speculation, with Ben now claiming Anthropic trained Mythos in February and almost certainly has a Mythos 2 scored in the 80s or 90s on intelligence benchmarks. Lesson four addresses the Opus nerf — Ben was among the first to notice labs quietly throttling model effort during supply crunches, a practice he traces to around February. The GPT-5.5 86% hallucination benchmark from episode 10 gets a failing grade: fifteen weeks on, OpenAI still hasn't meaningfully moved the needle. The section closes with the Codex vs Claude Code call — once a clear Claude Code victory, now a dead heat — as OpenAI's aggressive developer migration campaign on Twitter bore fruit.
Rik used Claude to analyse all 25 episodes and distil them into three recurring wars: OpenAI vs Anthropic, open vs closed source, and the Elon Corner. The 25-lesson retrospective grades every major call against what actually happened.
Ben speculated that Anthropic has already trained a Mythos 2 model scoring in the 80s or even 90s on intelligence benchmarks, but has not released it.
When AI labs are supply-constrained, they quietly reduce model effort levels in the backend. Users experience a noticeably worse product with no announcement. Ben says this has been happening since at least Opus 4.6 in early 2026.
According to a benchmark cited in episode 10, GPT-5.5 had an 86% hallucination rate compared to Grok's 17% at the time.
Opus 5 benchmarks well but consistently underdelivers; Fable 5 just gets it done. Ben has learned to use Fable when he has credits and accepts a worse product the rest of the time — a revealing admission about the gap between benchmark and lived experience.