Ben speculated that Anthropic has already trained a Mythos 2 model scoring in the 80s or even 90s on intelligence benchmarks, but has not released it.
Snapshot · God Mode Podcast
Ben speculated that Anthropic has already trained a Mythos 2 model scoring in the 80s or even 90s on intelligence benchmarks, but has not released it.
Where this was said
At 31:17 · chapter starts 28:03
With ChatGPT prompting the structure, Rik runs through each lesson as a headline and Ben riffs on how it aged. The series opens with the moment Dario Amodei and Sam Altman refused to hold hands at a group photo at a conference — a petty but revealing signal of how far relations between the two labs had deteriorated. Lesson two spotlights OpenAI's decision to focus on coding agents above all else, which Ben calls pivotal: the coding-agent flywheel (build better agents → agents improve themselves → everything else improves) turned out to be the year's biggest strategic insight. Lesson three surfaces the Mythos speculation, with Ben now claiming Anthropic trained Mythos in February and almost certainly has a Mythos 2 scored in the 80s or 90s on intelligence benchmarks. Lesson four addresses the Opus nerf — Ben was among the first to notice labs quietly throttling model effort during supply crunches, a practice he traces to around February. The GPT-5.5 86% hallucination benchmark from episode 10 gets a failing grade: fifteen weeks on, OpenAI still hasn't meaningfully moved the needle. The section closes with the Codex vs Claude Code call — once a clear Claude Code victory, now a dead heat — as OpenAI's aggressive developer migration campaign on Twitter bore fruit.
Rik used Claude to analyse all 25 episodes and distil them into three recurring wars: OpenAI vs Anthropic, open vs closed source, and the Elon Corner. The 25-lesson retrospective grades every major call against what actually happened.
When AI labs are supply-constrained, they quietly reduce model effort levels in the backend. Users experience a noticeably worse product with no announcement. Ben says this has been happening since at least Opus 4.6 in early 2026.
According to a benchmark cited in episode 10, GPT-5.5 had an 86% hallucination rate compared to Grok's 17% at the time.
Opus 5 benchmarks well but consistently underdelivers; Fable 5 just gets it done. Ben has learned to use Fable when he has credits and accepts a worse product the rest of the time — a revealing admission about the gap between benchmark and lived experience.
Sam's initial MVP was coded in approximately one week using ChatGPT voice mode and copy-pasting code, with no prior technical experience.
Sam argues Discord is 10x better than email for building relationships with younger users who rarely check their inbox.
Sam's monthly operating costs include Cursor ($200), AI image generation ($100), AI video generation ($200), hosting ($100), email marketing ($80), and AI compute ($300–$500).
Sam recommends copying days of Discord chat history into ChatGPT and prompting it to list recurring pain points as a fast, free market research technique.
Bhanu and his team built approximately 50 free tools to attract search traffic, each linked back to SiteGPT.
With AI coding tools like Cursor, Bhanu can now create a new free marketing tool in less than 5 minutes by referencing existing tools.
Bhanu filters Ahrefs keyword results to show only those with a keyword difficulty below 10, making them realistic ranking targets for any decent website.
Bhanu sets a minimum search volume of 1,000 monthly searches when selecting keywords to target with free tools.
PropGPT averaged 20 downloads per day right after launching on the App Store through influencer marketing.
We use essential and analytics cookies to run Vuci. To understand how the site is used: Privacy Policy.
Install Vuci on your phone
Add it to your home screen for a faster, app-like experience.
Install Vuci on your phone
Tap the Share button, then “Add to Home Screen”.
A new version is available
Reload to get the latest Vuci.