According to a benchmark cited in episode 10, GPT-5.5 had an 86% hallucination rate compared to Grok's 17% at the time.
Snapshot · God Mode Podcast
According to a benchmark cited in episode 10, GPT-5.5 had an 86% hallucination rate compared to Grok's 17% at the time.
Where this was said
At 31:53 · chapter starts 28:03
With ChatGPT prompting the structure, Rik runs through each lesson as a headline and Ben riffs on how it aged. The series opens with the moment Dario Amodei and Sam Altman refused to hold hands at a group photo at a conference — a petty but revealing signal of how far relations between the two labs had deteriorated. Lesson two spotlights OpenAI's decision to focus on coding agents above all else, which Ben calls pivotal: the coding-agent flywheel (build better agents → agents improve themselves → everything else improves) turned out to be the year's biggest strategic insight. Lesson three surfaces the Mythos speculation, with Ben now claiming Anthropic trained Mythos in February and almost certainly has a Mythos 2 scored in the 80s or 90s on intelligence benchmarks. Lesson four addresses the Opus nerf — Ben was among the first to notice labs quietly throttling model effort during supply crunches, a practice he traces to around February. The GPT-5.5 86% hallucination benchmark from episode 10 gets a failing grade: fifteen weeks on, OpenAI still hasn't meaningfully moved the needle. The section closes with the Codex vs Claude Code call — once a clear Claude Code victory, now a dead heat — as OpenAI's aggressive developer migration campaign on Twitter bore fruit.
Rik used Claude to analyse all 25 episodes and distil them into three recurring wars: OpenAI vs Anthropic, open vs closed source, and the Elon Corner. The 25-lesson retrospective grades every major call against what actually happened.
Ben speculated that Anthropic has already trained a Mythos 2 model scoring in the 80s or even 90s on intelligence benchmarks, but has not released it.
When AI labs are supply-constrained, they quietly reduce model effort levels in the backend. Users experience a noticeably worse product with no announcement. Ben says this has been happening since at least Opus 4.6 in early 2026.
Opus 5 benchmarks well but consistently underdelivers; Fable 5 just gets it done. Ben has learned to use Fable when he has credits and accepts a worse product the rest of the time — a revealing admission about the gap between benchmark and lived experience.
Quickly forming opinions on how the Twitter algorithm and platform worked allowed the speaker to grow rapidly on the platform.
There are more than 90,000 Flock surveillance cameras currently in use around the United States.
A 2023 report estimated that 10 million Americans own Ring cameras, roughly 1 in 5 households having a video-enabled doorbell.
The ImageNet dataset collected 15 million images to drive machine learning, becoming a cornerstone of the modern AI revolution.
A Stanford graduate student benchmarked human performance on the ImageNet 1,000-category challenge at roughly 4% error rate, a figure AI surpassed by 2016.
From the 2012 ImageNet breakthrough, it took only about 3–4 more years for AI algorithms to surpass human performance in naming 1,000 object categories.
Flock's surveillance network scans more than 20 billion license plates per month across the United States.
OpenAI's Sora, released in January 2024, demonstrated AI's ability to generate realistic video from text prompts, marking a key milestone in video generation.
AlphaGo's Move 37 against Lee Sedol was a move that human Go masters had never considered, illustrating a unique form of AI creativity within constrained mathematical rules.
We use essential and analytics cookies to run Vuci. To understand how the site is used: Privacy Policy.
Install Vuci on your phone
Add it to your home screen for a faster, app-like experience.
Install Vuci on your phone
Tap the Share button, then “Add to Home Screen”.
A new version is available
Reload to get the latest Vuci.