Quote · God Mode Podcast
EP21: The new AI stack — Claude Fable plans, GPT5.6 builds, Grok 4.5 grinds
Where this was said
The qualitative era: picking a model is taste now
At 10:55 · chapter starts 10:51
As benchmarks converge, Luca makes a case that resonates across the table: we've entered a qualitative era. He uses Claude for complex legal research — specifically, finding cross-jurisdictional loopholes he likens to hunting cybersecurity bugs — because Fable's depth is unmatched for that kind of pioneering work. But for quick website builds and SEO ranking tools, he reaches for Codex because Claude is simply too slow. Ben adds that Grok was missing from agentic coding conversations until this week, and that the Grok Build subscribers who bought at 60% off a few weeks ago made a surprisingly good deal. Rik notes that Grok 4.5's jump on Artificial Analysis benchmarks — from around 1,000 to 1,500 points, and 39% to 81% on TerminalBench [1] — Rik "Grok benchmark: 1,000 → 1,500 points: On Artificial Analysis benchmarks, Grok jumped from approximately 1,000 points (Grok 4.3) to 1,500 po…" 12:58 — is a serious leap, even if it's still nitpicked.
The quantitative gap between frontier models is closing fast. Luca argues we've entered a qualitative era where model selection is personal — you use Fable for legal research, Codex for quick builds, and Grok when cost matters. There's no universally right answer.
On Artificial Analysis benchmarks, Grok jumped from approximately 1,000 points (Grok 4.3) to 1,500 points (Grok 4.5), landing just behind Claude Opus and Sonnet.
Grok 4.5 jumped from 39% to 81% on TerminalBench compared to Grok 4.3, representing a dramatic improvement in agentic coding capability.
Grok 4.5 more than doubled its banking benchmark score from 12% to 32% compared to Grok 4.3, indicating significant gains on finance-domain tasks.
Grok 4.5's benchmark jump didn't come from architectural breakthroughs — it came from data. Once SpaceX AI plugged into Cursor's coding dataset, the model leapfrogged Composer 2.5. Ben's point lands hard: training data, not model design, is the real competitive moat.