On Artificial Analysis benchmarks, Grok jumped from approximately 1,000 points (Grok 4.3) to 1,500 points (Grok 4.5), landing just behind Claude Opus and Sonnet.
Snapshot · God Mode Podcast
On Artificial Analysis benchmarks, Grok jumped from approximately 1,000 points (Grok 4.3) to 1,500 points (Grok 4.5), landing just behind Claude Opus and Sonnet.
Where this was said
At 12:58 · chapter starts 10:51
As benchmarks converge, Luca makes a case that resonates across the table: we've entered a qualitative era. He uses Claude for complex legal research — specifically, finding cross-jurisdictional loopholes he likens to hunting cybersecurity bugs — because Fable's depth is unmatched for that kind of pioneering work. But for quick website builds and SEO ranking tools, he reaches for Codex because Claude is simply too slow. Ben adds that Grok was missing from agentic coding conversations until this week, and that the Grok Build subscribers who bought at 60% off a few weeks ago made a surprisingly good deal. Rik notes that Grok 4.5's jump on Artificial Analysis benchmarks — from around 1,000 to 1,500 points, and 39% to 81% on TerminalBench [1] — Rik "Grok benchmark: 1,000 → 1,500 points: On Artificial Analysis benchmarks, Grok jumped from approximately 1,000 points (Grok 4.3) to 1,500 po…" 12:58 — is a serious leap, even if it's still nitpicked.
The quantitative gap between frontier models is closing fast. Luca argues we've entered a qualitative era where model selection is personal — you use Fable for legal research, Codex for quick builds, and Grok when cost matters. There's no universally right answer.
Grok 4.5 jumped from 39% to 81% on TerminalBench compared to Grok 4.3, representing a dramatic improvement in agentic coding capability.
Grok 4.5 more than doubled its banking benchmark score from 12% to 32% compared to Grok 4.3, indicating significant gains on finance-domain tasks.
Grok 4.5's benchmark jump didn't come from architectural breakthroughs — it came from data. Once SpaceX AI plugged into Cursor's coding dataset, the model leapfrogged Composer 2.5. Ben's point lands hard: training data, not model design, is the real competitive moat.
Bhanu and his team built approximately 50 free tools to attract search traffic, each linked back to SiteGPT.
With AI coding tools like Cursor, Bhanu can now create a new free marketing tool in less than 5 minutes by referencing existing tools.
Bhanu filters Ahrefs keyword results to show only those with a keyword difficulty below 10, making them realistic ranking targets for any decent website.
Bhanu sets a minimum search volume of 1,000 monthly searches when selecting keywords to target with free tools.
PropGPT averaged 20 downloads per day right after launching on the App Store through influencer marketing.
Eyal and Yali shut down all marketing and spent 4 months completely rebuilding PropGPT from scratch.
PropGPT has accumulated over 40,000 total downloads since launch.
PropGPT's large language model (AI) operating costs are just $20 per month, and the cost is continually falling.
Ad-based monetization works well for game apps where users spend extended time in-session, as seen with Grid and Wordle.
We use essential and analytics cookies to run Vuci. To understand how the site is used: Privacy Policy.
Install Vuci on your phone
Add it to your home screen for a faster, app-like experience.
Install Vuci on your phone
Tap the Share button, then “Add to Home Screen”.
A new version is available
Reload to get the latest Vuci.