Anthropic's classifiers reduced jailbreak success from 95% to 4%, and then their latest two-stage system brought it down to under 0.5%.
Anthropic's classifiers reduced jailbreak success from 95% to 4%, and then their latest two-stage system brought it down to under 0.5%.
Where this was said
At 1:06:57 · chapter starts 1:02:50
The FUD around Fable being 'nerfed' after its return gets a systematic takedown. Ben notes that public benchmarks comparing Fable before and after the export ban are mostly hitting prompts that now trigger rerouting to Opus — the model itself hasn't changed. Theo then digs into the actual architecture behind rerouting, drawing on an Anthropic paper he'd reviewed that day. [1] — Theo "Anthropic's rerouting system has two stages: a cheap first-stage classifier (~0.05% compute) that watches which expert pathways activate in…" 1:02:50 The system is two-stage: a first-stage classifier monitors which expert pathways activate during inference (essentially watching the direction of the weights) and costs roughly 0.05% of full inference compute. When it flags a request, a far more expensive secondary classifier fires — one that adds 27–80% to the compute cost per request. By building the cheap first-stage gate, Anthropic can catch the vast majority of concerning requests without paying the full secondary cost on every token. The combined system has driven jailbreak pass rates from 95% down to under 0.5%, though Theo notes that since all of this runs on every token of output, there's an interesting lower bound related to how many tokens it takes to encode any given piece of dangerous information. Both hosts agree rerouting is annoying but rare in practical coding use — and recommend just sending a normal follow-up message to exit Opus and return to Fable.
Anthropic's rerouting system has two stages: a cheap first-stage classifier (~0.05% compute) that watches which expert pathways activate in the model, and an expensive second-stage classifier (27–80% extra compute) that only fires when the first flags a request. The result: jailbreak pass rates dropped from 95% to under 0.5%.
Anthropic's secondary safety classifier increases compute per request by 27% to 80%, prompting them to build a cheaper first-stage check costing only ~0.05% of compute.
GLM-5.2 is fast, cheap, and produces clean code easily controlled by a smarter orchestrator. Sonnet 5 edges it out only in one scenario: very long-running jobs requiring recursive self-calling and review loops. Outside that, GLM-5.2 wins on price-performance — and even Theo would pick it over Sonnet if he still had Fable.
Bhanu and his team built approximately 50 free tools to attract search traffic, each linked back to SiteGPT.
With AI coding tools like Cursor, Bhanu can now create a new free marketing tool in less than 5 minutes by referencing existing tools.
Bhanu filters Ahrefs keyword results to show only those with a keyword difficulty below 10, making them realistic ranking targets for any decent website.
Bhanu sets a minimum search volume of 1,000 monthly searches when selecting keywords to target with free tools.
PropGPT averaged 20 downloads per day right after launching on the App Store through influencer marketing.
Eyal and Yali shut down all marketing and spent 4 months completely rebuilding PropGPT from scratch.
PropGPT has accumulated over 40,000 total downloads since launch.
PropGPT's large language model (AI) operating costs are just $20 per month, and the cost is continually falling.
Ad-based monetization works well for game apps where users spend extended time in-session, as seen with Grid and Wordle.
We use essential and analytics cookies to run Vuci. To understand how the site is used: Privacy Policy.
Install Vuci on your phone
Add it to your home screen for a faster, app-like experience.
Install Vuci on your phone
Tap the Share button, then “Add to Home Screen”.
A new version is available
Reload to get the latest Vuci.