Nerd Snipe with Theo and Ben

Snapshot · Nerd Snipe with Theo and Ben

We Tested GPT 5.6 Sol Early

Explore episode Jul 9, 2026

Where this was said

Long-Running Tasks & What Actually Changed

At 36:05 · chapter starts 35:00

This chapter traces the evolution of how both hosts actually use these models. Fable was the catalyst: it was the first model that made genuinely long, complex runs feel worth attempting. When Fable was taken away, they found 5.6 could carry those same expanded workflows further than any previous OpenAI model. Theo explains the economics: a $200 Codex subscription, per Semianalysis measurements, delivers up to $14,000 worth of compute per period — and with two usage resets that restore a full week, this can reach $20,000 per month. He's careful to note that the outrageous dollar figures on their dashboards ($131,700 and $93,000 respectively) represent deliberate stress-testing experiments, not practical work. His real workflow tasks — auditing open PRs, rebasing, coordinating merges — ran on a single 5.6 thread spawning targeted subagents, and were both cheaper and more valuable.

Technology
How to Burn $65k on a Single Loop

We Tested GPT 5.6 Sol Early · Jul 9, 2026 Technology

Ben's single run to port the Executor project to Rust and Svelte ran to 100 billion tokens and cost $65,000. The real driver wasn't the orchestrating reasoning model — it was the dozens of subagents it spun up underneath, which is the only way to blow through API usage this fast.

Technology
$65k on a single loop run

We Tested GPT 5.6 Sol Early · Jul 9, 2026

Ben's single longest run — porting the Executor project to Rust and Svelte — cost $65,000 worth of API tokens via xHI reasoning model orchestrating massive subagent swarms.

Similar snapshots