Quote · All-In with Chamath, Jason, Sacks & Friedberg
Open Source Wins, AGI Is Here, and Scorsese's AI Toolkit with CEOs of Cerebras & Black Forest Labs
Where this was said
Reasoning, Inference, and Breaking Moore's Law
At 7:33 · chapter starts 1:50
Feldman describes a fundamental shift in human-AI interaction: early models did exactly what you told them — the classic 'computers are dumb, they do exactly what you say' formulation from a colleague two decades ago. But modern reasoning models like those in OpenAI's latest releases are beginning to infer intent: a user asks for a chart, and the model proactively offers both a line and a bar version. Calacanis illustrates this vividly with a live experiment — using an unrestricted model to build a real-time trend-scouting agent that debates with itself about where to look for emerging signals (Hacker News, Reddit, Instagram). He watches the model reason in real time, debating its own approach before collapsing into an answer. This is not ChatGPT summarizing a PDF — this is an agent doing strategic research, checking its own work, and asking the human what it missed.
Individual AI data centers under construction will draw more power than mid-sized cities, with aggregate demand set to exceed the previous 50 years of Earth's energy use.
Cerebras is sitting on $25 billion in backlog, and every hyperscaler from OpenAI to AWS faces the same problem: demand is fully booked, and they're racing to keep customers from leaving — not chasing speculative future adoption.
Cerebras has a $25 billion backlog of orders, reflecting demand that far outstrips the industry's ability to build and fill data centers.
Early AI adoption looked like everyone grabbing unlimited tokens with no strategy — exactly like someone wandering every aisle at Costco and walking out with $200 of impulse buys. Enterprises are now learning which AI models to use for which tasks, and routing intelligently between frontier and open source.
Modern reasoning AI consumes enormous numbers of tokens internally — essentially thinking out loud before responding. That internal computation is inference, and speed directly translates to more reasoning cycles per dollar. Run Cerebras for 24 hours and you get the equivalent of weeks of AI thinking.
Cerebras chips can run inference 15 times faster than competing hardware, meaning a 24-hour run on Cerebras could yield weeks or months worth of AI thinking.