Speaker
Alex Atallah
Appearances over time
1 episodes
Episodes
1Podcasts
Quotes & moments
OpenRouter launched 70 new models in July 2025, roughly one model every 10 hours, reflecting the explosive pace of AI model development.
After OpenAI cut GPT-5.6 Luna's price 10x on OpenRouter, usage grew 13x — a near-perfect real-world Jevons paradox demonstration.
US enterprises are more worried about frontier model data policies from US labs than about Chinese models, because they can't run frontier models on their own infrastructure.
Alex Atallah believes AI will be the biggest market not just in tech history but in all of human history, and no single model will capture all of it.
The overall AI inference market has been growing 10 to 15x per year, and OpenRouter's revenue is expected to continue being dominated by unplanned inference capacity needs.
Claude Sonnet is a partially distilled version of Opus, illustrating that even closed-weight frontier labs routinely use distillation to create smaller, cheaper models.
OpenRouter's central routing tech detects quality improvements, speedups, or price reductions every 5 minutes and immediately shifts traffic to better providers.
GPT-5.6 Luna is now in the top 3–5 models by token volume on OpenRouter, the first time an OpenAI model has reached that ranking in a very long time.
In the AI era, employee costs are becoming dynamic rather than static salaries, dependent on which AI models and tools they use and how efficiently they use them.
Distillation is vilified in public discourse, but Claude Sonnet is a distilled version of Opus. All major labs do it. The real question isn't whether to distill but whether you're using it to inspect and align the outputs — which is actually easier with open-weight models.
During the NFT boom, OpenSea's servers were melting and the site faced catastrophic outages. The fear of becoming 'the Twitter fail whale for crypto' forced Atallah to master infrastructure scaling at scale — a discipline he brought directly to OpenRouter.
America is very, very behind China on open-weight models. GLM 5.2 was a massive leap. KIMI K3 is catching up. Meanwhile, US open-source labs struggle to raise funding while competing against OpenAI and Anthropic on one side and state-backed Chinese labs on the other.
Reports of a $10B Stripe acquisition are swirling, and Alex Atallah won't deny them. His only comment: 'Whatever happens, we're going to execute on the vision.' Make of that what you will.
OpenAI cut GPT-5.6 Luna's price 10x on OpenRouter. Usage exploded 13x. This is the Jevons paradox playing out in live data, and it's the most bullish possible signal for the entire inference economy.
US enterprises are more scared of OpenAI and Anthropic than Chinese models. The reason: they can't see where their prompts go, and they can't run frontier models on their own infra. Chinese open-weight models, paradoxically, give them more control.
In July alone, OpenRouter added 70 models — one every 10 hours. And that's before agent labs like Cognition, Cursor, and Jeff Dean's new venture have even started releasing their own models in earnest.
Everyone is building a router because it's fashionable. But a router built as a side quest is months behind one built with 100% focus. And worse, partial routers reduce user leverage by limiting model access and flexibility.
Consolidation on a single AI model doesn't make sense. Creativity isn't verifiable, two models trained on different data will always produce ideas the other can't, and AI will be the biggest market in human history — too big for any one winner.
The hyperscalers haven't monopolized AI inference because NVIDIA actively wants market heterogeneity. Preventing customer concentration among cloud providers is a top NVIDIA priority — and it's created space for inference startups to consistently outperform Google, Amazon, and Azure.
OpenRouter treats AI safety like internet safety: you don't ban the internet, you build guardrails. It offers prompt injection protection, PII redaction, and works with model labs on safety practices — because it's the ideal choke point for deploying safety across an entire enterprise.
The fear that agent frameworks will absorb the routing layer misses something key: as models get smarter, bloated system prompts become a handicap, not a feature. Anthropic's own research showed removing prompt clutter improved model performance. Harnesses will evolve, not disappear.
In the AI age, employee cost is no longer a static salary — it's a dynamic variable driven by which models workers use and how efficiently. Companies should map employees on a quadrant: high productivity vs. cost effectiveness, and address the 'AI psychosis' in the danger zone.
Memory will be a key AI retention mechanism, but no single layer can own all of it. The model has the best intelligence context, the app has the richest behavioral data, and the router sits in between. The labs will need to incentivize app developers to share context they don't currently have.
To compete with Chinese open models, the US needs to make compute accessible to the right talent, leverage distillation of Chinese models to bootstrap American training, and build a NeoLab ecosystem. NVIDIA is already doing some of this — but TPUs, Trainium, and NeoChips need to be part of the answer.
Analysis
What they talk about
- Technology 85%
- Business 15%
Connections
Shows they appear on and people they share episodes with. Drag to explore.