Speaker
Lin Qiao
Appearances over time
1 episodes
Episodes
1Podcasts
Quotes & moments
Fireworks processes more than 40 trillion tokens per day, the majority from customised rather than off-the-shelf models.
Lin Qiao projects Fireworks' token throughput could grow 20x to 100x by end of next year as AI adoption accelerates.
Lin Qiao predicts token costs will fall 10x over the next three years due to supply chain improvements and infrastructure efficiency.
Lin Qiao argued that a 10x cost reduction in AI tokens will trigger a 100x explosion in usage as AI becomes a utility.
Fireworks AI expects to at least double from $800M–$1B ARR by the end of 2026, targeting roughly $2B ARR.
Cursor, an early Fireworks customer when it was a single-digit-million company, grew approximately 1,000x over two years.
Lin Qiao believes the future will feature millions of specialised AI models, one per application or use case, rather than a single dominant AGI.
Meta has been building its own custom AI chips (MTIA) since at least 2018, well before the generative AI era.
Traditional hardware depreciation cycles of 6 years are being upended as multiple GPU SKUs launch within a single year, making older hardware obsolete faster.
Fireworks AI achieves zero KL-divergence between training and inference, meaning model weights transfer with full numerical equivalence and no quality loss.
After Jensen Huang told Lin Qiao that every company must be special to justify its existence, Lin realised the implication was profound: all of a company's product design, data, and user relationships encode irreplaceable private intelligence that no external model can learn.
The top six open-weight models globally are currently Chinese-built. Lin Qiao argues that once a model is open, enterprises can wrap their own guardrails around it, but the deeper point is that every model — Chinese or American — encodes its creator's judgment and taste, which always needs tuning.
Fireworks doesn't hire for competence first. They hire for extreme ownership — people who claim end-to-end problems without being asked, see them through to delivery, and treat every outcome as their own responsibility. That mindset, Lin Qiao says, compounds faster than any skill.
Everyone talks about HBM or energy as the AI bottleneck. Lin Qiao's answer is different: the industry still has no great system designed for 10-trillion-parameter models. Closing that gap requires co-design from model to serving platform to chip system — a level of integration that doesn't yet exist.
The AGI believers assume one model will solve everything. Lin Qiao thinks that's both technically wrong and philosophically depressing. The future is millions of specialised models — one per application, per use case, per company.
Most of the world's valuable data sits locked inside enterprise applications, never touching a general model's training set. Fireworks AI was built on the conviction that activating this private data through specialised models is the real frontier of AI.
When Fireworks was founded, open models were in their infancy. Betting on them was a huge gamble. The PyTorch roots gave the team conviction in open ecosystems, and the payoff came as open models crossed quality thresholds that now rival closed models for the vast majority of enterprise use cases.
Fireworks processes more than 40 trillion tokens a day today. Lin Qiao projects that number could be 20x to 100x higher by end of next year. At those volumes, worries about a CapEx bubble look completely backwards.
Fireworks CTO Dima embedded at Cursor for months to build a distributed reinforcement learning infrastructure that decouples the trainer from RL rollout across six global data centre regions. This let a capital-constrained startup run training jobs that previously required 100,000 interconnected chips at a hyperscaler.
Lin Qiao suggests that as open models handle 90% of enterprise use cases at a fraction of the cost, the market is beginning to re-examine whether frontier model companies are priced for a world that will actually materialise. The power-line metaphor applies: important, yes — irreplaceable, no.
Watching TikTok get briefly banned in the US crystallised a broader truth: if your critical intelligence infrastructure sits on another country's AI, it can be cut off overnight. Lin Qiao argues every nation — and every company — needs sovereign AI the same way they need sovereign electricity grids.
Fireworks AI hit $1 billion in ARR with just 200 employees by betting on specialised inference when everyone else was chasing training. The company processes over 40 trillion tokens a day, mostly from customised rather than off-the-shelf models.
In the SaaS era, finding product-market fit was the hard part — scaling was cheap. In the AI era, companies with genuine demand can still destroy themselves by growing, because AI infrastructure costs don't scale gracefully. Lin Qiao calls it 'scaling to bankruptcy.'
Traditional GPU hardware depreciated over six years. Now multiple new SKUs launch annually, and the newest models always prefer the newest chips. This collapse in depreciation cycles fundamentally changes whether companies should own or rent data centre capacity.
Jensen Huang replies to emails within one minute. Lin Qiao spent years marvelling at the sheer volume before understanding the wisdom: in a fast-moving company, information loss between layers is guaranteed. The only antidote is a leader who stays close to the ground and makes precise, fast judgements.
Analysis
What they talk about
- Business 67%
- Technology 25%
- Society & Culture 8%
Connections
Shows they appear on and people they share episodes with. Drag to explore.