Speaker
Lin Qiao
Appearances over time
1 episodes
Episodes
1Podcasts
Quotes & moments
Fireworks processes more than 40 trillion tokens per day, the majority from customised rather than off-the-shelf models.
Lin Qiao projects Fireworks' token throughput could grow 20x to 100x by end of next year as AI adoption accelerates.
Lin Qiao predicts token costs will fall 10x over the next three years due to supply chain improvements and infrastructure efficiency.
Lin Qiao argued that a 10x cost reduction in AI tokens will trigger a 100x explosion in usage as AI becomes a utility.
Fireworks AI expects to at least double from $800M–$1B ARR by the end of 2026, targeting roughly $2B ARR.
Cursor, an early Fireworks customer when it was a single-digit-million company, grew approximately 1,000x over two years.
Lin Qiao believes the future will feature millions of specialised AI models, one per application or use case, rather than a single dominant AGI.
Meta has been building its own custom AI chips (MTIA) since at least 2018, well before the generative AI era.
Traditional hardware depreciation cycles of 6 years are being upended as multiple GPU SKUs launch within a single year, making older hardware obsolete faster.
Fireworks AI achieves zero KL-divergence between training and inference, meaning model weights transfer with full numerical equivalence and no quality loss.
Fireworks AI hit $1 billion in ARR with just 200 employees by betting on specialised inference when everyone else was chasing training. The company processes over 40 trillion tokens a day, mostly from customised rather than off-the-shelf models.
After Jensen Huang told Lin Qiao that every company must be special to justify its existence, Lin realised the implication was profound: all of a company's product design, data, and user relationships encode irreplaceable private intelligence that no external model can learn.
In the SaaS era, finding product-market fit was the hard part — scaling was cheap. In the AI era, companies with genuine demand can still destroy themselves by growing, because AI infrastructure costs don't scale gracefully. Lin Qiao calls it 'scaling to bankruptcy.'
Lin Qiao suggests that as open models handle 90% of enterprise use cases at a fraction of the cost, the market is beginning to re-examine whether frontier model companies are priced for a world that will actually materialise. The power-line metaphor applies: important, yes — irreplaceable, no.
The top six open-weight models globally are currently Chinese-built. Lin Qiao argues that once a model is open, enterprises can wrap their own guardrails around it, but the deeper point is that every model — Chinese or American — encodes its creator's judgment and taste, which always needs tuning.
Token costs are artificially high because of supply chain constraints. Once competition and infrastructure catch up, Lin Qiao expects a 10x cost reduction in three years. Cheaper tokens will unlock use cases that are currently economically impossible, driving a 100x surge in usage.
Fireworks achieves zero KL-divergence between training and inference — meaning model weights transfer with bit-exact numerical equivalence. This matters enormously at scale: even tiny quality drops at the train-inference boundary waste customers' entire training investment.
Traditional GPU hardware depreciated over six years. Now multiple new SKUs launch annually, and the newest models always prefer the newest chips. This collapse in depreciation cycles fundamentally changes whether companies should own or rent data centre capacity.
Watching TikTok get briefly banned in the US crystallised a broader truth: if your critical intelligence infrastructure sits on another country's AI, it can be cut off overnight. Lin Qiao argues every nation — and every company — needs sovereign AI the same way they need sovereign electricity grids.
Fireworks doesn't hire for competence first. They hire for extreme ownership — people who claim end-to-end problems without being asked, see them through to delivery, and treat every outcome as their own responsibility. That mindset, Lin Qiao says, compounds faster than any skill.
Jensen Huang replies to emails within one minute. Lin Qiao spent years marvelling at the sheer volume before understanding the wisdom: in a fast-moving company, information loss between layers is guaranteed. The only antidote is a leader who stays close to the ground and makes precise, fast judgements.
Just as every company builds its own software stack because its problems are unique, every company will eventually own its own AI intelligence stack. Renting general intelligence from a handful of providers will be seen as a competitive liability — not a convenience.
Lin Qiao sees 2024 as the year of coding AI and 2025 as the year of co-work. Co-work is dramatically more diverse than coding — spanning legal, finance, healthcare, customer support, and consumer-facing AI — and Fireworks is already landing customers across all of those verticals.
Lin Qiao met George Hu, the former President of Salesforce, a year before hiring him — and turned him down because Fireworks had only 50 people. The relationship evolved through board-level advice before Hu formally joined as Fireworks crossed into hypergrowth. The lesson: build the relationship early, even if the timing isn't right yet.
Everyone talks about HBM or energy as the AI bottleneck. Lin Qiao's answer is different: the industry still has no great system designed for 10-trillion-parameter models. Closing that gap requires co-design from model to serving platform to chip system — a level of integration that doesn't yet exist.
Analysis
What they talk about
- Business 67%
- Technology 25%
- Society & Culture 8%
Connections
Shows they appear on and people they share episodes with. Drag to explore.