Fireworks AI raised $1.5 billion at a $17 billion valuation, a remarkable outcome for a 200-person company.
Fireworks AI hit $1B ARR with 200 people by betting that millions of specialised AI models will beat one AGI — and Lin Qiao says token costs will fall 10x while usage explodes 100x in three years.
The Twenty Minute VC (20VC): Venture Capital | Startup Funding | The Pitch
Fireworks AI hit $1B ARR with 200 people by betting that millions of specialised AI models will beat one AGI — and Lin Qiao says token costs will fall 10x while usage explodes 100x in three years.
TL;DR
Lin Qiao, founder and CEO of Fireworks AI, makes the case that the future of AI isn't one AGI ruling everything — it's millions of specialised models, one per application [1] — Lin Qiao "The AGI believers assume one model will solve everything. Lin Qiao thinks that's both technically wrong and philosophically depressing. The…" 08:35 . Fireworks has hit $1B in ARR with just 200 people and processes over 40 trillion tokens a day, the vast majority from customised rather than off-the-shelf models [2] — Harry Stebbings "$1B ARR, 200 people: Fireworks AI reached $1 billion in annual recurring revenue with only 200 employees, scaling in roughly 4 years." 00:55 . Lin predicts a 10x reduction in token costs over three years will unlock 100x more usage [3] — Lin Qiao "10x cost drop drives 100x usage: Lin Qiao argued that a 10x cost reduction in AI tokens will trigger a 100x explosion in usage as AI become…" 44:50 , and argues every company will eventually own its own intelligence stack the same way every company owns its own software stack. The single most useful takeaway: don't wait for AGI to solve your problem — own your model now or risk losing competitive differentiation.
Lin Qiao, Co-Founder and CEO of Fireworks AI, discusses why the company bet on inference over training, the open-source AI revolution, enterprise trust in Chinese models, the multi-model future, token cost trajectories, and whether Fireworks will eventually need to build its own data centres.
The episode opens with a punchy teaser from Lin Qiao before Harry Stebbings delivers a rare investor endorsement: a $10 million check written after just a 15-minute meeting. Harry breaks down the five reasons — a world-class team, a fast-growing inference market, triple-digit ARR growth to $1 billion in four years, the ability to hire stars like ex-Salesforce President George Hu, and the sheer upside potential of a company he thinks could reach $500 billion. The framing sets the episode's bullish tone before three sponsor integrations (JPMorgan, Navan, Base44) round out the opening block.
Harry asks the valuation question directly: if 90% of enterprise workflows can now be handled by open-weight models at a fraction of the cost, are Anthropic and OpenAI dramatically overvalued? Lin agrees the market is beginning to realise this, then introduces a more uncomfortable concept — 'scaling to bankruptcy.' In the SaaS era, finding product-market fit was the hard problem; scaling was cheap. In the AI era, these are decoupled. Companies with genuine customer demand and willingness to pay can destroy themselves by growing because the AI infrastructure COGS spiral faster than revenue. This creates a structural pull toward open-weight models that companies can control, customise, and optimise.
Harry — declaring a conflict of interest as a Lagoora investor — uses the Harvey-Lagoora dynamic as a test case. A year ago, not building your own model seemed fine because frontier models were improving so fast. Now the picture is reversed. Lin reframes the question: in legal, where error tolerance is near-zero and workflows are deeply specialised, the competitive advantage isn't whether you own a model but whether you've encoded your proprietary knowledge into the orchestration layer and the fine-tuned models that power it. He points to Cursor as evidence — they were the first coding AI company to tune their own models, and almost all coding companies have followed.
Harry asks whether the relentless pace of model launches is sustainable. Lin separates two layers: base general IQ advancement, which comes in occasional large step functions (like the introduction of chain-of-thought reasoning in early 2024), and specialisation, which builds rapidly on top of each step function. As base model quality improves, the number of specialised branches that can be grafted onto the trunk multiplies. If anything, Lin expects the specialisation wave to accelerate faster than general intelligence progress — a view with direct implications for Fireworks' market opportunity.
One of the most technically rich chapters in the episode. Lin explains that early AI tech adopters are hackers who want control, and Cursor — flush with frontier lab researchers — is the prime example. The problem they solved together is profound: hyperscalers run RL training on 100,000 interconnected chips. Cursor and Fireworks had no such cluster. Their solution was to decouple the trainer (which updates model weights) from the RL rollout (which deploys the model into a synthetic environment to collect rewards), distributing the system across five or six global data centre regions and syncing fresh model weights efficiently enough that the reward signal stays numerically sound. This distributed design enabled Cursor's recent model launches without a hyperscaler-level budget.
Harry opens with the customer concentration question — can Fireworks survive a Cursor churn given a potential SpaceX acquisition? Lin pivots to the broader portfolio: 2024 was the year of coding, 2025 is the year of co-work, and co-work customers span legal, finance, healthcare, consumer-facing recommendation systems, and beyond. He then drops a striking operational disclosure: Fireworks processes more than 40 trillion tokens per day, the majority from customised models. When Harry asks where that number goes in a year, Lin offers a range of 20x to 100x — and argues that at those growth rates, the idea that AI CapEx is in a bubble is simply wrong. The world is bottlenecked by physical supply chains, not by lack of demand.
This is the most quotable chapter in the episode. Lin explains that current token prices are artificially elevated by supply chain constraints, not by any fundamental cost floor. As competition increases and infrastructure scales, prices will fall sharply. He quantifies three levers: solving tasks requires fewer tokens as models become more precise; Fireworks' platform optimises inference unit economics for each customised workload; and underlying GPU and memory infrastructure will improve structurally over two to three years. The combined effect: a 10x cost reduction over three years that unlocks a 100x usage explosion. He then dives into zero KLD — Fireworks' commitment to bit-exact numerical equivalence from training to inference — as the quality guarantee that justifies its premium over commoditised inference providers.
Harry observes that China's lack of regulatory friction means data centres go up in weeks versus years in the US, and asks whether that translates into a structural AI advantage. Lin acknowledges the construction velocity differential but argues the US has comparable specialisation — the bottleneck is global supply chain constraints (electricians, transistors, materials) rather than policy alone. The conversation then broadens to sovereign models: the brief TikTok ban illustrated the existential vulnerability of depending on foreign-controlled infrastructure. Lin's analogy is stark — if your country's AI is like electricity, another country should never control the switch. The principle applies equally to individual enterprises: owning your intelligence is not optional.
The quickfire round delivers some of the episode's most memorable moments. Lin admits he was wrong to fear growing too fast — aggressive AI tool adoption and a refined hiring filter for extreme ownership changed his calculus. His Jensen Huang insight crystallises into a philosophy: replying to emails in one minute isn't ego, it's the only way to maintain the information freshness needed for precise, fast leadership judgement. His biggest regret is underinvesting in marketing early, framing it not as fluff but as the discipline of educating customers on the right direction. He closes with his sharpest three-year prediction: owning your own intelligence stack will become mandatory for every company, just as owning your own software stack is today. Sponsor reads for JPMorgan, Navan, and Base44 close the episode.
Chapter 1 · 00:07
The episode opens with a punchy teaser from Lin Qiao before Harry Stebbings delivers a rare investor endorsement: a $10 million check written after just a 15-minute meeting. Harry breaks down the five reasons — a world-class team, a fast-growing inference market, triple-digit ARR growth to $1 billion in four years, the ability to hire stars like ex-Salesforce President George Hu, and the sheer upside potential of a company he thinks could reach $500 billion. The framing sets the episode's bullish tone before three sponsor integrations (JPMorgan, Navan, Base44) round out the opening block.
Fireworks AI raised $1.5 billion at a $17 billion valuation, a remarkable outcome for a 200-person company.
Fireworks AI hit $1 billion in ARR with just 200 employees by betting on specialised inference when everyone else was chasing training. The company processes over 40 trillion tokens a day, mostly from customised rather than off-the-shelf models.
Fireworks AI reached $1 billion in annual recurring revenue with only 200 employees, scaling in roughly 4 years.
Navan claims the industry average for booking a business trip is 45 minutes, versus 7 minutes on their platform.
Lin Qiao founded Fireworks AI at 48 years old, after 7 years at Meta and earlier stints at LinkedIn and in academia.
Most of the world's valuable data sits locked inside enterprise applications, never touching a general model's training set. Fireworks AI was built on the conviction that activating this private data through specialised models is the real frontier of AI.
The AGI believers assume one model will solve everything. Lin Qiao thinks that's both technically wrong and philosophically depressing. The future is millions of specialised models — one per application, per use case, per company.
After Jensen Huang told Lin Qiao that every company must be special to justify its existence, Lin realised the implication was profound: all of a company's product design, data, and user relationships encode irreplaceable private intelligence that no external model can learn.
Chapter 2 · 13:00
Harry asks the valuation question directly: if 90% of enterprise workflows can now be handled by open-weight models at a fraction of the cost, are Anthropic and OpenAI dramatically overvalued? Lin agrees the market is beginning to realise this, then introduces a more uncomfortable concept — 'scaling to bankruptcy.' In the SaaS era, finding product-market fit was the hard problem; scaling was cheap. In the AI era, these are decoupled. Companies with genuine customer demand and willingness to pay can destroy themselves by growing because the AI infrastructure COGS spiral faster than revenue. This creates a structural pull toward open-weight models that companies can control, customise, and optimise.
When Fireworks was founded, open models were in their infancy. Betting on them was a huge gamble. The PyTorch roots gave the team conviction in open ecosystems, and the payoff came as open models crossed quality thresholds that now rival closed models for the vast majority of enterprise use cases.
Lin Qiao suggests that as open models handle 90% of enterprise use cases at a fraction of the cost, the market is beginning to re-examine whether frontier model companies are priced for a world that will actually materialise. The power-line metaphor applies: important, yes — irreplaceable, no.
In the SaaS era, finding product-market fit was the hard part — scaling was cheap. In the AI era, companies with genuine demand can still destroy themselves by growing, because AI infrastructure costs don't scale gracefully. Lin Qiao calls it 'scaling to bankruptcy.'
Chapter 3 · 19:00
Harry — declaring a conflict of interest as a Lagoora investor — uses the Harvey-Lagoora dynamic as a test case. A year ago, not building your own model seemed fine because frontier models were improving so fast. Now the picture is reversed. Lin reframes the question: in legal, where error tolerance is near-zero and workflows are deeply specialised, the competitive advantage isn't whether you own a model but whether you've encoded your proprietary knowledge into the orchestration layer and the fine-tuned models that power it. He points to Cursor as evidence — they were the first coding AI company to tune their own models, and almost all coding companies have followed.
The top six open-weight models globally are currently Chinese-built. Lin Qiao argues that once a model is open, enterprises can wrap their own guardrails around it, but the deeper point is that every model — Chinese or American — encodes its creator's judgment and taste, which always needs tuning.
Lin Qiao believes the future will feature millions of specialised AI models, one per application or use case, rather than a single dominant AGI.
Chapter 5 · 28:00
One of the most technically rich chapters in the episode. Lin explains that early AI tech adopters are hackers who want control, and Cursor — flush with frontier lab researchers — is the prime example. The problem they solved together is profound: hyperscalers run RL training on 100,000 interconnected chips. Cursor and Fireworks had no such cluster. Their solution was to decouple the trainer (which updates model weights) from the RL rollout (which deploys the model into a synthetic environment to collect rewards), distributing the system across five or six global data centre regions and syncing fresh model weights efficiently enough that the reward signal stays numerically sound. This distributed design enabled Cursor's recent model launches without a hyperscaler-level budget.
Fireworks CTO Dima embedded at Cursor for months to build a distributed reinforcement learning infrastructure that decouples the trainer from RL rollout across six global data centre regions. This let a capital-constrained startup run training jobs that previously required 100,000 interconnected chips at a hyperscaler.
Lin Qiao sees 2024 as the year of coding AI and 2025 as the year of co-work. Co-work is dramatically more diverse than coding — spanning legal, finance, healthcare, customer support, and consumer-facing AI — and Fireworks is already landing customers across all of those verticals.
Fireworks processes more than 40 trillion tokens a day today. Lin Qiao projects that number could be 20x to 100x higher by end of next year. At those volumes, worries about a CapEx bubble look completely backwards.
Fireworks processes more than 40 trillion tokens per day, the majority from customised rather than off-the-shelf models.
Chapter 6 · 37:00
Harry opens with the customer concentration question — can Fireworks survive a Cursor churn given a potential SpaceX acquisition? Lin pivots to the broader portfolio: 2024 was the year of coding, 2025 is the year of co-work, and co-work customers span legal, finance, healthcare, consumer-facing recommendation systems, and beyond. He then drops a striking operational disclosure: Fireworks processes more than 40 trillion tokens per day, the majority from customised models. When Harry asks where that number goes in a year, Lin offers a range of 20x to 100x — and argues that at those growth rates, the idea that AI CapEx is in a bubble is simply wrong. The world is bottlenecked by physical supply chains, not by lack of demand.
Lin Qiao projects Fireworks' token throughput could grow 20x to 100x by end of next year as AI adoption accelerates.
Cursor, an early Fireworks customer when it was a single-digit-million company, grew approximately 1,000x over two years.
Marc Benioff reportedly spends about 3.8% of Salesforce developer salaries on Anthropic's Claude Code product.
Chapter 7 · 43:00
This is the most quotable chapter in the episode. Lin explains that current token prices are artificially elevated by supply chain constraints, not by any fundamental cost floor. As competition increases and infrastructure scales, prices will fall sharply. He quantifies three levers: solving tasks requires fewer tokens as models become more precise; Fireworks' platform optimises inference unit economics for each customised workload; and underlying GPU and memory infrastructure will improve structurally over two to three years. The combined effect: a 10x cost reduction over three years that unlocks a 100x usage explosion. He then dives into zero KLD — Fireworks' commitment to bit-exact numerical equivalence from training to inference — as the quality guarantee that justifies its premium over commoditised inference providers.
Token costs are artificially high because of supply chain constraints. Once competition and infrastructure catch up, Lin Qiao expects a 10x cost reduction in three years. Cheaper tokens will unlock use cases that are currently economically impossible, driving a 100x surge in usage.
Lin Qiao predicts token costs will fall 10x over the next three years due to supply chain improvements and infrastructure efficiency.
Lin Qiao argued that a 10x cost reduction in AI tokens will trigger a 100x explosion in usage as AI becomes a utility.
Fireworks achieves zero KL-divergence between training and inference — meaning model weights transfer with bit-exact numerical equivalence. This matters enormously at scale: even tiny quality drops at the train-inference boundary waste customers' entire training investment.
Fireworks AI achieves zero KL-divergence between training and inference, meaning model weights transfer with full numerical equivalence and no quality loss.
Chapter 8 · 49:00
Harry observes that China's lack of regulatory friction means data centres go up in weeks versus years in the US, and asks whether that translates into a structural AI advantage. Lin acknowledges the construction velocity differential but argues the US has comparable specialisation — the bottleneck is global supply chain constraints (electricians, transistors, materials) rather than policy alone. The conversation then broadens to sovereign models: the brief TikTok ban illustrated the existential vulnerability of depending on foreign-controlled infrastructure. Lin's analogy is stark — if your country's AI is like electricity, another country should never control the switch. The principle applies equally to individual enterprises: owning your intelligence is not optional.
Traditional GPU hardware depreciated over six years. Now multiple new SKUs launch annually, and the newest models always prefer the newest chips. This collapse in depreciation cycles fundamentally changes whether companies should own or rent data centre capacity.
Traditional hardware depreciation cycles of 6 years are being upended as multiple GPU SKUs launch within a single year, making older hardware obsolete faster.
Watching TikTok get briefly banned in the US crystallised a broader truth: if your critical intelligence infrastructure sits on another country's AI, it can be cut off overnight. Lin Qiao argues every nation — and every company — needs sovereign AI the same way they need sovereign electricity grids.
Meta has been building its own custom AI chips (MTIA) since at least 2018, well before the generative AI era.
Chapter 9 · 1:00:15
The quickfire round delivers some of the episode's most memorable moments. Lin admits he was wrong to fear growing too fast — aggressive AI tool adoption and a refined hiring filter for extreme ownership changed his calculus. His Jensen Huang insight crystallises into a philosophy: replying to emails in one minute isn't ego, it's the only way to maintain the information freshness needed for precise, fast leadership judgement. His biggest regret is underinvesting in marketing early, framing it not as fluff but as the discipline of educating customers on the right direction. He closes with his sharpest three-year prediction: owning your own intelligence stack will become mandatory for every company, just as owning your own software stack is today. Sponsor reads for JPMorgan, Navan, and Base44 close the episode.
Everyone talks about HBM or energy as the AI bottleneck. Lin Qiao's answer is different: the industry still has no great system designed for 10-trillion-parameter models. Closing that gap requires co-design from model to serving platform to chip system — a level of integration that doesn't yet exist.
Fireworks AI expects to at least double from $800M–$1B ARR by the end of 2026, targeting roughly $2B ARR.
Lin Qiao met George Hu, the former President of Salesforce, a year before hiring him — and turned him down because Fireworks had only 50 people. The relationship evolved through board-level advice before Hu formally joined as Fireworks crossed into hypergrowth. The lesson: build the relationship early, even if the timing isn't right yet.
Fireworks doesn't hire for competence first. They hire for extreme ownership — people who claim end-to-end problems without being asked, see them through to delivery, and treat every outcome as their own responsibility. That mindset, Lin Qiao says, compounds faster than any skill.
Jensen Huang replies to emails within one minute. Lin Qiao spent years marvelling at the sheer volume before understanding the wisdom: in a fast-moving company, information loss between layers is guaranteed. The only antidote is a leader who stays close to the ground and makes precise, fast judgements.
Just as every company builds its own software stack because its problems are unique, every company will eventually own its own AI intelligence stack. Renting general intelligence from a handful of providers will be seen as a competitive liability — not a convenience.
No indexed bits in this chapter.
This episode
Factual claims made this episode, and whether a source was named.
Fireworks AI has scaled to $1 billion in ARR in approximately 4 years.
Fireworks AI processes more than 40 trillion tokens per day, with the majority coming from customised rather than off-the-shelf models.
AI token costs will fall 10x over the next 3 years, and this cost reduction will drive a 100x increase in usage.
Cursor grew approximately 1,000x in revenue over two years, starting from single-digit millions when it first partnered with Fireworks AI.
Marc Benioff spends approximately 3.8% of Salesforce developer salaries on Anthropic's Claude Code.
The top 6 open-source AI models globally are currently Chinese-built.
Fireworks AI raised $1.5 billion at a $17 billion valuation.
Fireworks AI expects to at least double its ARR by the end of 2026.
The industry average time to book a business trip is 45 minutes, versus 7 minutes on Navan's platform.
Fireworks AI was founded when Lin Qiao was 48 years old, after 7 years at Meta.
Meta's MTIA custom AI chip programme has been running since at least 2018, long before the generative AI era.
Traditional GPU hardware used to be depreciated over 6-year cycles; now multiple new GPU SKUs from NVIDIA launch within a single year.
Fireworks AI achieves zero KL-divergence (zero KLD) between training and inference, meaning model weights transfer with bit-exact numerical equivalence.
Fireworks AI's distributed RL training system for Cursor runs across 5 to 6 global data centre regions, tapping into scattered GPUs.
Fireworks AI had approximately 50 employees a year before the episode was recorded, and now has approximately 200.
This episode
Nvidia CEO cited multiple times; Lin Qiao described a key conversation with him about specialisation and his leadership style.
Former President of Salesforce hired by Lin Qiao as a key executive at Fireworks AI to lead GTM at scale.
Benchmark investor who broke his own rule against investing in big-tech directors by backing Lin Qiao and Fireworks AI.
Salesforce founder and investor in Fireworks' new round; cited for spending 3.8% of Salesforce developer salaries on Claude Code.
The central company discussed — an AI inference and specialised intelligence platform that hit $1B ARR with 200 employees.
An AI coding tool and major Fireworks customer that grew 1,000x in 2 years; Fireworks CTO co-built its RL training infrastructure.
Lin Qiao's former employer; discussed as a model for growing into infrastructure ownership and for its MTIA custom chip programme.
Referenced as a company fully committed to AGI and building enterprise AI products like Claude Code; also a Navan customer.
Discussed as the dominant chip provider and as a model training participant (Nemo-chan); also referenced via Jensen Huang's 5-layer AI cake framework.
Discussed as a frontier model lab pursuing AGI, contrasted with Fireworks' specialised intelligence thesis.
George Hu's former employer; Marc Benioff cited for spending 3.8% of developer salaries on Claude Code, illustrating AI tooling spend scale.
SRAM-based ASIC accelerator company recently acquired by NVIDIA; discussed for its complementary role in AI inference alongside GPUs.
Chinese AI lab referenced as part of the open-source AI model landscape and as a company building its own chips.
AI legal tech company discussed as an example of a vertical AI firm that chose to build its own foundation model.
AI legal tech company (a Harry Stebbings portfolio investment) discussed as a counterpoint to Harvey — one that chose not to build its own model.
Lin Qiao's employer before Meta; where he built large-scale data systems and honed product instincts before deciding to start a company.
Open-source ML framework; Lin Qiao was on its founding team at Meta, shaping Fireworks' commitment to open ecosystems.
Referenced as an example of what happens when a foreign-owned digital platform is threatened with a ban, illustrating AI sovereignty risk.
Stats
We scan show notes for social handles, websites and apps. Nothing matched on this episode.
We use essential and analytics cookies to run Vuci. To understand how the site is used: Privacy Policy.
Install Vuci on your phone
Add it to your home screen for a faster, app-like experience.
Install Vuci on your phone
Tap the Share button, then “Add to Home Screen”.
A new version is available
Reload to get the latest Vuci.