Quote · The Twenty Minute VC (20VC): Venture Capital | Startup Funding | The Pitch
20VC: Are OpenAI and Anthropic Overvalued? The Open-Source AI Reality | How Token Costs Will Fall 10x And Usage Will Explode 100x | The Future Is Not One AGI; It's Millions of Specialised Models with Lin Qiao, Founder and CEO @ Fireworks
Where this was said
Will Token Costs Fall 10x—and Unleash 100x More Demand?
At 44:15 · chapter starts 43:00
This is the most quotable chapter in the episode. Lin explains that current token prices are artificially elevated by supply chain constraints, not by any fundamental cost floor. As competition increases and infrastructure scales, prices will fall sharply. He quantifies three levers: solving tasks requires fewer tokens as models become more precise; Fireworks' platform optimises inference unit economics for each customised workload; and underlying GPU and memory infrastructure will improve structurally over two to three years. The combined effect: a 10x cost reduction over three years that unlocks a 100x usage explosion. He then dives into zero KLD — Fireworks' commitment to bit-exact numerical equivalence from training to inference — as the quality guarantee that justifies its premium over commoditised inference providers.
Token costs are artificially high because of supply chain constraints. Once competition and infrastructure catch up, Lin Qiao expects a 10x cost reduction in three years. Cheaper tokens will unlock use cases that are currently economically impossible, driving a 100x surge in usage.
Lin Qiao predicts token costs will fall 10x over the next three years due to supply chain improvements and infrastructure efficiency.
Lin Qiao argued that a 10x cost reduction in AI tokens will trigger a 100x explosion in usage as AI becomes a utility.
Fireworks achieves zero KL-divergence between training and inference — meaning model weights transfer with bit-exact numerical equivalence. This matters enormously at scale: even tiny quality drops at the train-inference boundary waste customers' entire training investment.
Fireworks AI achieves zero KL-divergence between training and inference, meaning model weights transfer with full numerical equivalence and no quality loss.