Lin Qiao predicts token costs will fall 10x over the next three years due to supply chain improvements and infrastructure efficiency.
Snapshot · The Twenty Minute VC (20VC): Venture Capital | Startup Funding | The Pitch
Lin Qiao predicts token costs will fall 10x over the next three years due to supply chain improvements and infrastructure efficiency.
Where this was said
At 44:15 · chapter starts 43:00
This is the most quotable chapter in the episode. Lin explains that current token prices are artificially elevated by supply chain constraints, not by any fundamental cost floor. As competition increases and infrastructure scales, prices will fall sharply. He quantifies three levers: solving tasks requires fewer tokens as models become more precise; Fireworks' platform optimises inference unit economics for each customised workload; and underlying GPU and memory infrastructure will improve structurally over two to three years. The combined effect: a 10x cost reduction over three years that unlocks a 100x usage explosion. He then dives into zero KLD — Fireworks' commitment to bit-exact numerical equivalence from training to inference — as the quality guarantee that justifies its premium over commoditised inference providers.
Token costs are artificially high because of supply chain constraints. Once competition and infrastructure catch up, Lin Qiao expects a 10x cost reduction in three years. Cheaper tokens will unlock use cases that are currently economically impossible, driving a 100x surge in usage.
Lin Qiao argued that a 10x cost reduction in AI tokens will trigger a 100x explosion in usage as AI becomes a utility.
Fireworks achieves zero KL-divergence between training and inference — meaning model weights transfer with bit-exact numerical equivalence. This matters enormously at scale: even tiny quality drops at the train-inference boundary waste customers' entire training investment.
Fireworks AI achieves zero KL-divergence between training and inference, meaning model weights transfer with full numerical equivalence and no quality loss.
The founder attached their SaaS demo to the trending debate about whether AI coding is actually good enough to build a full SaaS product.
Quickly forming opinions on how the Twitter algorithm and platform worked allowed the speaker to grow rapidly on the platform.
Sam's initial MVP was coded in approximately one week using ChatGPT voice mode and copy-pasting code, with no prior technical experience.
Sam argues Discord is 10x better than email for building relationships with younger users who rarely check their inbox.
Sam's monthly operating costs include Cursor ($200), AI image generation ($100), AI video generation ($200), hosting ($100), email marketing ($80), and AI compute ($300–$500).
Sam recommends copying days of Discord chat history into ChatGPT and prompting it to list recurring pain points as a fast, free market research technique.
Bhanu and his team built approximately 50 free tools to attract search traffic, each linked back to SiteGPT.
With AI coding tools like Cursor, Bhanu can now create a new free marketing tool in less than 5 minutes by referencing existing tools.
Bhanu filters Ahrefs keyword results to show only those with a keyword difficulty below 10, making them realistic ranking targets for any decent website.
We use essential and analytics cookies to run Vuci. To understand how the site is used: Privacy Policy.
Install Vuci on your phone
Add it to your home screen for a faster, app-like experience.
Install Vuci on your phone
Tap the Share button, then “Add to Home Screen”.
A new version is available
Reload to get the latest Vuci.