Starting with the B3-100, NVIDIA began acknowledging the quadratic scaling and now specs FP4 at 3× the throughput of FP8, though the theoretical maximum is 4×.
Snapshot · Dwarkesh Podcast
Starting with the B3-100, NVIDIA began acknowledging the quadratic scaling and now specs FP4 at 3× the throughput of FP8, though the theoretical maximum is 4×.
Where this was said
At 15:15 · chapter starts 0:00
Reiner explains why multiply-accumulate is the fundamental AI chip operation. Demonstrates partial product generation with AND gates and introduces the DADA multiplier using full adders (3-to-2 compressors). Area scales as p×q. [1] — Reiner Pope "The fundamental AI chip operation — multiply-accumulate — requires exactly p×q AND gates to generate partial products, then p×q full adders…" 03:28
The fundamental AI chip operation — multiply-accumulate — requires exactly p×q AND gates to generate partial products, then p×q full adders (3-to-2 compressors) to sum them in a DADA tree. Every atomic step in long multiplication maps directly to a physical logic gate. This is why area scales quadratically with bit width.
A p-bit by q-bit integer multiplier requires exactly p×q AND gates to generate partial products, scaling quadratically with bit width.
A DADA multiplier uses p×q full adders because each full adder eliminates one bit from the partial product tree, reducing 24 input bits to 8 output bits.
Multiplier circuit area scales quadratically with bit width, so halving precision from FP8 to FP4 should yield 4× more throughput. NVIDIA historically reported only 2×. The B3-100 finally moved to 3×, but the theoretical max is 4×. This quadratic scaling is the single biggest reason low-precision AI arithmetic works so well.
Halving the bit precision of a multiplier reduces circuit area quadratically, meaning FP4 should theoretically be 4× faster than FP8, not 2×.
Sam's initial MVP was coded in approximately one week using ChatGPT voice mode and copy-pasting code, with no prior technical experience.
Sam argues Discord is 10x better than email for building relationships with younger users who rarely check their inbox.
Sam's monthly operating costs include Cursor ($200), AI image generation ($100), AI video generation ($200), hosting ($100), email marketing ($80), and AI compute ($300–$500).
Sam recommends copying days of Discord chat history into ChatGPT and prompting it to list recurring pain points as a fast, free market research technique.
Bhanu and his team built approximately 50 free tools to attract search traffic, each linked back to SiteGPT.
With AI coding tools like Cursor, Bhanu can now create a new free marketing tool in less than 5 minutes by referencing existing tools.
Bhanu filters Ahrefs keyword results to show only those with a keyword difficulty below 10, making them realistic ranking targets for any decent website.
Bhanu sets a minimum search volume of 1,000 monthly searches when selecting keywords to target with free tools.
PropGPT averaged 20 downloads per day right after launching on the App Store through influencer marketing.
We use essential and analytics cookies to run Vuci. To understand how the site is used: Privacy Policy.
Install Vuci on your phone
Add it to your home screen for a faster, app-like experience.
Install Vuci on your phone
Tap the Share button, then “Add to Home Screen”.
A new version is available
Reload to get the latest Vuci.