Dwarkesh Podcast

Quote · Dwarkesh Podcast

Reiner Pope – Chip design from the bottom up

Explore episode May 22, 2026

Where this was said

Building a multiply-accumulate from logic gates

At 15:10 · chapter starts 0:00

Reiner explains why multiply-accumulate is the fundamental AI chip operation. Demonstrates partial product generation with AND gates and introduces the DADA multiplier using full adders (3-to-2 compressors). Area scales as p×q.

Technology
Multiply-accumulate from first principles: AND gates and DADA multipliers

Reiner Pope – Chip design from the bottom up · May 22, 2026 Technology

The fundamental AI chip operation — multiply-accumulate — requires exactly p×q AND gates to generate partial products, then p×q full adders (3-to-2 compressors) to sum them in a DADA tree. Every atomic step in long multiplication maps directly to a physical logic gate. This is why area scales quadratically with bit width.

Technology
FP4 should be 4× faster than FP8 — not 2×

Reiner Pope – Chip design from the bottom up · May 22, 2026 Technology

Multiplier circuit area scales quadratically with bit width, so halving precision from FP8 to FP4 should yield 4× more throughput. NVIDIA historically reported only 2×. The B3-100 finally moved to 3×, but the theoretical max is 4×. This quadratic scaling is the single biggest reason low-precision AI arithmetic works so well.

Similar quotes