Speaker

Reiner Pope

1 podcast 27 moments 2026
1 episodes
1 podcasts
12 quotes
15 snapshots
1 years active

Appearances over time

1 episodes

Less
More

Episodes

1

Podcasts

Quotes & moments

Technology
Why almost all chip area is wasted on moving data

Reiner Pope – Chip design from the bottom up · May 22, 2026 Technology

In a classic CPU or CUDA core, 7/8 of the circuit area is consumed by MUX circuits just to read and write the register file — not by actual computation. The multiply-accumulate unit you care about is a tiny fraction of the silicon. This is the fundamental problem that Tensor Cores and systolic arrays were invented to solve.

Technology
Multiply-accumulate from first principles: AND gates and DADA multipliers

Reiner Pope – Chip design from the bottom up · May 22, 2026 Technology

The fundamental AI chip operation — multiply-accumulate — requires exactly p×q AND gates to generate partial products, then p×q full adders (3-to-2 compressors) to sum them in a DADA tree. Every atomic step in long multiplication maps directly to a physical logic gate. This is why area scales quadratically with bit width.

Technology
FP4 should be 4× faster than FP8 — not 2×

Reiner Pope – Chip design from the bottom up · May 22, 2026 Technology

Multiplier circuit area scales quadratically with bit width, so halving precision from FP8 to FP4 should yield 4× more throughput. NVIDIA historically reported only 2×. The B3-100 finally moved to 3×, but the theoretical max is 4×. This quadratic scaling is the single biggest reason low-precision AI arithmetic works so well.

Technology
Clock cycles: synchronizing 100 billion transistors every nanosecond

Reiner Pope – Chip design from the bottom up · May 22, 2026 Technology

A chip's global clock forces every transistor to synchronize in lockstep every nanosecond. Without it, two paths through logic could produce outputs that arrive at different times, corrupting results. Pipeline registers let you split logic to raise clock speed, but a feedback loop in the adder shows the hard limit: you can't pipeline your way out of a recurrence.

Technology
Why GPU cores are much smaller than CPU cores

Reiner Pope – Chip design from the bottom up · May 22, 2026 Technology

Most CPU die area goes to cache and branch predictor hardware — not ALUs. The branch predictor must guess a branch outcome 5 clock cycles before it's evaluated so the pipeline doesn't stall. GPUs strip out branch prediction and simplify register files, which is why you can fit thousands of GPU cores where a CPU has a hundred.

Technology
Brains vs chips: the batch size of one

Reiner Pope – Chip design from the bottom up · May 22, 2026 Technology

A GPU runs at GHz clock speeds because it processes batch sizes of 1,000 simultaneously. The brain runs at a much slower clock because it only ever processes one instance of itself. Running a GPU at MHz instead of GHz would give roughly 1,000× less energy consumption — but not a 1,000× improvement in energy efficiency per operation.

Technology
A GPU is just a bunch of tiny TPUs

Reiner Pope – Chip design from the bottom up · May 22, 2026 Technology

A TPU has a few large systolic arrays (MXUs) with a shared vector unit. A GPU is the same architecture tiled many times in miniature: each SM has a Tensor Core (mini-MXU) plus its own vector unit. The GPU trades larger systolic array amortization for more parallel data paths between vector and matrix units.

Analysis

What they talk about

  • Technology 100%