Quote · Moonshots with Peter Diamandis
Urgent Update- AI Sputnik Moment: Kimi K3 Released w/ Emad Mostaque | Ep. 272
Where this was said
Sub-1-Bit Quantization and the Future of Computing Substrates
At 1:09:25 · chapter starts 1:07:40
Alex Wissner-Gross takes the quantization story to its logical extreme: Samsung's recently published NanoQuant model already achieves sub-1-effective-bit-per-weight through sparsity, quantization, and low-rank factorization [1] — Alex Wissner-Gross "Alex Wissner-Gross reports that Samsung's NanoQuant model has already broken the 1-effective-bit-per-weight barrier using sparsity, quantiz…" 1:07:48 . His extrapolation puts sub-1-bit quantization going mainstream within the next year. Emad Mostaque puts a specific marker down: 0.78 bits per weight is his predicted floor, at which point models could be physically etched into silicon — the gate is just absent or present, no power required to store the weight. Dave Blundin adds that if the computation is just multiply-by-1, multiply-by-0, or multiply-by-minus-1, an entirely new class of computing substrates becomes viable — photonic chips, crystals, biological systems, anything that can represent three states. He projects 100 to 10,000x improvement in raw compute within 3 years through quantization and new computing methods, multiplicative with algorithmic improvements. Dave Blundin notes this vaulted to his favorite piece of media ever recorded on the pod.
Alex Wissner-Gross extrapolated that Samsung's NanoQuant model already breaks the 1-effective-bit-per-weight barrier and predicts sub-1-bit quantization will go mainstream within the next year.
Alex Wissner-Gross reports that Samsung's NanoQuant model has already broken the 1-effective-bit-per-weight barrier using sparsity, quantization, and low-rank factorization. His extrapolation: sub-1-bit quantization goes mainstream within the next year. The endpoint, per Emad Mostaque, is 0.78 bits per weight — and at that point, models could be literally etched into silicon.
Dave Blundin projected a 100 to 10,000x increase in raw compute efficiency within 3 years through quantization and new computing methods, multiplicative with algorithmic improvements.