Moonshots with Peter Diamandis

Snapshot · Moonshots with Peter Diamandis

Urgent Update- AI Sputnik Moment: Kimi K3 Released w/ Emad Mostaque | Ep. 272

Explore episode Jul 19, 2026

Where this was said

Bonsai 27B and On-Device AI: Frontier Intelligence in Your Pocket

At 1:03:40 · chapter starts 1:02:05

Emad Mostaque presents the counterpart to KIMI K3's datacenter scale: PrismML's Bonsai 27B, a US startup from Caltech backed by Khosla Ventures that achieved what was previously considered impossible — running a 27-billion-parameter, GPT-5-class model entirely on a smartphone. The technique is ternary quantization: reducing model weights from 16-bit floating point to three values (approximately 1.58 bits), shrinking the model to 6 gigabytes with only a 5% accuracy drop and 4 gigabytes with a 15% drop. The speed bonus is equally remarkable — ternary is 5x faster than 16-bit. Dave Blundin and Emad Mostaque excitedly note that binary and ternary computation opens the door to entirely new computing substrates beyond CMOS silicon — crystals, photonic chips, anything capable of representing -1, 0, or 1. Tencent's parallel achievement with the HiDream-3 team (formerly WizardLM at Microsoft) achieves binary compression of a 300-billion-parameter model with a 5% performance drop. The endpoint: a device the size of a smartphone running KIMI K3-class intelligence by end of 2027.

Similar snapshots