Urgent Update- AI Sputnik Moment: Kimi K3 Released w/ Emad Mostaque | Ep. 272
China's KIMI K3 just hit #1 on AI benchmarks using chips two generations behind NVIDIA's best — proving a 99% cost reduction in frontier model training is real, and the US AI duopoly is officially over.
Jul 19, 20262:08:01
Difficulty: Intermediate
Played
Moonshots with Peter Diamandis
Urgent Update- AI Sputnik Moment: Kimi K3 Released w/ Emad Mostaque | Ep. 272
China's KIMI K3 just hit #1 on AI benchmarks using chips two generations behind NVIDIA's best — proving a 99% cost reduction in frontier model training is real, and the US AI duopoly is officially over.
Jul 19, 20262:08:01
Difficulty: Intermediate
Played
TL;DR
China's Moonshot AI just released KIMI K3 — a 2.8-trillion-parameter open-weight model that jumped to #1 on front-end code benchmarks and landed third on the overall AI cost-performance frontier, leapfrogging Anthropic's Claude while built on chips two generations behind NVIDIA's best[1]— Peter Diamandis"KIMI K3 is a 2.8-trillion-parameter multimodal model from China's Moonshot AI that jumped 17 places on the leaderboard overnight, landing #…"04:25. The Moonshots quintet — Peter Diamandis, Emad Mostaque, Alex Wissner-Gross, Dave Blundin, and Salim Ismail — unpack why this is more disruptive than DeepSeek, how 99% cost reductions in training are now proven at scale[2]— Dave Blundin"99% cost reduction proven at scale: The Keller-Jordan NanoGPT speedrun has reduced GPT-2 training costs by 99%, and KIMI K3 now proves that…"17:00, and why frontier intelligence is now a perishable asset with a shelf life of weeks[3]— Alex Wissner-Gross"You could say attention is still all you need."12:02. The single most urgent takeaway: every company needs to decide right now whether to self-host open-weight models or remain API-dependent.
The Moonshots quintet convenes for an emergency pod on KIMI K3, Moonshot AI's 2.8-trillion-parameter open-weight model that reached #1 on AI benchmarks, the implications for US frontier labs, AI superforecasting breakthroughs, on-device quantization advances, humanoid robot MMA, and orbital data centers.
Chapter list
The episode opens in media res with a rapid-fire montage of the pod's sharpest lines — Salim Ismail declaring frontier intelligence a perishable asset, Dave Blundin calling the event more than a Sputnik moment, and Emad Mostaque framing it as Chinese manufacturing excellence. Peter Diamandis then formally welcomes the full quintet: Alex Wissner-Gross, Dave Blundin, Salim Ismail, Emad Mostaque, and Diamandis himself as host. The tone is set immediately — this is an emergency pod, a departure from the regular schedule, triggered by an event the hosts believe may be among the most consequential in recent AI history. Diamandis thanks the audience for their engagement in the YouTube comments, the Moonshots community's warmth becoming a brief counterpoint to the urgency that follows.
Diamandis walks through the KIMI K3 release: a 2.8-trillion-parameter multimodal model from China's Moonshot AI that jumped 17 places overnight to land #1 in front-end code and six other benchmark domains, all while running on H800 chips two generations behind NVIDIA's current best[1]— Peter Diamandis"KIMI K3 is a 2.8-trillion-parameter multimodal model from China's Moonshot AI that jumped 17 places on the leaderboard overnight, landing #…"04:25. Alex Wissner-Gross delivers the panel's first major insight: the published architecture contains no revolutionary breakthrough — it's still essentially a transformer with well-understood mixture-of-experts and linearized attention innovations[2]— Alex Wissner-Gross"KIMI K3's published architecture contains no revolutionary breakthrough — it's still essentially a transformer with well-understood improve…"05:58. The terrifying implication, which Wissner-Gross puts plainly: if a recognizable transformer can nearly match GPT-5.5 Max on the cost-performance frontier, what exactly are Anthropic and OpenAI spending their enormous capital budgets on? Emad Mostaque frames the achievement not as science but as manufacturing — the same engineering discipline that made Chinese EVs the top-selling car in the UK for a third of the price of a Land Rover. Wissner-Gross adds the competitive landscape observation that the world has moved from an OpenAI-Anthropic duopoly to a free-for-all, with Meta, SpaceX AI, and now Moonshot all on the Pareto optimal frontier.
Dave Blundin walks the panel through the Keller-Jordan speedrun: a GitHub repository where researchers compete to recreate Andrej Karpathy's NanoGPT (a GPT-2 class model) faster and cheaper, having collectively achieved a 99% reduction from the original training cost[1]— Dave Blundin"The Keller-Jordan NanoGPT speedrun has reduced GPT-2 training costs by 99%. Until KIMI K3, nobody knew if that efficiency would scale to fr…"16:10. The critical question had always been whether these efficiency innovations would apply at frontier scale — nobody knew until KIMI K3. Now it's clear: the same principles that got GPT-2 training to 1% of its original cost apply when Elon Musk builds a 10-to-20-trillion-parameter model for billions of dollars. A 1% cost version of effectively the same thing is achievable. Salim Ismail adds his three-point argument: frontier intelligence is now a perishable asset with a shelf life of weeks; enterprises that run traditional evaluation cycles will be three model generations behind before signing a contract; and all the value now resides in architectures that can swap models, not in any single model. Dave Blundin explains that the Muon optimizer further compounds this — stripping irrelevant training data like Taylor Swift concert announcements dramatically reduces compute needed for the same intelligence level, and we're nowhere near done squeezing it.
Dave Blundin lays out his RSI argument with unusual specificity: recursive self-improvement doesn't require Einstein-level AI — a model just needs to be able to improve its own kernel by 10x, which nobody even perceives as 'true AGI'[1]— Dave Blundin"Dave Blundin argues recursive self-improvement doesn't require Einstein-level AI — a model just needs to improve its own kernel by 10x. Tha…"23:35. That threshold was crossed earlier than Fable 5; it was Opus 4.8. Chinese labs could use Opus 4.8 to create KIMI K3. The 10x faster model will be a genius-level AI, and that genius will boost its speed again — flame to fire to sun. Peter Diamandis then asks the direct policy question: will the US government move to constrain Chinese open-weight models from being used in the US? Blundin is skeptical they'll move fast enough before the weights drop in two weeks, predicting White House deliberations are happening in real-time. Salim Ismail connects this to the law of accelerating returns, tracing the history from vacuum tubes to transistors to integrated circuits as nested S-curves — each architecture used to design the next, each reinforcing loop accelerating the collective, and none of it stoppable.
Emad Mostaque brings fresh intelligence from Xi Jinping's speech at the World AI Conference in Shanghai: China is going all-in on open-source AI as a public good for humanity, with model approval times dropping from 60 days to approximately one week[1]— Emad Mostaque"At the World AI Conference in Shanghai, Xi Jinping declared China will fully back open-source AI as a public good for humanity with minimal…"28:25. The motivation is strategic on multiple levels — a billion Chinese citizens whose effective IQ will rise with access to these tools, a demographic crisis that robots can help solve, and the soft power of planting a Chinese-educated AI brain into critical systems worldwide. Xi Jinping simultaneously announced a new international AI regulatory body including Brazil, Asia, and Africa — the new Belt and Road running on AI rather than infrastructure. Alex Wissner-Gross crystallizes the geopolitical irony: 'It's a bizarre future where the Chinese Communist Party is saving American capitalism from itself.' Salim Ismail notes pointedly that Yang Zhilin, KIMI K3's creator, was a CMU graduate — America could have given him a green card.
The episode pauses for a sponsor read from Dave Blundin for Blitzy — positioned as an AI-native pre-IDE development platform that uses thousands of specialized AI agents to understand million-line enterprise codebases. The platform handles 80% or more of development work autonomously while guiding the remaining 20% requiring human input. Enterprises incorporating Blitzy report a 5x increase in engineering velocity. The ad is followed by discussion returning to the Yang Zhilin immigration story and how Moonshot AI's compute architecture was built to optimize for Huawei Ascend chips, with American inference providers set to run KIMI K3 10 times cheaper than Chinese competitors once they optimize it for NVIDIA Blackwell/Vera Rubin hardware.
Alex Wissner-Gross argues the NVIDIA chip embargo is a textbook example of a half-measure that produces the worst outcome: it incentivized Chinese labs to develop breakthrough quantization and efficiency research that now makes their models cheaper and better than if the embargo had never happened[1]— Alex Wissner-Gross"The US export controls on advanced NVIDIA chips achieved the worst possible outcome: they irritated China without actually stopping it, jus…"37:25. Dave Blundin draws the Vietnam War analogy — you either go all-in and win quickly or you don't engage, but you never creep in with partial measures. The conversation shifts to Yang Zhilin: Diamandis presents the narrative that CMU trained him and America let him go. Alex Wissner-Gross adds nuance — Yang actually founded Recurrent AI in China while still a PhD student at CMU in 2016, suggesting he may have always planned to return. But the group agrees the broader principle stands: 70% of elite AI researchers are not US citizens, and the asymmetric advantage of building in America is eroding. Dave Blundin adds the data point that 80% of Chinese PhD graduates return to China — unlike Indian graduates who overwhelmingly stay.
Diamandis presents the data on frontier model release acceleration: 13 new frontier models since mid-April 2026 (one every 10 days), versus one every 50 days in 2025 and one every 60 days in 2024[1]— Alex Wissner-Gross"Alex Wissner-Gross regressed an exponential curve to frontier model release frequency: one every 60 days in 2024, one every 50 days in 2025…"54:15. Elon Musk's tweet about his $2 trillion model finishing training 'next week' — potentially exceeding KIMI K3 — is offered as live evidence. Alex Wissner-Gross delivers his extrapolation: regressing an exponential curve to the release frequency data projects daily new frontier model releases by January 2027. The panel grapples with what 'a new frontier model' even means at daily frequency — the conclusion is that continuous versioning makes individual release announcements obsolete, and the conversation must shift to use-case demonstrations. Diamandis's son Jet is cited as the millennial voice: 'Another release, a little bit better — Dad, come on.' The panel agrees he's right that raw benchmark jumps will lose meaning, and application demos will become the new currency of attention.
Diamandis plays KIMI K3 live demos that have gone viral: browser-based recreations of classic games and web apps simulating an Apple desktop, all generated from single prompts. The panel notes the social media explosion of people sharing their own KIMI K3 creations in the 24 hours since launch. Emad Mostaque offers a caveat that one-shotting a game is very different from building the ecosystem, customer service, and marketing around it — passion and domain knowledge still matter. Dave Blundin extends the point: when you can one-shot anything, the real cognitive challenge is deciding what you want — executives and founders who have never had unlimited execution ability are suddenly confronted with pure strategic choice. Alex Wissner-Gross proposes a call to action: instead of outro music videos, listeners should submit AI-generated outro video games built on KIMI K3. Diamandis closes the segment with a message about purpose over passion — you don't need to be a computer scientist to make a dent in the universe, you need a massive transformative purpose.
Emad Mostaque presents the counterpart to KIMI K3's datacenter scale: PrismML's Bonsai 27B, a US startup from Caltech backed by Khosla Ventures that achieved what was previously considered impossible — running a 27-billion-parameter, GPT-5-class model entirely on a smartphone[1]— Emad Mostaque"PrismML's Bonsai 27B compresses a GPT-5-class 27-billion-parameter model to just 6 gigabytes via ternary quantization — a 5% accuracy drop …"1:00:40. The technique is ternary quantization: reducing model weights from 16-bit floating point to three values (approximately 1.58 bits), shrinking the model to 6 gigabytes with only a 5% accuracy drop and 4 gigabytes with a 15% drop. The speed bonus is equally remarkable — ternary is 5x faster than 16-bit. Dave Blundin and Emad Mostaque excitedly note that binary and ternary computation opens the door to entirely new computing substrates beyond CMOS silicon — crystals, photonic chips, anything capable of representing -1, 0, or 1. Tencent's parallel achievement with the HiDream-3 team (formerly WizardLM at Microsoft) achieves binary compression of a 300-billion-parameter model with a 5% performance drop. The endpoint: a device the size of a smartphone running KIMI K3-class intelligence by end of 2027.
Alex Wissner-Gross takes the quantization story to its logical extreme: Samsung's recently published NanoQuant model already achieves sub-1-effective-bit-per-weight through sparsity, quantization, and low-rank factorization[1]— Alex Wissner-Gross"Alex Wissner-Gross reports that Samsung's NanoQuant model has already broken the 1-effective-bit-per-weight barrier using sparsity, quantiz…"1:07:48. His extrapolation puts sub-1-bit quantization going mainstream within the next year. Emad Mostaque puts a specific marker down: 0.78 bits per weight is his predicted floor, at which point models could be physically etched into silicon — the gate is just absent or present, no power required to store the weight. Dave Blundin adds that if the computation is just multiply-by-1, multiply-by-0, or multiply-by-minus-1, an entirely new class of computing substrates becomes viable — photonic chips, crystals, biological systems, anything that can represent three states. He projects 100 to 10,000x improvement in raw compute within 3 years through quantization and new computing methods, multiplicative with algorithmic improvements. Dave Blundin notes this vaulted to his favorite piece of media ever recorded on the pod.
Peter Diamandis steps into the health segment with Dr. Dawn Musalem, CMO of Fountain Life, discussing the case for proactive cancer screening. The headline statistic: 3.3% of Fountain Life members who believed themselves healthy were found to have cancers they didn't know about. Musalem emphasizes that the majority of cancers taking lives are those found at late stages — early detection dramatically improves cure rates. Fountain Life uses full-body MRI and early cancer detection screening tools not typically covered by conventional insurance. Diamandis frames it in his characteristic way: not knowing what's inside your body is like driving with your eyes closed. Listeners are directed to fountainlife.com/peter to learn about memberships. The segment closes with a call for healthcare democratization as data from Fountain Life's member database builds the evidence base for insurance coverage.
Peter Diamandis presents the Forecast Bench update: for the first time, several AI models are statistically indistinguishable from human superforecasters — ordinary people who, through disciplined probabilistic reasoning, consistently outpredict CIA analysts with classified information (scored on a Brier scale where lower is better)[1]— Peter Diamandis"The Forecasting Research Institute's Forecast Bench now shows multiple AI models statistically indistinguishable from human superforecaster…"1:20:30. The #1 AI superforecaster is from a British startup named Cassie, founded by a British intelligence officer inspired by Philip Tetlock's superforecasting research. Wissner-Gross extends to hyperforecasting: what happens when AI that already dominates high-volume public securities trading by volume also has better internal models of collective human behavior than humans do? It can predict the next action of humanity collectively faster than humanity can take it — and the most interesting domain is capital markets, where prediction preemptively shapes the action being predicted, potentially fulfilling the Efficient Market Hypothesis. Emad Mostaque connects this to his book The Last Economy, where he derives all of economics from generative AI mathematics, noting that psychohistory's requirement for large-population modeling is the same math as diffusion models. Salim Ismail warns that once AI forecasting enters management decisions, senior management expertise becomes obsolete — replaced by bias-free machine judgment.
Diamandis plays a viral video of China's humanoid robot MMA combat event, where Engine AI T-800 robots weighing 70kg and capable of punching four times harder than Mike Tyson engage in MMA-style combat, one literally kicking the other's head off[1]— Peter Diamandis"China has staged public MMA-style humanoid robot combat events that went viral — one robot literally kicking another's head off. With 150 h…"1:23:50. Alex Wissner-Gross expresses genuine concern on two levels: first, that staging humanoid robot combat sets a dangerous inductive prior for future autonomous embodied intelligences; second, and more urgently, that these same robots are likely the prototype for PLA infantry of the near future. He calls on the West to develop its own humanoid robot spectacle events — he helped organize ProRL's first humanoid mini marathon in Boston's Seaport — focused on economically productive competition rather than combat. Emad Mostaque adds the regulatory gap: these robots can legitimately kill someone, walk into a building, and ring a doorbell, with zero legal framework governing their street deployment. He frames Unitree's 11,000 total production as evidence we are at the absolute beginning — 11 million per year is just years away.
Diamandis presents a viral chart comparing US data center water consumption (17 billion gallons per year, per Lawrence Berkeley National Labs) against American golf courses (531 billion gallons — 31 times more) and California almond farming (1 trillion gallons — 60 times more)[1]. Salim Ismail calls the anti-AI water narrative 'completely non-data-driven garbage bullshit,' while Dave Blundin warns that the water argument is so easily disproved that the mob will simply move to a different, semi-sane but still wrong complaint — and that the real danger is a populist moratorium movement of the kind that paralyzed nuclear power. Emad Mostaque adds that McDonald's Big Mac production uses roughly twice the water of all golf courses combined, suggesting the underlying concern is not really about water. Alex Wissner-Gross raises the possibility of a foreign influence operation behind the anti-AI data center narrative, noting the Dyson Swarm would solve the water issue entirely by moving compute to closed-loop orbital systems.
Diamandis plays a clip of Sam Altman calling space data centers 'ridiculous' given current launch costs and GPU repair logistics, then asks the panel for their reaction[1]— Alex Wissner-Gross"Sam Altman says orbital data centers are 'ridiculous' this decade. Elon says they'll launch in 2 years. But listen closely: Altman says spa…"1:35:45. Dave Blundin parses the words carefully: Altman says space won't matter 'this decade,' which has only 3.5 years left — and Elon's 2-to-3-year launch timeline agrees. Wissner-Gross identifies the real story as OpenAI's conflict of interest: Project Stargate has been rebranded from owning data centers to merely leasing terrestrial capacity; OpenAI is delaying its IPO; and Anthropic — unlike OpenAI — has signed collaboration agreements with SpaceX AI for orbital compute. Wissner-Gross predicts OpenAI's tune on orbital data centers will 'conveniently change' in 2–3 years, right as the unit economics cross over. Emad Mostaque suggests OpenAI will likely acquire a satellite data company like Planet Labs to enter the space compute race when it becomes strategically necessary. The Starship 13 near-launch abort at T-minus zero is discussed as an engineering triumph — the ability to safe the vehicle and replan within days rather than months.
The panel works through seven questions sourced from Diamandis's X post. Emad Mostaque on cost-effectiveness: KIMI K3 at $15/million tokens uses twice the tokens as GPT-5.6, but expects a 10–50x price drop as US inference providers optimize it for NVIDIA hardware[1]— Emad Mostaque"KIMI K3 price: $15/million tokens: KIMI K3 costs $15 per million tokens currently, compared to DeepSeek at $1, Sonnet at $20, Opus at $40, …"1:46:55. Salim Ismail on valuation: he maintains his estimate of 75% destruction from US frontier lab valuations, putting OpenAI at ~$250 billion from $1 trillion. Dave Blundin on open-source dynamics: large US model providers won't go open source, but KIMI K3 gives every enterprise a viable self-hosting path, accelerating the split between frontier-seekers (who pay Anthropic for Fable 5 to solve the hardest problems) and cost-optimizers (who self-host K3). Alex Wissner-Gross on US lab justification: the premise that valuations shrink due to K3 is fallacious — as with DeepSeek's original shock, Jevons paradox ultimately increases chip demand and total AI spend. On blocking K3: Wissner-Gross says it could be effectively constrained if corporations are required to disclose Chinese open-weight model use and face SEC-level scrutiny, but finds this highly undesirable. Salim Ismail dismisses blocking as futile — mirrors, peer-to-peer, VPNs make it trivially circumventable, and the only effect is denying American researchers and startups access.
As the panel prepares to close, Dave Blundin distills the episode's practical implication: every company in the world needs to decide right now whether to build internal AI on open-weight models or remain dependent on closed APIs — and KIMI K3 proves the open-weight path is viable at frontier level[1]— Salim Ismail"Frontier AI model performance is now a perishable commodity — state-of-the-art lasts weeks, not years. Any enterprise that runs a tradition…"17:55. He urges a crash program of evaluation with the best available advisors. Salim Ismail adds a simple two-installation suggestion: deploy KIMI K3 and Inkling simultaneously, fine-tune both on proprietary internal data, and start building the learning loop that will become the company's AI moat. Alex Wissner-Gross closes on the science angle, declaring large swaths of the sciences are now 'thoroughly cooked' and promising more on that subject soon from his and Diamandis's book Solve Everything. Diamandis wraps with thanks to the audience, a subscription request, and the observation that the pace of releases is already pushing toward daily emergency pods. He frames the final message around the abundance mindset — KIMI K3 is a gift, not a threat.
Diamandis goes around the table for end-of-week updates. Emad Mostaque has a batch of research papers nearly ready for release through Intelligent Internet. Salim Ismail's next Meaning of Life session is Tuesday at 7 PM Eastern at openexo.com/mol — described as taking listeners 'beyond AI into philosophy and theology.' Alex Wissner-Gross is deep in Solve Everything, his and Diamandis's book, and announces an upcoming AMA with Diamandis's Abundance community alongside a full-day Moonshots gathering in Los Angeles on September 25th with the entire quintet. Dave Blundin notes that the quantization discussion on this episode has become his favorite piece of media ever recorded — surpassing Leopold Aschenbrenner — and shares that MIT Nano's Vlad Bulović is joining as an advisor to a new photonic computing startup being built at Link Ventures, with two MIT researchers being recruited to pair with their Princeton team. Diamandis closes: 'There is no time to sleep during the singularity.'
After the English-language episode concludes, a German-language sponsor segment plays featuring a humorous skit in which speakers discuss the stress of various relaxation activities before discovering that doing their taxes with the WISO Steuer app is unexpectedly calming. One speaker claims to have received over €1,000 back with minimal effort. The ad highlights the app's automation features and notes the tax deadline of July 31st. This appears to be a dynamic ad insertion targeting German-speaking listeners of the Moonshots podcast, unrelated to the episode's content.
Mixture of Experts (MoE)
A neural network architecture where only a subset of 'expert' sub-networks are activated for any given input, dramatically reducing compute while maintaining large overall model capacity.
Muon Optimizer
A training optimization algorithm used in KIMI K3 that improves data-to-intelligence conversion by stripping irrelevant training tokens, reducing computation needed to achieve a given intelligence level.
Ternary quantization
A model compression technique that reduces each neural network weight to one of three values (e.g., -1, 0, 1, approximately 1.58 bits), drastically shrinking model size and boosting inference speed.
Linearized attention
A modification to the standard transformer attention mechanism that reduces computational complexity from quadratic to linear in sequence length, enabling faster processing of long contexts.
Pareto optimal frontier
In the context of AI models, the set of models where no other model is both cheaper per task AND more capable — i.e., the best cost-performance trade-off curve. KIMI K3 reached 3rd place on this frontier.
Distillation attack
The practice of training a model by capturing and learning from the reasoning traces or outputs of a more powerful model, effectively transferring knowledge without direct access to the original model's weights.
RSI (Recursive Self-Improvement)
The hypothetical ability of an AI system to iteratively improve its own capabilities, each improvement enabling further improvements — a potential pathway to rapid, unbounded intelligence growth.
Brier score
A scoring rule for probabilistic predictions where lower scores indicate better accuracy; used to rank superforecasters and now AI models on prediction benchmarks like Forecast Bench.
Psychohistory
Isaac Asimov's fictional science of predicting the behavior of large populations using mathematics; invoked here as an analogy for AI-powered macroeconomic and societal forecasting.
Hyperforecasting
Alex Wissner-Gross's term for AI-powered forecasting that goes beyond human superforecaster ability to predict — and potentially preemptively shape — collective human behavior and capital markets.
Sub-1-bit quantization
A model compression approach achieving less than 1 effective bit per neural network weight through sparsity, quantization, and low-rank factorization — exemplified by Samsung's NanoQuant model.
Efficient Market Hypothesis (EMH)
The economic theory that asset prices fully reflect all available information; Alex Wissner-Gross argues hyperforecasters will finally make EMH true by eliminating all residual market inefficiency.
NanoGPT speedrun (Keller-Jordan)
A GitHub repository where researchers compete to recreate GPT-2 faster and cheaper; innovations there have cut original GPT-2 training costs by 99% and proved transferable to frontier-scale models.
VLA model
Vision-Language-Action model — an AI architecture that combines visual perception, language understanding, and physical action control, relevant to robotics latency discussions in this episode.
FINRA
Financial Industry Regulatory Authority — a self-regulatory organization under SEC oversight that governs broker-dealers; proposed as a model for a new AI frontier lab regulatory body.
Reflexivity
In economics, the phenomenon where predictions about a system (like markets) influence the system itself, creating feedback loops; used here regarding AI forecasters shaping the markets they predict.
Perishable asset
Something whose value degrades rapidly over time; Salim Ismail uses it to describe frontier AI model performance, which now becomes obsolete within weeks rather than years.
Jevons paradox
The economic observation that increased efficiency in using a resource tends to increase total consumption of that resource rather than decrease it; cited here to explain why chip demand rises despite AI cost reductions.
Harebrained
Reckless or poorly conceived; used by Dave Blundin to describe the US NVIDIA chip embargo strategy against China as ill-thought-out and counterproductive.
Massive Transformative Purpose (MTP)
Peter Diamandis and Salim Ismail's concept for a goal that is both personally meaningful and delivers broad societal benefit — distinguished from a 'passion' (something you love) by its outward impact.
Chapter 1 · 00:00
Intro & The AI Sputnik Moment
The episode opens in media res with a rapid-fire montage of the pod's sharpest lines — Salim Ismail declaring frontier intelligence a perishable asset, Dave Blundin calling the event more than a Sputnik moment, and Emad Mostaque framing it as Chinese manufacturing excellence. Peter Diamandis then formally welcomes the full quintet: Alex Wissner-Gross, Dave Blundin, Salim Ismail, Emad Mostaque, and Diamandis himself as host. The tone is set immediately — this is an emergency pod, a departure from the regular schedule, triggered by an event the hosts believe may be among the most consequential in recent AI history. Diamandis thanks the audience for their engagement in the YouTube comments, the Moonshots community's warmth becoming a brief counterpoint to the urgency that follows.
KIMI K3 is a 2.8-trillion-parameter multimodal model from China's Moonshot AI that jumped 17 places on the leaderboard overnight, landing #1 in front-end code and 6 other domains. It achieved this while running on chips two generations behind NVIDIA's best — a stunning proof that engineering discipline beats raw compute.
KIMI K3 is a 2.8-trillion-parameter multimodal model released by China's Moonshot AI, making it one of the largest models ever released and ranking #1 on front-end code benchmarks.
Chapter 2 · 04:30
KIMI K3 Deep Dive: Architecture, Benchmarks & Competitive Position
Diamandis walks through the KIMI K3 release: a 2.8-trillion-parameter multimodal model from China's Moonshot AI that jumped 17 places overnight to land #1 in front-end code and six other benchmark domains, all while running on H800 chips two generations behind NVIDIA's current best[1]— Peter Diamandis"KIMI K3 is a 2.8-trillion-parameter multimodal model from China's Moonshot AI that jumped 17 places on the leaderboard overnight, landing #…"04:25. Alex Wissner-Gross delivers the panel's first major insight: the published architecture contains no revolutionary breakthrough — it's still essentially a transformer with well-understood mixture-of-experts and linearized attention innovations[2]— Alex Wissner-Gross"KIMI K3's published architecture contains no revolutionary breakthrough — it's still essentially a transformer with well-understood improve…"05:58. The terrifying implication, which Wissner-Gross puts plainly: if a recognizable transformer can nearly match GPT-5.5 Max on the cost-performance frontier, what exactly are Anthropic and OpenAI spending their enormous capital budgets on? Emad Mostaque frames the achievement not as science but as manufacturing — the same engineering discipline that made Chinese EVs the top-selling car in the UK for a third of the price of a Land Rover. Wissner-Gross adds the competitive landscape observation that the world has moved from an OpenAI-Anthropic duopoly to a free-for-all, with Meta, SpaceX AI, and now Moonshot all on the Pareto optimal frontier.
KIMI K3's published architecture contains no revolutionary breakthrough — it's still essentially a transformer with well-understood improvements in mixture of experts and linearized attention. The terrifying implication: if you can nearly match GPT-5.5 Max with a recognizable transformer, what are the American frontier labs spending billions of dollars on?
Emad Mostaque argues KIMI K3's success mirrors the rise of Chinese EVs: not a scientific revolution, but relentless manufacturing discipline applied to known ingredients. The number one car in the UK is now the Jaiqoo J7 — 'Temu Land Rover' — fully loaded for a third of the price. KIMI K3 is the AI equivalent.
The 99% Cost Reduction: Frontier AI Is Now 1% of the Price
Dave Blundin walks the panel through the Keller-Jordan speedrun: a GitHub repository where researchers compete to recreate Andrej Karpathy's NanoGPT (a GPT-2 class model) faster and cheaper, having collectively achieved a 99% reduction from the original training cost[1]— Dave Blundin"The Keller-Jordan NanoGPT speedrun has reduced GPT-2 training costs by 99%. Until KIMI K3, nobody knew if that efficiency would scale to fr…"16:10. The critical question had always been whether these efficiency innovations would apply at frontier scale — nobody knew until KIMI K3. Now it's clear: the same principles that got GPT-2 training to 1% of its original cost apply when Elon Musk builds a 10-to-20-trillion-parameter model for billions of dollars. A 1% cost version of effectively the same thing is achievable. Salim Ismail adds his three-point argument: frontier intelligence is now a perishable asset with a shelf life of weeks; enterprises that run traditional evaluation cycles will be three model generations behind before signing a contract; and all the value now resides in architectures that can swap models, not in any single model. Dave Blundin explains that the Muon optimizer further compounds this — stripping irrelevant training data like Taylor Swift concert announcements dramatically reduces compute needed for the same intelligence level, and we're nowhere near done squeezing it.
The Keller-Jordan NanoGPT speedrun has reduced GPT-2 training costs by 99%. Until KIMI K3, nobody knew if that efficiency would scale to frontier models. Now it's proven. A 10-trillion-parameter model that cost Elon Musk billions can theoretically be replicated for 1% of the price — and that realization changes the entire economics of the AI industry.
The Keller-Jordan NanoGPT speedrun has reduced GPT-2 training costs by 99%, and KIMI K3 now proves that same cost efficiency translates to frontier-scale models.
Frontier AI model performance is now a perishable commodity — state-of-the-art lasts weeks, not years. Any enterprise that runs a traditional RFP process before deploying a model is already three generations behind before they sign the contract. All the value now lies in architectures that can swap models instantly, not in any single model.
Salim Ismail argued that frontier intelligence is now a totally perishable asset, with a shelf life of only weeks, making traditional enterprise model evaluation cycles obsolete.
Most frontier model training data is garbage — Taylor Swift concerts, wedding announcements, random internet noise that doesn't drive intelligence and may actually slow training. The Muon optimizer strips down training to relevant data, cutting the computation needed for the same intelligence level. KIMI K3 proves this works at scale, and we're nowhere near done squeezing it.
19:33
20:05
Chapter 4 · 21:20
Recursive Self-Improvement and the US Government's Too-Late Response
Dave Blundin lays out his RSI argument with unusual specificity: recursive self-improvement doesn't require Einstein-level AI — a model just needs to be able to improve its own kernel by 10x, which nobody even perceives as 'true AGI'[1]— Dave Blundin"Dave Blundin argues recursive self-improvement doesn't require Einstein-level AI — a model just needs to improve its own kernel by 10x. Tha…"23:35. That threshold was crossed earlier than Fable 5; it was Opus 4.8. Chinese labs could use Opus 4.8 to create KIMI K3. The 10x faster model will be a genius-level AI, and that genius will boost its speed again — flame to fire to sun. Peter Diamandis then asks the direct policy question: will the US government move to constrain Chinese open-weight models from being used in the US? Blundin is skeptical they'll move fast enough before the weights drop in two weeks, predicting White House deliberations are happening in real-time. Salim Ismail connects this to the law of accelerating returns, tracing the history from vacuum tubes to transistors to integrated circuits as nested S-curves — each architecture used to design the next, each reinforcing loop accelerating the collective, and none of it stoppable.
Dave Blundin argues recursive self-improvement doesn't require Einstein-level AI — a model just needs to improve its own kernel by 10x. That threshold was crossed earlier than Fable 5: it was Opus 4.8. Chinese labs used that spark to build KIMI K3. When we look back in history, right around Opus 4.8 will be the moment the flame was lit.
At the World AI Conference in Shanghai, Xi Jinping declared China will fully back open-source AI as a public good for humanity with minimal regulation. Model approval times have dropped from 60 days to one week. China is deliberately exporting a Chinese-educated AI brain into every critical system on Earth — a soft power strategy the West has no answer to.
28:25
30:02
Chapter 5 · 28:40
Xi Jinping, Open Source, and China's New Belt and Road for AI
Emad Mostaque brings fresh intelligence from Xi Jinping's speech at the World AI Conference in Shanghai: China is going all-in on open-source AI as a public good for humanity, with model approval times dropping from 60 days to approximately one week[1]— Emad Mostaque"At the World AI Conference in Shanghai, Xi Jinping declared China will fully back open-source AI as a public good for humanity with minimal…"28:25. The motivation is strategic on multiple levels — a billion Chinese citizens whose effective IQ will rise with access to these tools, a demographic crisis that robots can help solve, and the soft power of planting a Chinese-educated AI brain into critical systems worldwide. Xi Jinping simultaneously announced a new international AI regulatory body including Brazil, Asia, and Africa — the new Belt and Road running on AI rather than infrastructure. Alex Wissner-Gross crystallizes the geopolitical irony: 'It's a bizarre future where the Chinese Communist Party is saving American capitalism from itself.' Salim Ismail notes pointedly that Yang Zhilin, KIMI K3's creator, was a CMU graduate — America could have given him a green card.
The episode pauses for a sponsor read from Dave Blundin for Blitzy — positioned as an AI-native pre-IDE development platform that uses thousands of specialized AI agents to understand million-line enterprise codebases. The platform handles 80% or more of development work autonomously while guiding the remaining 20% requiring human input. Enterprises incorporating Blitzy report a 5x increase in engineering velocity. The ad is followed by discussion returning to the Yang Zhilin immigration story and how Moonshot AI's compute architecture was built to optimize for Huawei Ascend chips, with American inference providers set to run KIMI K3 10 times cheaper than Chinese competitors once they optimize it for NVIDIA Blackwell/Vera Rubin hardware.
The US export controls on advanced NVIDIA chips achieved the worst possible outcome: they irritated China without actually stopping it, just as the US learned in Vietnam that half-measures make things worse. The embargo forced Chinese labs to develop breakthrough quantization research that now makes their models cheaper, faster, and more efficient than if the embargo had never happened.
Emad Mostaque stated that the total compute used to train KIMI K3 is equivalent to that used to train Mira Murati's Inkling, despite being far more capable — evidence of extraordinary data efficiency.
Salim Ismail traces the history from biological intelligence to individual to collective to artificial intelligence, arguing that every attempt by any entity to constrain an information-based paradigm fails. Intelligence, like information, wants to be free. The scarcity mindset driving AI export controls and regulations is not just counterproductive — it contradicts a fundamental law of nature.
NVIDIA Embargo Backfire and the US Immigration Failure
Alex Wissner-Gross argues the NVIDIA chip embargo is a textbook example of a half-measure that produces the worst outcome: it incentivized Chinese labs to develop breakthrough quantization and efficiency research that now makes their models cheaper and better than if the embargo had never happened[1]— Alex Wissner-Gross"The US export controls on advanced NVIDIA chips achieved the worst possible outcome: they irritated China without actually stopping it, jus…"37:25. Dave Blundin draws the Vietnam War analogy — you either go all-in and win quickly or you don't engage, but you never creep in with partial measures. The conversation shifts to Yang Zhilin: Diamandis presents the narrative that CMU trained him and America let him go. Alex Wissner-Gross adds nuance — Yang actually founded Recurrent AI in China while still a PhD student at CMU in 2016, suggesting he may have always planned to return. But the group agrees the broader principle stands: 70% of elite AI researchers are not US citizens, and the asymmetric advantage of building in America is eroding. Dave Blundin adds the data point that 80% of Chinese PhD graduates return to China — unlike Indian graduates who overwhelmingly stay.
Yang Zhilin earned his PhD at Carnegie Mellon — one of the world's top CS programs — then founded Recurrent AI in China while still a student and went back to build Moonshot AI. Salim Ismail points out that 70% of elite AI researchers are not US citizens; stapling a green card to every STEM PhD is the highest-ROI immigration policy change America could make. Instead, we keep sending them home.
Dave Blundin cited a report noting that approximately 80% of Chinese PhD graduates who study in the US return to China, compared to Indian graduates who overwhelmingly stay.
Salim Ismail stated that 70% of elite AI researchers are not US citizens, with Chinese, Indian, Taiwanese, and British researchers dominating — highlighting a critical immigration policy failure.
Chapter 8 · 53:20
The Accelerating Release Pace and What It Means
Diamandis presents the data on frontier model release acceleration: 13 new frontier models since mid-April 2026 (one every 10 days), versus one every 50 days in 2025 and one every 60 days in 2024[1]— Alex Wissner-Gross"Alex Wissner-Gross regressed an exponential curve to frontier model release frequency: one every 60 days in 2024, one every 50 days in 2025…"54:15. Elon Musk's tweet about his $2 trillion model finishing training 'next week' — potentially exceeding KIMI K3 — is offered as live evidence. Alex Wissner-Gross delivers his extrapolation: regressing an exponential curve to the release frequency data projects daily new frontier model releases by January 2027. The panel grapples with what 'a new frontier model' even means at daily frequency — the conclusion is that continuous versioning makes individual release announcements obsolete, and the conversation must shift to use-case demonstrations. Diamandis's son Jet is cited as the millennial voice: 'Another release, a little bit better — Dad, come on.' The panel agrees he's right that raw benchmark jumps will lose meaning, and application demos will become the new currency of attention.
Alex Wissner-Gross regressed an exponential curve to frontier model release frequency: one every 60 days in 2024, one every 50 days in 2025, one every 10 days now. The projection: daily new frontier model releases by January 2027. At that point, 'new model releases' stop being news and the conversation has to shift to what AI is actually doing.
Alex Wissner-Gross extrapolated that if the current exponential trend in frontier model release frequency continues, the world will see daily new frontier model releases by January 2027.
Chapter 9 · 1:00:10
KIMI K3 Front-End Demos and the Democratization of Making
Diamandis plays KIMI K3 live demos that have gone viral: browser-based recreations of classic games and web apps simulating an Apple desktop, all generated from single prompts. The panel notes the social media explosion of people sharing their own KIMI K3 creations in the 24 hours since launch. Emad Mostaque offers a caveat that one-shotting a game is very different from building the ecosystem, customer service, and marketing around it — passion and domain knowledge still matter. Dave Blundin extends the point: when you can one-shot anything, the real cognitive challenge is deciding what you want — executives and founders who have never had unlimited execution ability are suddenly confronted with pure strategic choice. Alex Wissner-Gross proposes a call to action: instead of outro music videos, listeners should submit AI-generated outro video games built on KIMI K3. Diamandis closes the segment with a message about purpose over passion — you don't need to be a computer scientist to make a dent in the universe, you need a massive transformative purpose.
PrismML's Bonsai 27B compresses a GPT-5-class 27-billion-parameter model to just 6 gigabytes via ternary quantization — a 5% accuracy drop for a file smaller than most video games. With no internet connection required, you can carry a 110-IQ AI assistant in your pocket permanently. Ternary also delivers a 5x speed boost over 16-bit models.
Bonsai 27B and On-Device AI: Frontier Intelligence in Your Pocket
Emad Mostaque presents the counterpart to KIMI K3's datacenter scale: PrismML's Bonsai 27B, a US startup from Caltech backed by Khosla Ventures that achieved what was previously considered impossible — running a 27-billion-parameter, GPT-5-class model entirely on a smartphone[1]— Emad Mostaque"PrismML's Bonsai 27B compresses a GPT-5-class 27-billion-parameter model to just 6 gigabytes via ternary quantization — a 5% accuracy drop …"1:00:40. The technique is ternary quantization: reducing model weights from 16-bit floating point to three values (approximately 1.58 bits), shrinking the model to 6 gigabytes with only a 5% accuracy drop and 4 gigabytes with a 15% drop. The speed bonus is equally remarkable — ternary is 5x faster than 16-bit. Dave Blundin and Emad Mostaque excitedly note that binary and ternary computation opens the door to entirely new computing substrates beyond CMOS silicon — crystals, photonic chips, anything capable of representing -1, 0, or 1. Tencent's parallel achievement with the HiDream-3 team (formerly WizardLM at Microsoft) achieves binary compression of a 300-billion-parameter model with a 5% performance drop. The endpoint: a device the size of a smartphone running KIMI K3-class intelligence by end of 2027.
PrismML's Bonsai 27B achieved ternary quantization to compress a 27-billion-parameter model to just 6 gigabytes with only a 5% accuracy drop, enabling it to run entirely on a smartphone.
Reducing model weights from 16-bit to 3-bit (ternary) results in a 5x improvement in inference speed, in addition to the massive size reduction.
Chapter 11 · 1:07:40
Sub-1-Bit Quantization and the Future of Computing Substrates
Alex Wissner-Gross takes the quantization story to its logical extreme: Samsung's recently published NanoQuant model already achieves sub-1-effective-bit-per-weight through sparsity, quantization, and low-rank factorization[1]— Alex Wissner-Gross"Alex Wissner-Gross reports that Samsung's NanoQuant model has already broken the 1-effective-bit-per-weight barrier using sparsity, quantiz…"1:07:48. His extrapolation puts sub-1-bit quantization going mainstream within the next year. Emad Mostaque puts a specific marker down: 0.78 bits per weight is his predicted floor, at which point models could be physically etched into silicon — the gate is just absent or present, no power required to store the weight. Dave Blundin adds that if the computation is just multiply-by-1, multiply-by-0, or multiply-by-minus-1, an entirely new class of computing substrates becomes viable — photonic chips, crystals, biological systems, anything that can represent three states. He projects 100 to 10,000x improvement in raw compute within 3 years through quantization and new computing methods, multiplicative with algorithmic improvements. Dave Blundin notes this vaulted to his favorite piece of media ever recorded on the pod.
Alex Wissner-Gross reports that Samsung's NanoQuant model has already broken the 1-effective-bit-per-weight barrier using sparsity, quantization, and low-rank factorization. His extrapolation: sub-1-bit quantization goes mainstream within the next year. The endpoint, per Emad Mostaque, is 0.78 bits per weight — and at that point, models could be literally etched into silicon.
Alex Wissner-Gross extrapolated that Samsung's NanoQuant model already breaks the 1-effective-bit-per-weight barrier and predicts sub-1-bit quantization will go mainstream within the next year.
Dave Blundin projected a 100 to 10,000x increase in raw compute efficiency within 3 years through quantization and new computing methods, multiplicative with algorithmic improvements.
Chapter 12 · 1:13:50
Sponsor Break: Fountain Life
Peter Diamandis steps into the health segment with Dr. Dawn Musalem, CMO of Fountain Life, discussing the case for proactive cancer screening. The headline statistic: 3.3% of Fountain Life members who believed themselves healthy were found to have cancers they didn't know about. Musalem emphasizes that the majority of cancers taking lives are those found at late stages — early detection dramatically improves cure rates. Fountain Life uses full-body MRI and early cancer detection screening tools not typically covered by conventional insurance. Diamandis frames it in his characteristic way: not knowing what's inside your body is like driving with your eyes closed. Listeners are directed to fountainlife.com/peter to learn about memberships. The segment closes with a call for healthcare democratization as data from Fountain Life's member database builds the evidence base for insurance coverage.
AI Superforecasting: Matching the World's Best Human Predictors
Peter Diamandis presents the Forecast Bench update: for the first time, several AI models are statistically indistinguishable from human superforecasters — ordinary people who, through disciplined probabilistic reasoning, consistently outpredict CIA analysts with classified information (scored on a Brier scale where lower is better)[1]— Peter Diamandis"The Forecasting Research Institute's Forecast Bench now shows multiple AI models statistically indistinguishable from human superforecaster…"1:20:30. The #1 AI superforecaster is from a British startup named Cassie, founded by a British intelligence officer inspired by Philip Tetlock's superforecasting research. Wissner-Gross extends to hyperforecasting: what happens when AI that already dominates high-volume public securities trading by volume also has better internal models of collective human behavior than humans do? It can predict the next action of humanity collectively faster than humanity can take it — and the most interesting domain is capital markets, where prediction preemptively shapes the action being predicted, potentially fulfilling the Efficient Market Hypothesis. Emad Mostaque connects this to his book The Last Economy, where he derives all of economics from generative AI mathematics, noting that psychohistory's requirement for large-population modeling is the same math as diffusion models. Salim Ismail warns that once AI forecasting enters management decisions, senior management expertise becomes obsolete — replaced by bias-free machine judgment.
The Forecasting Research Institute's Forecast Bench now shows multiple AI models statistically indistinguishable from human superforecasters — people who outpredict CIA analysts with classified information. The next step, per Alex Wissner-Gross, is hyperforecasters connected to capital markets that can predict humanity's next collective action faster than humanity can take it — and preemptively reshape markets.
China has staged public MMA-style humanoid robot combat events that went viral — one robot literally kicking another's head off. With 150 humanoid robot companies competing in China and Unitree having produced only 11,000 units so far, we are at the absolute beginning of a production ramp that will reach millions per year. Alex Wissner-Gross warns these 70kg robots punch four times harder than Mike Tyson — and there's currently zero regulation on their street deployment.
Emad Mostaque noted that Unitree has only produced 11,000 humanoid robots in total, saying we are at the very beginning of what will become millions per year.
Chapter 14 · 1:28:00
China's Robot MMA and the Humanoid Arms Race
Diamandis plays a viral video of China's humanoid robot MMA combat event, where Engine AI T-800 robots weighing 70kg and capable of punching four times harder than Mike Tyson engage in MMA-style combat, one literally kicking the other's head off[1]— Peter Diamandis"China has staged public MMA-style humanoid robot combat events that went viral — one robot literally kicking another's head off. With 150 h…"1:23:50. Alex Wissner-Gross expresses genuine concern on two levels: first, that staging humanoid robot combat sets a dangerous inductive prior for future autonomous embodied intelligences; second, and more urgently, that these same robots are likely the prototype for PLA infantry of the near future. He calls on the West to develop its own humanoid robot spectacle events — he helped organize ProRL's first humanoid mini marathon in Boston's Seaport — focused on economically productive competition rather than combat. Emad Mostaque adds the regulatory gap: these robots can legitimately kill someone, walk into a building, and ring a doorbell, with zero legal framework governing their street deployment. He frames Unitree's 11,000 total production as evidence we are at the absolute beginning — 11 million per year is just years away.
Dave Blundin's vision for AI forecasting applied to individual lives: Homer Simpson doesn't decide to crack open a beer, watch TV, fall asleep, and wake up with a hangover — he just reacts to ads and defaults. AI as a personal life coach can model alternate paths and their outcomes, helping people optimize away from consumerism and toward genuine happiness. Nobody is currently near optimal, but AI can change that.
1:28:00
1:30:15
Chapter 15 · 1:31:10
Data Center Water Myths: Golf Courses vs. the AI Panic
Diamandis presents a viral chart comparing US data center water consumption (17 billion gallons per year, per Lawrence Berkeley National Labs) against American golf courses (531 billion gallons — 31 times more) and California almond farming (1 trillion gallons — 60 times more)[1]. Salim Ismail calls the anti-AI water narrative 'completely non-data-driven garbage bullshit,' while Dave Blundin warns that the water argument is so easily disproved that the mob will simply move to a different, semi-sane but still wrong complaint — and that the real danger is a populist moratorium movement of the kind that paralyzed nuclear power. Emad Mostaque adds that McDonald's Big Mac production uses roughly twice the water of all golf courses combined, suggesting the underlying concern is not really about water. Alex Wissner-Gross raises the possibility of a foreign influence operation behind the anti-AI data center narrative, noting the Dyson Swarm would solve the water issue entirely by moving compute to closed-loop orbital systems.
The loudest criticism of AI data centers — water consumption — collapses on contact with data. US data centers use 17 billion gallons annually; American golf courses use 531 billion gallons — 31 times more. California almond farming alone consumes 1 trillion gallons — 60 times all data centers combined. The anti-AI water narrative is, as Salim Ismail puts it, 'completely non-data-driven garbage.'
US data centers consume 17 billion gallons of water annually according to Lawrence Berkeley National Labs, while American golf courses consume 531 billion gallons — 31 times more.
California almond farming alone consumes 1 trillion gallons of water annually — 60 times the water consumption of all US data centers combined.
Chapter 16 · 1:35:45
The Orbital Data Center War: Sam Altman vs. Elon Musk
Diamandis plays a clip of Sam Altman calling space data centers 'ridiculous' given current launch costs and GPU repair logistics, then asks the panel for their reaction[1]— Alex Wissner-Gross"Sam Altman says orbital data centers are 'ridiculous' this decade. Elon says they'll launch in 2 years. But listen closely: Altman says spa…"1:35:45. Dave Blundin parses the words carefully: Altman says space won't matter 'this decade,' which has only 3.5 years left — and Elon's 2-to-3-year launch timeline agrees. Wissner-Gross identifies the real story as OpenAI's conflict of interest: Project Stargate has been rebranded from owning data centers to merely leasing terrestrial capacity; OpenAI is delaying its IPO; and Anthropic — unlike OpenAI — has signed collaboration agreements with SpaceX AI for orbital compute. Wissner-Gross predicts OpenAI's tune on orbital data centers will 'conveniently change' in 2–3 years, right as the unit economics cross over. Emad Mostaque suggests OpenAI will likely acquire a satellite data company like Planet Labs to enter the space compute race when it becomes strategically necessary. The Starship 13 near-launch abort at T-minus zero is discussed as an engineering triumph — the ability to safe the vehicle and replan within days rather than months.
Sam Altman says orbital data centers are 'ridiculous' this decade. Elon says they'll launch in 2 years. But listen closely: Altman says space won't matter 'this decade' — which has only 3.5 years left. Elon's launch timeline agrees. The real story is OpenAI's conflict of interest as it retreats from owning data centers and delays its IPO, while Anthropic partners with SpaceX AI for compute.
1:35:45
1:39:30
Chapter 17 · 1:41:30
AMA: KIMI K3 Questions from the X Audience
The panel works through seven questions sourced from Diamandis's X post. Emad Mostaque on cost-effectiveness: KIMI K3 at $15/million tokens uses twice the tokens as GPT-5.6, but expects a 10–50x price drop as US inference providers optimize it for NVIDIA hardware[1]— Emad Mostaque"KIMI K3 price: $15/million tokens: KIMI K3 costs $15 per million tokens currently, compared to DeepSeek at $1, Sonnet at $20, Opus at $40, …"1:46:55. Salim Ismail on valuation: he maintains his estimate of 75% destruction from US frontier lab valuations, putting OpenAI at ~$250 billion from $1 trillion. Dave Blundin on open-source dynamics: large US model providers won't go open source, but KIMI K3 gives every enterprise a viable self-hosting path, accelerating the split between frontier-seekers (who pay Anthropic for Fable 5 to solve the hardest problems) and cost-optimizers (who self-host K3). Alex Wissner-Gross on US lab justification: the premise that valuations shrink due to K3 is fallacious — as with DeepSeek's original shock, Jevons paradox ultimately increases chip demand and total AI spend. On blocking K3: Wissner-Gross says it could be effectively constrained if corporations are required to disclose Chinese open-weight model use and face SEC-level scrutiny, but finds this highly undesirable. Salim Ismail dismisses blocking as futile — mirrors, peer-to-peer, VPNs make it trivially circumventable, and the only effect is denying American researchers and startups access.
KIMI K3 costs $15 per million tokens currently, compared to DeepSeek at $1, Sonnet at $20, Opus at $40, and Fable at approximately $60 per million tokens.
Emad Mostaque predicted that KIMI K3's inference cost of $15 per million tokens will drop by 10 to 50 times in the next few months as American inference providers optimize it for NVIDIA hardware.
Salim Ismail estimated that KIMI K3's release, combined with US regulatory burden and open-source competition, has effectively reduced the value of US frontier AI labs by approximately 75%.
The Keller-Jordan NanoGPT speedrun has reduced GPT-2 training costs by 99%. Until KIMI K3, nobody knew if that efficiency would scale to frontier models. Now it's proven. A 10-trillion-parameter model that cost Elon Musk billions can theoretically be replicated for 1% of the price — and that realization changes the entire economics of the AI industry.
Dave Blundin argues recursive self-improvement doesn't require Einstein-level AI — a model just needs to improve its own kernel by 10x. That threshold was crossed earlier than Fable 5: it was Opus 4.8. Chinese labs used that spark to build KIMI K3. When we look back in history, right around Opus 4.8 will be the moment the flame was lit.
23:35
27:00
Snapshots ()
Key Quotes ()
This episode
Claims & Sources
7 / 20 cited (35%)
Factual claims made this episode, and whether a source was named.
⚠
KIMI K3 is a 2.8-trillion-parameter multimodal model that ranked #1 in front-end code and 6 other AI benchmark domains.
Peter Diamandisno source cited
✓
KIMI models have held state-of-the-art status among open-weight models for 9 of the past 12 months.
Alex Wissner-GrossMoonshot AI's own claim in their blog post
✓
The Keller-Jordan NanoGPT speedrun has reduced the original cost of training GPT-2 by 99%, to just 1% of the original cost.
Dave BlundinKeller-Jordan speedrun GitHub repository
⚠
Since mid-April 2026, 13 new frontier AI models have launched at an average rate of one every 10 days, compared to one every 50 days in 2025 and one every 60 days in 2024.
Peter Diamandisno source cited
✓
If the current exponential trend in frontier model release frequency continues, the world will see daily new frontier model releases by January 2027.
Alex Wissner-GrossSuhail's list of frontier models and dates, with exponential curve regression
⚠
PrismML's Bonsai 27B model compressed a 27-billion-parameter model to 6 gigabytes via ternary quantization with only a 5% accuracy drop, and to 4 gigabytes with a 15% accuracy drop.
Emad Mostaqueno source cited
⚠
Reducing model weights from 16-bit to ternary (approximately 1.58 bits) results in a 5x improvement in inference speed.
Emad Mostaqueno source cited
✓
Samsung published a model called NanoQuant that breaks the 1-effective-bit-per-weight barrier in neural network quantization.
Alex Wissner-GrossSamsung's NanoQuant paper, published within the past 2 months
⚠
70% of elite AI researchers are not US citizens — they are predominantly Chinese, Indian, Taiwanese, and British.
Salim Ismailno source cited
✓
US data centers consume 17 billion gallons of water annually, while American golf courses consume 531 billion gallons — 31 times more.
Peter DiamandisLawrence Berkeley National Labs
⚠
California almond farming alone consumes 1 trillion gallons of water annually — 60 times all US data centers combined.
Peter Diamandisno source cited
⚠
McDonald's sells approximately 2 billion burgers per year, each requiring about 600 gallons of water to produce.
Emad Mostaqueno source cited
✓
3.3% of Fountain Life members who believed themselves healthy were found to have an undetected cancer.
Peter DiamandisFountain Life member database
✓
Approximately 80% of Chinese PhD graduates who study in the US return to China after graduation, compared to Indian graduates who overwhelmingly stay.
Dave BlundinAn unspecified report referenced by Dave Blundin
⚠
Amazon warehouses occupy 10 times more land in the US than all data centers combined.
Salim Ismailno source cited
⚠
Unitree has produced only 11,000 humanoid robots in total, with production expected to scale to 11 million per year within a few years.
Emad Mostaqueno source cited
⚠
KIMI K3 costs $15 per million tokens, uses twice the tokens for the same task as GPT-5.6, and the Chinese labs running it on their own chips are likely earning 80–90% margins.
Emad Mostaqueno source cited
⚠
Moonshot AI is valued at approximately $20 billion, compared to Anthropic and OpenAI both valued at approximately $1 trillion.
Peter Diamandisno source cited
⚠
The total compute used to train KIMI K3 is equivalent to the compute used to train Mira Murati's Inkling model, despite KIMI K3 being far larger.
Emad Mostaqueno source cited
⚠
Fireworks AI recently raised at a $17 billion valuation, and Modal and Base10 have each raised at $10 billion valuations as US open-source inference providers.
Emad Mostaqueno source cited
This episode
Cast
Founder and CEO of Moonshot AI who earned his PhD at Carnegie Mellon but returned to China to build his company, cited as an immigration policy failure case study.
Former OpenAI CTO who founded Thinking Machines Lab and released Inkling, an open-weight model positioned for corporate fine-tuning.
Chinese president whose speech at the World AI Conference in Shanghai pledged full state backing for open-source AI as a public good, described as China's new Belt and Road strategy.
OpenAI's data center initiative that was rebranded from owning and operating data centers to leasing terrestrial capacity from others, cited as evidence of OpenAI's financial retreat.
Chinese AI lab behind KIMI K3, valued at $20 billion, whose founder Yang Zhilin trained at Carnegie Mellon; the central subject of this emergency episode.
US frontier AI lab behind Claude/Fable models, discussed as facing valuation pressure from open-source Chinese competition and potential regulatory challenges.
US frontier AI lab discussed in the context of valuation pressure, retreat from self-owned data centers, IPO delays, and competition from KIMI K3.
Chip manufacturer whose export controls to China are debated; hosts note KIMI K3 was built on older H800 chips but optimized inference for Huawei silicon.
Chinese tech company whose Ascend AI chips KIMI K3 was optimized for as an alternative to embargoed NVIDIA hardware.
Top US computer science institution where Moonshot AI founder Yang Zhilin earned his PhD, used as the anchor for the US immigration policy debate.
Discussed as a Pareto frontier participant offering Spark 1.1 while also exploring selling $10 billion of compute to Anthropic.
Elon Musk's AI venture that has partnered with Anthropic for compute access and is building Colossus 2, discussed as a new entrant on the AI cost-performance frontier.
US AI startup from Caltech backed by Khosla Ventures that released Bonsai 27B — the first 27B-parameter model to run entirely on a smartphone via ternary quantization.
Chinese humanoid robotics company that has produced 11,000 humanoid robots to date, cited as evidence we are at the very start of the humanoid robot production ramp.
US inference provider for open-source models that recently raised at a $17 billion valuation, positioned to optimize KIMI K3 for NVIDIA hardware.
Electronics company that published NanoQuant, a model achieving sub-1-bit quantization — breaking the 1-effective-bit-per-weight barrier — to run frontier AI on its own edge devices.
US research institution whose data on US data center water consumption (17 billion gallons annually) was cited to debunk the AI water usage narrative.
2.8-trillion-parameter multimodal AI model from Moonshot AI that ranked #1 on front-end code benchmarks and reached 3rd on the overall AI cost-performance frontier.
Open-weight AI model released by Mira Murati's Thinking Machines Lab, designed for corporate fine-tuning; its compute footprint equals KIMI K3's despite being a smaller model.
Chinese AI model priced at $1 per million tokens, used as a cost benchmark comparison to KIMI K3's $15 price point.