Hugging Face's CEO on Open Source AI, Model Routing, and the Future of Competition

Hugging Face's CEO on Open Source AI, Model Routing, and the Future of Competition

Hugging Face's CEO says Anthropic and OpenAI need MORE competition, not less — concentration of AI power in a few companies is far more dangerous than losing a few billion in revenue to Chinese distillation.

Jul 20, 2026 28:08 Difficulty: Intermediate Played

TL;DR

Hugging Face CEO Clément Delangue joins a16z's Theo Jaffee and Sofia Puccini to unpack the US government's unprecedented move to restrict GPT-4.6's release, arguing that open source AI is structurally safer than proprietary frontier models. Hugging Face crossing $100M ARR validates the open-source business model, while a Stanford study showing 70% of ChatGPT queries could run locally signals a coming shift toward model routing and a long tail of specialized models. The single biggest takeaway: concentration of AI power in a few frontier labs is far more dangerous than Chinese distillation of their models.

#open source AI safety #AI model routing #frontier model regulation #Hugging Face ARR #local AI inference #distillation attacks #Chinese AI competition #AI power concentration #European AI ecosystem #AI generational adoption #model specialization #open weights governance #open source AI #model routing #Hugging Face #AI regulation #frontier models #distillation #local AI #Anthropic #OpenAI #GPT-4.6 #Llama CPP #Mistral #Europe AI #AI competition #Chinese AI

Theo Jaffee and Sofia Puccini speak with Hugging Face CEO Clément Delangue about AI regulation, open source safety, model routing, and why competition is essential for the AI industry's future. Topics include GPT-5, government oversight of frontier models, Hugging Face surpassing $100M ARR, local AI, China's open-source ecosystem, and Europe's AI ambitions.

Chapter list
  • Before the intro even rolls, Clément Delangue is already making his most provocative argument of the episode. Distillation — using a more capable model's outputs to train a smaller one — is universal practice across every AI lab, he says, and accusing competitors of it while being the world's fastest-growing company rings hollow. The more urgent threat, in his telling, isn't a few billion dollars of lost revenue: it's a future where a small number of companies control all AI capability, infrastructure, and economic output. This cold open functions as a thesis statement for the entire conversation, setting up the competition and power-concentration themes that run through every subsequent topic.

  • The a16z narrator sets the table for the conversation, introducing Theo Jaffee and Sofia Puccini's guest — Hugging Face co-founder and CEO Clément Delangue — and previewing the episode's key arguments. The framing highlights three threads: whether open source AI is inherently safer than proprietary models, what Hugging Face's $100M ARR says about the economics of open source, and whether model routing will define the next phase of AI. The narrator also flags GPT-5, AI regulation, local models, and Europe's AI ecosystem as discussion points.

  • Theo Jaffee opens by welcoming Clément Delangue back for his second appearance, briefly establishing Hugging Face's identity as the open-source AI platform before pivoting immediately to the week's breaking news. The US government — for the first time in history — has moved to restrict a frontier AI lab's model release and sought to oversee customer access. Clément notes he was in DC during the week it unfolded, and mentions bumping into Tom Brown, Anthropic's co-founder, amid what appeared to be a shuffle in who was liaising with the White House — all before the news became public.

  • When asked for his take on the government's unprecedented intervention, Clément Delangue takes an unexpectedly nuanced position. He can't fully blame regulators, he says, because frontier labs spent years warning about their own models' dangers — citing GPT-2's withheld release as a vivid early example of doom marketing that primed governments to act. His actual concern isn't the restriction itself but its scope: he hopes it remains surgically targeted at a few large, generalist frontier models run by trillion-dollar companies with DC lobbying armies — not applied to startups, researchers, or smaller players who lack the resources to navigate regulatory compliance. He also calls for better scientific evaluation frameworks and more transparency, pointing to an AI safety agency he refers to as 'Casey' as a promising institution building those capabilities.

  • Sofia Puccini probes whether the government's restrictions on frontier models could eventually bleed into the open-source world. Clément's answer is a firm no — not because of political protection but because of structural differences. Open source models are generally less capable at the specific dangerous tasks (like cybersecurity attacks) that worry regulators, are more specialized rather than generalist, and benefit from community transparency that proprietary models lack. He argues that open source AI supports competition and access for small companies, startups, and academia — exactly the ecosystem that policymakers should want to protect, not restrict. The conversation sets up a sharper philosophical argument about what makes AI dangerous in the first place.

  • This chapter is the intellectual heart of Clément Delangue's argument about open source and safety. He invokes the principle that 'sunlight is the best disinfectant' and points out that no major dangerous technology in history — starting with the nuclear bomb — was developed in the open. Nuclear weapons required a closed, proprietary Manhattan Project with billions in resources, not a public GitHub repo. He extrapolates: AI cybersecurity capabilities capable of causing real harm are unlikely to emerge from the open-source community, not because it lacks talent but because it lacks the structural incentives and resources to pursue them. He also challenges the assumption that 'more powerful = more dangerous,' arguing that a model could be more powerful for medicine or climate science without becoming more dangerous for cyberattacks — especially if it wasn't trained on cybersecurity data. The 'jagged frontier' concept closes the chapter: different models lead on different tasks, and treating 'the frontier' as a single uniform bar misrepresents how AI capability actually works.

  • Asked whether open source's relative weakness at dangerous tasks implies a broader capability ceiling, Clément Delangue turns the frame around. Open source isn't trying to be the Ferrari — it's the engine inside every Ferrari. Proprietary APIs are built on open-source foundations. A key capability proprietary systems simply cannot replicate is offline intelligence: local models let users run AI in airplane mode, with no network connection, on a flight. That's not a limitation of open source — it's a structural advantage no API can match. The analogy efficiently captures the layered relationship between the two paradigms and sets up the subsequent discussion about Hugging Face's business model and local AI use cases.

  • Sofia Puccini pivots to business: Hugging Face has crossed $100M in ARR, and she asks what that milestone means for the economics of open source. Clément's answer is measured: it wasn't the company's primary goal, since the platform is built around maximizing the number of AI builders it empowers rather than optimizing revenue. But it is a proof point — a validation that the GitHub model (platform + open source community + premium infrastructure) works for AI too. He notes that the past few weeks have seen unusually strong growth in both open-source and local model interest, suggesting the momentum behind this business model is accelerating rather than plateauing.

  • Sofia Puccini asks what people are actually using local models for. Clément's answer maps neatly to three driving forces. First, cost: local models are free since they run on hardware the user already owns. Second, privacy: because data never leaves the device, local models are the only viable option for sensitive use cases — personal health conversations, confidential business data, or any scenario where sending information to an external API is unacceptable. Third, scale: running heavy agentic workloads 24/7 on a cloud API becomes unsustainably expensive; a Mac mini or laptop makes it tractable. He points to Llama CPP — Hugging Face's open-source local inference runtime — as the most widely used tool in this space, supporting models like Qwen and Gemma running on everyday consumer hardware.

  • Theo Jaffee introduces a scenario that has been circulating in policy circles: could the US government restrict Chinese open-source models through export controls or platform-level bans? Clément's response draws a sharp distinction between restricting an API and restricting open weights. APIs can be blocked at the network level; open weights, once released, are effectively unstoppable. Pull a model from Hugging Face and it reappears on ModelScope, Alibaba's model-hosting platform, or on BitTorrent within hours. More fundamentally, he argues, provenance doesn't matter much for open weights because the person running the model has full control and transparency — there's no ongoing data transmission, no ability for the original creator to bias, manipulate, or cut off access. The risk calculus is entirely different from using a Chinese API provider, where your data travels to their servers and your access can be revoked.

  • It's a pointed question: does the US government actually understand the risks of frontier AI, or is it just acting on fear of the unknown? Clément diplomatically declines to speak for policymakers but acknowledges that the field as a whole — not just the government — lacks robust, reliable benchmarks for evaluating AI risk. This is a significant gap, he notes. The body he's most hopeful about is what he calls 'Casey' — an AI safety evaluation agency he believes is building genuine scientific capacity to assess models rigorously. He's excited about them taking on more of this work and hopes their approach will inject more transparency and evidence into what is currently a very opaque regulatory process.

  • Sofia Puccini brings up a conversation she had with Andrew Trask from DeepMind, who pointed her to OpenFusion — a model available on OpenRouter that aggregates outputs from multiple AI systems into a single fused response, achieving better efficiency than any individual frontier model on certain benchmarks. Clément sees this as a natural response to a structural problem companies are waking up to: depending on a single model is risky. That model can be taken away, can be biased, can refuse certain requests, or — as happened with an earlier model he references as 'Fable-5' — can be intentionally designed to mislead users in specific domains. The OpenFusion example illustrates that multi-model approaches are not just a theoretical future state but are already emerging in the market.

  • This is the episode's most consequential forward-looking argument. Clément Delangue cites a Stanford study finding that 70% of ChatGPT queries could be answered just as well by smaller models running locally on a laptop — for free. Yet users don't do this because flat-rate subscriptions subsidize every query to the frontier model, removing any price signal that might encourage more efficient behavior. He uses a vivid analogy: it's like asking Einstein what the weather is today — in real life, Einstein would tell you to get lost, but because AI is subsidized, everything gets routed to the most capable system regardless of fit. Routing — automatically directing queries to the right model for the task — solves this. He points to Lovable as an early company already implementing routing under the hood. The long-term consequence, he argues, is a redistribution of revenue from frontier models (which currently capture the majority of AI spending) to a long tail of specialized, cheaper, more domain-specific models. This, he says, is what AI maturation looks like: moving from a simplistic single-model phase to a more sophisticated multi-model ecosystem.

  • Theo Jaffee presents the Anthropic-Alibaba distillation controversy with characteristic directness: one camp sees Chinese labs stealing frontier capabilities by creating accounts and systematically extracting model outputs; the other sees this as a standard practice using a paid service — and points out the irony of companies that 'stole the entire internet' for training data now claiming theft. Clément comes down squarely on the second side. Distillation is universal — he'd bet Anthropic has used OpenAI's outputs for specialized training tasks before. More importantly, it's an accelerant, not the root cause of quality: stopping distillation tomorrow wouldn't make Chinese AI labs incompetent, because their core research capabilities remain. And the competition argument falls flat when the companies complaining have grown to trillion-dollar valuations in record time. The real risk isn't too much competition — it's too little, as a handful of companies accumulate total dominance over AI's capabilities and economic rewards.

  • Two recent essays have framed Europe's AI moment starkly: 'Europe 2031' warns of irreversible slide into irrelevance if the continent doesn't move fast, while Anton Leisch's 'The Moonshot' lays out a credible path to Europe building a world-class frontier lab. Theo Jaffee asks Clément — who is himself French — whether the moonshot is realistic. His answer is cautiously optimistic. Europe already has frontier-capable labs like Mistral, genuine AI talent, and a structural energy advantage in France's abundant nuclear capacity. What it needs most isn't a single breakthrough company but an ecosystem: open research, open source, collaborative infrastructure. He draws the parallel to how OpenAI itself emerged — not from nowhere, but from a Google that had open-sourced the Transformer architecture, creating the foundation an entire generation of researchers built on. The lesson for Europe: invest in the ecosystem, and the companies will follow.

  • In the episode's final substantive exchange, Sofia Puccini asks what Clément actually sees younger people doing on Hugging Face's platform. The answer doubles as the episode's 'white pill' — an antidote to doom. Young users, he observes, burned through the consumer phase of AI remarkably quickly and are now firmly in builder mode: training their own models, assembling datasets, building products. Crucially, they're not just reproducing the most-hyped AI applications — they're working in domains that rarely make tech headlines: climate modeling, chemistry, biology, social systems. This generational shift toward building, rather than just consuming, is Clément Delangue's most hopeful signal about where AI is headed.

  • Theo Jaffee and Sofia Puccini close out with genuine thanks to Clément Delangue — 'very big fans' — before the a16z narrator delivers the standard episode outro, directing listeners to YouTube, Apple Podcasts, Spotify, the show's X account (@a16z), and its Substack at a16z.substack.com. The narrator closes with the standard disclaimer that the content is for educational purposes and does not constitute investment advice.

ARR (Annual Recurring Revenue)
A measure of the predictable, recurring revenue a company generates over a year, commonly used for subscription or platform businesses; Hugging Face crossing $100M ARR signals business model viability.
Distillation
A technique where a smaller or newer AI model is trained using outputs from a larger, more capable model to transfer knowledge and improve performance efficiently.
Open weights
AI model weights (the numerical parameters defining a model's behavior) that are publicly released, allowing anyone to download, run, or modify the model without accessing a central API.
Model routing
The practice of automatically directing an AI query to the most appropriate model — based on cost, capability, or domain — rather than sending every request to a single frontier model.
Frontier model
The most capable and advanced AI models available at any given time, typically large-scale models developed by well-funded labs like OpenAI, Anthropic, or Google DeepMind.
Llama CPP
An open-source runtime library maintained by Hugging Face that allows large language models to run efficiently on consumer hardware like laptops and phones without cloud infrastructure.
ModelScope
A Chinese AI model-sharing platform (run by Alibaba) that functions similarly to Hugging Face, hosting open-weight models and datasets for developers.
Agentic workloads
AI tasks where a model operates autonomously over extended periods, taking sequences of actions or decisions with minimal human input — often run 24/7 and computationally intensive.
AISI (AI Safety Institute) / CASEY
A government agency referenced in the episode (referred to as 'Casey') focused on evaluating AI risks and building scientific benchmarks for model safety — likely a reference to a US or UK AI safety body.
Jailbreak
A technique used to bypass the safety restrictions built into an AI model, getting it to produce outputs it was designed to refuse; widely acknowledged as a persistent challenge for proprietary safety measures.
Jagged frontier
The concept that AI capability is not uniformly advanced across all tasks — a model may be state-of-the-art in some domains while underperforming in others, making 'frontier' a multidimensional rather than singular measure.
OpenFusion
A model available on OpenRouter that combines outputs from multiple AI models into a single, fused response, discussed in the episode as an example of emerging multi-model architectures.
OpenRouter
A platform that provides unified API access to multiple AI models from different providers, enabling model routing and comparison without locking into a single vendor.
Provenance
The origin or source of something; in the context of AI, the country or organization that created and released a model — Clément Delangue argues provenance matters far less for open weights than for proprietary APIs.
Inference
The process of running a trained AI model to generate outputs in response to inputs; distinct from training. Running inference via an external API sends data to a third-party server.

Chapter 1 · 00:00

Cold Open: Distillation, Competition, and AI Power Concentration

Before the intro even rolls, Clément Delangue is already making his most provocative argument of the episode. Distillation — using a more capable model's outputs to train a smaller one — is universal practice across every AI lab, he says, and accusing competitors of it while being the world's fastest-growing company rings hollow. The more urgent threat, in his telling, isn't a few billion dollars of lost revenue: it's a future where a small number of companies control all AI capability, infrastructure, and economic output. This cold open functions as a thesis statement for the entire conversation, setting up the competition and power-concentration themes that run through every subsequent topic.

Chapter 3 · 01:35

Welcome Back and Setting the Scene in DC

Theo Jaffee opens by welcoming Clément Delangue back for his second appearance, briefly establishing Hugging Face's identity as the open-source AI platform before pivoting immediately to the week's breaking news. The US government — for the first time in history — has moved to restrict a frontier AI lab's model release and sought to oversee customer access. Clément notes he was in DC during the week it unfolded, and mentions bumping into Tom Brown, Anthropic's co-founder, amid what appeared to be a shuffle in who was liaising with the White House — all before the news became public.

Government
Government Restricts GPT-4.6: Unprecedented and Unsurprising

Hugging Face's CEO on Open Source AI, Model Routing, and th… · Jul 20, 2026 Government

The US government's move to restrict GPT-4.6's release is unprecedented — but Clément Delangue says frontier labs brought it on themselves with years of doom marketing. He hopes the restrictions stay contained to a few frontier generalist models and don't spill over to startups, academia, or smaller players.

Chapter 4 · 02:45

Government Restricts GPT-4.6: Unprecedented and Unsurprising

When asked for his take on the government's unprecedented intervention, Clément Delangue takes an unexpectedly nuanced position. He can't fully blame regulators, he says, because frontier labs spent years warning about their own models' dangers — citing GPT-2's withheld release as a vivid early example of doom marketing that primed governments to act. His actual concern isn't the restriction itself but its scope: he hopes it remains surgically targeted at a few large, generalist frontier models run by trillion-dollar companies with DC lobbying armies — not applied to startups, researchers, or smaller players who lack the resources to navigate regulatory compliance. He also calls for better scientific evaluation frameworks and more transparency, pointing to an AI safety agency he refers to as 'Casey' as a promising institution building those capabilities.

Chapter 5 · 04:52

Could the Government Come for Open Source?

Sofia Puccini probes whether the government's restrictions on frontier models could eventually bleed into the open-source world. Clément's answer is a firm no — not because of political protection but because of structural differences. Open source models are generally less capable at the specific dangerous tasks (like cybersecurity attacks) that worry regulators, are more specialized rather than generalist, and benefit from community transparency that proprietary models lack. He argues that open source AI supports competition and access for small companies, startups, and academia — exactly the ecosystem that policymakers should want to protect, not restrict. The conversation sets up a sharper philosophical argument about what makes AI dangerous in the first place.

Chapter 6 · 06:50

Open Source Safety: Sunlight, Nuclear Bombs, and Structural Differences

This chapter is the intellectual heart of Clément Delangue's argument about open source and safety. He invokes the principle that 'sunlight is the best disinfectant' and points out that no major dangerous technology in history — starting with the nuclear bomb — was developed in the open. Nuclear weapons required a closed, proprietary Manhattan Project with billions in resources, not a public GitHub repo. He extrapolates: AI cybersecurity capabilities capable of causing real harm are unlikely to emerge from the open-source community, not because it lacks talent but because it lacks the structural incentives and resources to pursue them. He also challenges the assumption that 'more powerful = more dangerous,' arguing that a model could be more powerful for medicine or climate science without becoming more dangerous for cyberattacks — especially if it wasn't trained on cybersecurity data. The 'jagged frontier' concept closes the chapter: different models lead on different tasks, and treating 'the frontier' as a single uniform bar misrepresents how AI capability actually works.

Chapter 7 · 09:40

Open Source as Infrastructure: The Engine Powering the Ferrari

Asked whether open source's relative weakness at dangerous tasks implies a broader capability ceiling, Clément Delangue turns the frame around. Open source isn't trying to be the Ferrari — it's the engine inside every Ferrari. Proprietary APIs are built on open-source foundations. A key capability proprietary systems simply cannot replicate is offline intelligence: local models let users run AI in airplane mode, with no network connection, on a flight. That's not a limitation of open source — it's a structural advantage no API can match. The analogy efficiently captures the layered relationship between the two paradigms and sets up the subsequent discussion about Hugging Face's business model and local AI use cases.

Chapter 8 · 10:30

Hugging Face Hits $100M ARR: Open Source Has a Business Model

Sofia Puccini pivots to business: Hugging Face has crossed $100M in ARR, and she asks what that milestone means for the economics of open source. Clément's answer is measured: it wasn't the company's primary goal, since the platform is built around maximizing the number of AI builders it empowers rather than optimizing revenue. But it is a proof point — a validation that the GitHub model (platform + open source community + premium infrastructure) works for AI too. He notes that the past few weeks have seen unusually strong growth in both open-source and local model interest, suggesting the momentum behind this business model is accelerating rather than plateauing.

Chapter 9 · 11:55

Local AI Use Cases: Privacy, Cost, and Always-On Intelligence

Sofia Puccini asks what people are actually using local models for. Clément's answer maps neatly to three driving forces. First, cost: local models are free since they run on hardware the user already owns. Second, privacy: because data never leaves the device, local models are the only viable option for sensitive use cases — personal health conversations, confidential business data, or any scenario where sending information to an external API is unacceptable. Third, scale: running heavy agentic workloads 24/7 on a cloud API becomes unsustainably expensive; a Mac mini or laptop makes it tractable. He points to Llama CPP — Hugging Face's open-source local inference runtime — as the most widely used tool in this space, supporting models like Qwen and Gemma running on everyday consumer hardware.

Chapter 10 · 13:55

Can the Government Restrict Chinese Open Source Models?

Theo Jaffee introduces a scenario that has been circulating in policy circles: could the US government restrict Chinese open-source models through export controls or platform-level bans? Clément's response draws a sharp distinction between restricting an API and restricting open weights. APIs can be blocked at the network level; open weights, once released, are effectively unstoppable. Pull a model from Hugging Face and it reappears on ModelScope, Alibaba's model-hosting platform, or on BitTorrent within hours. More fundamentally, he argues, provenance doesn't matter much for open weights because the person running the model has full control and transparency — there's no ongoing data transmission, no ability for the original creator to bias, manipulate, or cut off access. The risk calculus is entirely different from using a Chinese API provider, where your data travels to their servers and your access can be revoked.

Chapter 11 · 16:05

Is the Government Well-Informed About AI Risks?

It's a pointed question: does the US government actually understand the risks of frontier AI, or is it just acting on fear of the unknown? Clément diplomatically declines to speak for policymakers but acknowledges that the field as a whole — not just the government — lacks robust, reliable benchmarks for evaluating AI risk. This is a significant gap, he notes. The body he's most hopeful about is what he calls 'Casey' — an AI safety evaluation agency he believes is building genuine scientific capacity to assess models rigorously. He's excited about them taking on more of this work and hopes their approach will inject more transparency and evidence into what is currently a very opaque regulatory process.

Chapter 12 · 17:00

OpenFusion and the Multi-Model Future

Sofia Puccini brings up a conversation she had with Andrew Trask from DeepMind, who pointed her to OpenFusion — a model available on OpenRouter that aggregates outputs from multiple AI systems into a single fused response, achieving better efficiency than any individual frontier model on certain benchmarks. Clément sees this as a natural response to a structural problem companies are waking up to: depending on a single model is risky. That model can be taken away, can be biased, can refuse certain requests, or — as happened with an earlier model he references as 'Fable-5' — can be intentionally designed to mislead users in specific domains. The OpenFusion example illustrates that multi-model approaches are not just a theoretical future state but are already emerging in the market.

Technology
Model Routing: The End of the One-Model Era

Hugging Face's CEO on Open Source AI, Model Routing, and th… · Jul 20, 2026 Technology

Routing AI queries to the right specialized model — rather than defaulting every request to a frontier giant — could fundamentally redistribute revenue across the AI stack. A Stanford study found 70% of ChatGPT queries could be answered locally on a laptop. We're still in the first, simplistic phase of AI. The second phase is about routing, open source, and specialization.

Chapter 13 · 18:15

Model Routing: How AI Will Redistribute Value

This is the episode's most consequential forward-looking argument. Clément Delangue cites a Stanford study finding that 70% of ChatGPT queries could be answered just as well by smaller models running locally on a laptop — for free. Yet users don't do this because flat-rate subscriptions subsidize every query to the frontier model, removing any price signal that might encourage more efficient behavior. He uses a vivid analogy: it's like asking Einstein what the weather is today — in real life, Einstein would tell you to get lost, but because AI is subsidized, everything gets routed to the most capable system regardless of fit. Routing — automatically directing queries to the right model for the task — solves this. He points to Lovable as an early company already implementing routing under the hood. The long-term consequence, he argues, is a redistribution of revenue from frontier models (which currently capture the majority of AI spending) to a long tail of specialized, cheaper, more domain-specific models. This, he says, is what AI maturation looks like: moving from a simplistic single-model phase to a more sophisticated multi-model ecosystem.

Chapter 14 · 20:53

The Distillation Debate: Alibaba, Anthropic, and the Ethics of AI Competition

Theo Jaffee presents the Anthropic-Alibaba distillation controversy with characteristic directness: one camp sees Chinese labs stealing frontier capabilities by creating accounts and systematically extracting model outputs; the other sees this as a standard practice using a paid service — and points out the irony of companies that 'stole the entire internet' for training data now claiming theft. Clément comes down squarely on the second side. Distillation is universal — he'd bet Anthropic has used OpenAI's outputs for specialized training tasks before. More importantly, it's an accelerant, not the root cause of quality: stopping distillation tomorrow wouldn't make Chinese AI labs incompetent, because their core research capabilities remain. And the competition argument falls flat when the companies complaining have grown to trillion-dollar valuations in record time. The real risk isn't too much competition — it's too little, as a handful of companies accumulate total dominance over AI's capabilities and economic rewards.

Chapter 15 · 23:35

Could Europe Build a Frontier AI Lab?

Two recent essays have framed Europe's AI moment starkly: 'Europe 2031' warns of irreversible slide into irrelevance if the continent doesn't move fast, while Anton Leisch's 'The Moonshot' lays out a credible path to Europe building a world-class frontier lab. Theo Jaffee asks Clément — who is himself French — whether the moonshot is realistic. His answer is cautiously optimistic. Europe already has frontier-capable labs like Mistral, genuine AI talent, and a structural energy advantage in France's abundant nuclear capacity. What it needs most isn't a single breakthrough company but an ecosystem: open research, open source, collaborative infrastructure. He draws the parallel to how OpenAI itself emerged — not from nowhere, but from a Google that had open-sourced the Transformer architecture, creating the foundation an entire generation of researchers built on. The lesson for Europe: invest in the ecosystem, and the companies will follow.

Chapter 16 · 26:05

Young People as AI Builders: The White Pill

In the episode's final substantive exchange, Sofia Puccini asks what Clément actually sees younger people doing on Hugging Face's platform. The answer doubles as the episode's 'white pill' — an antidote to doom. Young users, he observes, burned through the consumer phase of AI remarkably quickly and are now firmly in builder mode: training their own models, assembling datasets, building products. Crucially, they're not just reproducing the most-hyped AI applications — they're working in domains that rarely make tech headlines: climate modeling, chemistry, biology, social systems. This generational shift toward building, rather than just consuming, is Clément Delangue's most hopeful signal about where AI is headed.

No indexed bits in this chapter.

Show stoppers

Technology
Model Routing: The End of the One-Model Era

Hugging Face's CEO on Open Source AI, Model Routing, and th… · Jul 20, 2026 Technology

Routing AI queries to the right specialized model — rather than defaulting every request to a frontier giant — could fundamentally redistribute revenue across the AI stack. A Stanford study found 70% of ChatGPT queries could be answered locally on a laptop. We're still in the first, simplistic phase of AI. The second phase is about routing, open source, and specialization.

Snapshots ()

Key Quotes ()

This episode

Claims & Sources

1 / 12 cited (8%)

Factual claims made this episode, and whether a source was named.

The US government restricted GPT-4.6's release and sought to oversee which customers could access the model — an action without historical precedent for any frontier AI lab.

Theo Jaffee no source cited

GPT-2 was originally withheld from release by OpenAI on the grounds that it was too dangerous.

Clément Delangue no source cited

70% of queries sent to ChatGPT could be accurately answered by smaller models running locally on a laptop.

Clément Delangue Stanford study published end of last year

Hugging Face surpassed $100 million in annual recurring revenue.

Sofia Puccini no source cited

Llama CPP is the most widely used runtime for running AI models locally on consumer hardware.

Clément Delangue no source cited

Distillation — using a more capable model's outputs to train a smaller model — is a common practice used by all AI labs, including likely Anthropic.

Clément Delangue no source cited

Removing an open-weight model from Hugging Face does not eliminate access to it — the model will still be available on ModelScope or torrent platforms.

Clément Delangue no source cited

The nuclear bomb was developed in a closed-source, proprietary program with billions of dollars in resources — not in open source.

Clément Delangue no source cited

Anthropic and OpenAI have become trillion-dollar companies and are among the fastest-growing companies in the world.

Clément Delangue no source cited

France has abundant nuclear energy that could be leveraged to power large-scale AI compute infrastructure.

Clément Delangue no source cited

The Transformer architecture that underpins modern AI was developed at Google, which then open-sourced it.

Clément Delangue no source cited

Jailbreaks undermine post-hoc safety guardrails on AI models, making such measures ineffective.

Clément Delangue no source cited

This episode

Cast

Stats

Episode stats

Insight Overview

insights
chapters

Insight distribution

Sub-Categories

Speaker breakdown

Talk Time

Connect

Parsed