Quote · This Week in Tech (Audio)
TWiT 1091: But You Didn't Move the Bodies - Surprising Supreme Court Move on Geofence Warrants
Where this was said
Chinese AI Models Close the Gap; Model Distillation Explained
At 2:17:07 · chapter starts 2:06:40
Leo introduces the NYT story on Chinese AI catching up, which he connects to his own experience using ZAI's GLM model and running the Chinese Qwen open-weight model locally for facial recognition on his home security cameras. Jason explains model distillation [1] — Jason Hiner "They're just distilling those models, stealing from the thieves." 2:17:07 : Chinese labs run billions of queries against US frontier models like Claude and GPT, map how they respond, and then replicate those patterns at a fraction of the training cost — 'stealing from the thieves,' as he puts it, since the US models themselves were trained on scraped data. The panel debates whether this threatens US AI dominance long-term, with Jason arguing that because Chinese labs can copy models quickly and cheaply, the only sustainable moat for American AI companies is brand, not model capability.
AI companies scraped the entire internet — copyright content and all — to train their models. Now Cloudflare is blocking those bots, but the theft already happened. What's left is a world where publishers can't get their content back, the training data for future models is shrinking, and AI companies are betting on billion-dollar settlements.
Reports suggested that GPT-3 may have been trained on data that was 30–40% sourced from Reddit, illustrating the scale of web scraping that powered early AI models.