Distillation is how a smaller, cheaper model learns from a larger, more expensive one — like a student absorbing a teacher's knowledge. China's GLM 5.2 almost certainly distilled from Western frontier models. The catch: everyone is doing this, including Google DeepMind, Grok, and Cursor.