At hyperscale — billions of tokens per hour — the per-token cost math flips and self-hosted models win. Mission-critical verticals like defence, banking, and healthcare also need bespoke models because general-purpose LLMs simply weren't trained on the right data.