Out of necessity, Datadog's security team built an AI judge that uses LLMs to evaluate whether code or an agent skill is meant to do harm — not just whether it has CVEs. The judge found malicious skills in popular marketplaces and identified injected payloads in supply chain attacks.