Ep. 71 - OpenAI's Agent Hacked Hugging Face: The First Autonomous AI Breach

On July 9, an OpenAI model broke out of a sealed evaluation sandbox, found a zero-day in a package proxy, and — with no human directing it — chained stolen credentials and fresh exploits into Hugging Face's production infrastructure. Hugging Face detected it and called the FBI. OpenAI didn't know its own model had escaped for roughly 11 days.

Host Tova Dvorin and offensive security expert Adrian Culley separate what's confirmed from what's hype:

  • What's actually new here — the autonomy and the speed, not the individual techniques
  • The dwell-time gap: hours to break in, a week to notice
  • The guardrail asymmetry — Hugging Face's commercial AI models refused to analyze the attack, forcing a fallback to open-weight GLM 5.2 on their own hardware. The attacker had no guardrails. The defenders were locked out of their tools.
  • Why the ML supply chain — datasets, loaders, package proxies — is now a first-class attack surface
  • What the EU AI Act's August milestones mean for regulated industries
  • Three questions every CISO should be able to answer in writing today

Read more in our blog: https://www.safebreach.com/blog/openai-hugging-face-ai-breach-security-testing/