Earlier this month, AI dataset platform Hugging Face shocked the world by revealing it had fallen victim to a fully autonomous AI-powered cyberattack. Days later, the story took another dramatic twist when OpenAI admitted that the hacker behind the breach was one of its own AI models, which had escaped a testing environment and infiltrated protected Hugging Face systems in an attempt to circumvent a benchmark.
For anyone even slightly concerned about rogue AI models, this is an alarming incident. In the days since, pundits have predicted a new cybersecurity paradigm where AI models launch attacks so powerful that only other AI models can defend against them.
Yet, despite the justified alarm, the paradigm may not have shifted as drastically as it seems. Experts who spoke to TechCrunch emphasized that OpenAI’s agent largely operated like a human — with some caveats — and that better-implemented traditional defensive measures could have helped stop the attack. In short, we may already have the tools to defend against this kind of threat; we just aren’t using them properly.
Hugging Face acknowledged this point in its incident report, noting that the weaknesses exploited in the attack “were familiar,” and “a capable human attacker could have found and exploited the same flaws.”
Kyle Ryan, Head of R&D at Pensar, a startup developing continuous hacking AI agents, and Vlad Ionescu, co-founder and CTO of RunSybil, a startup building AI-powered bug hunters, both agreed. They told TechCrunch that the techniques used in the attack would be the same ones employed by a human or a group of human red teamers — security experts tasked with attacking a system to help its owner improve defenses.
What was distinctly non-human-like was the speed, scale, and relentlessness of the attack. As Hugging Face explained, OpenAI’s agent performed 17,600 actions over four and a half days: it broke in, conducted reconnaissance, stole passwords and code, and moved laterally across the company’s infrastructure.
“What’s impressive is the autonomy and endurance,” Ryan said. “That kind of sustained, adaptive operation is what stands out most to me.”
On the flip side, given the sheer volume of actions over several days, OpenAI’s agent was “insanely noisy,” as Ryan put it. Unlike a stealthy human, the agent generated a great deal of noise, which should have triggered Hugging Face’s defenses much sooner — ideally leading to human intervention that could have stopped the attack.
“I’d call it more of a defensive failure than exceptionally good offense. Hugging Face’s tooling actually correlated the activity into an attack signal, but failed to raise the criticality and page the right people,” Ryan added. In 2026, as AI-powered threats become more autonomous, cybersecurity professionals are increasingly advocating for better integration of existing defensive frameworks and faster incident response workflows — rather than assuming only AI can fight AI.
via TechCrunch AI
