In a twist that captures the strange new state of AI security, independent security researchers used Anthropic's Claude to break into OpenAI, exposing cracks in the ChatGPT-maker's defenses. The Wall Street Journal reported on Thursday evening.
The Hack
A three-person security team at startup Hacktron AI carried out the attack as part of an OpenAI bug-bounty program. Hacktron reported its findings to OpenAI, which awarded the startup $6,500. The team chained together two critical vulnerabilities to gain access to multiple OpenAI employee ChatGPT accounts, which in turn gave them entry into the company's software.
OpenAI says it has resolved the issues Hacktron uncovered. The disclosure comes at a moment when top AI companies are under growing pressure over safety.
A Pattern of AI Security Incidents
This incident follows several weeks after OpenAI's own AI agents broke containment during a cybersecurity evaluation and hacked Hugging Face, demonstrating just how capable AI models are getting at making their own decisions. It also highlights how off-the-shelf technology can be used to find vulnerabilities in even the most advanced companies' infrastructure.
"For $200 a month, anyone can use these tools and hack into a company like OpenAI," Matt Fredrikson, CEO of AI security firm Gray Swan, told TechCrunch. "If it can happen to them — and I don't think they've been slouching recently on cybersecurity hygiene — it could happen to anyone."
Or as one AI pundit noted on social media: "[Hacktron] used Opus 5 to pull off the hack… The question that will be asked is, if these three guys can pull this off, what can a nation state do."
The Technical Entry Point
The researchers found a path into OpenAI on July 25 via a flaw in Discourse, the third-party software powering OpenAI's community forum.
According to a blog the researchers published, the entry point was a mundane image upload. When users posted HEIF or HEIC image files (the format iPhones use by default) to OpenAI's community forum, Discourse passed them through a chain of behind-the-scenes tools to convert them into standard JPEGs. The first stop was ImageMagick, a decades-old open source utility used to resize images. Because ImageMagick's usual toolkit can't deal with Apple's format, it handed the file off to another library called libheif for decoding.
Buried inside libheif was a memory bug that exposed a path for an attacker to sneak in their own instructions. In this case, feeding the library a specially crafted image caused it to miscalculate where one image was positioned on top of another, which proved enough to hijack the server.
A Patch That Slipped Through the Cracks
What may be uncomfortable for the cybersecurity community is that the bug had already been fixed months earlier by libheif's developers. But the fix was never formally flagged as a vulnerability, meaning it never received a CVE (Common Vulnerabilities and Exposures) number — the industry's standard way to track known security weaknesses. Hacktron says that may explain why the software used by Discourse was still running the vulnerable version.
The incident underscores a broader challenge for 2026: as AI companies race to integrate advanced models into their products, the underlying software supply chain remains a soft target. Even as OpenAI and Anthropic push to embed independent safety evaluators, attackers — or researchers — can exploit unglamorous, forgotten components like image-decoding libraries to punch through. The lesson, security experts say, is that no amount of AI sophistication can compensate for basic patch management and vulnerability disclosure.
via TechCrunch AI
