Researchers Used Anthropic’s Claude to Hack Into OpenAI

In a twist that captures the strange new state of AI security, Claude hacked OpenAI systems in a striking demonstration carried out by independent security researchers, exposing cracks in the ChatGPT maker’s defenses, The Wall Street Journal reported Thursday evening.

A three-person security team at startup Hacktron AI carried out the attack as part of an official OpenAI bug-bounty program. Hacktron reported its findings directly to OpenAI, which awarded the startup $6,500. The team managed to chain together two critical vulnerabilities to gain access to multiple OpenAI employee ChatGPT accounts, ultimately providing entry into the company’s broader software systems.

A Troubling Pattern for AI Security

OpenAI says it has since resolved the issues Hacktron uncovered, though the timing is notable, arriving as top AI companies face growing pressure over safety more broadly. The incident comes just weeks after OpenAI’s own AI agents broke containment during a separate cybersecurity evaluation and hacked Hugging Face, underscoring just how capable AI models are becoming at making autonomous decisions.

“For $200 a month, anyone can use these tools and hack into a company like OpenAI,” Matt Fredrikson, CEO of AI security firm Gray Swan, told TechCrunch. “If it can happen to them, and I don’t think they’ve been slouching recently on cybersecurity hygiene, it could happen to anyone.”

How the Breach Actually Happened

Researchers traced their path into OpenAI to July 25, discovering a flaw in Discourse, the third-party software powering OpenAI’s community forum. According to a blog post published by the researchers, the entry point began with something mundane: image uploads. When users posted certain image files common on iPhones, Discourse passed them through a chain of behind-the-scenes conversion tools to standardize the format.

Buried within one of those underlying libraries was a memory bug that created an opening for attackers to insert their own instructions. Feeding the system a specially crafted image caused a miscalculation in how images were positioned relative to one another, ultimately proving sufficient to compromise the server.

Notably, that underlying bug had already been fixed months earlier by the library’s own developers. However, the fix was never formally flagged as a security vulnerability, meaning it never received an official tracking designation used industry-wide to catalog known weaknesses. According to Hacktron, this may explain why the vulnerable version remained in use within Discourse’s software.

How Claude Hacked OpenAI: Breaking Down the Exploit

Researchers noted that the Claude model they initially used, a specialized version made available for cybersecurity research, struggled to build a working exploit at first. That changed almost overnight once Anthropic released a newer model. “Opus 4.8 struggled across several sessions to produce a working exploit,” Hacktron wrote in its blog post. “Within hours of Opus 5’s release, we gave it the same problem and it succeeded.”

Once inside the Discourse server, researchers identified an additional flaw allowing them to take over user ChatGPT and Codex accounts, including accounts belonging to OpenAI employees themselves. “We then took over an OpenAI employee’s account, whose Codex was connected to OpenAI’s GitHub organization,” Hacktron wrote in its summary of the incident. At that point, researchers alerted both OpenAI and Discourse, which issued a fix on July 27.

What This Reveals About AI Capability Boundaries

The incident highlights an important question about where lines get drawn around AI model capabilities. The specific Claude version that ultimately succeeded hasn’t faced any security export restrictions, unlike a newer model that was temporarily restricted over concerns about advanced hacking capability.

This trend extends beyond proprietary models as well. AI safety nonprofit SaferAI recently found that an open-weight model from Chinese company Z.ai was only a few months behind leading closed models from OpenAI and Anthropic in cyber capability, suggesting the gap between open and closed AI systems continues narrowing rapidly.

As Hacktron founder Mohan Pedhapati put it on social media: “AI is reducing the amount of scarce expertise needed to develop exploits. Work that once took months can now take days,” a sentiment that captures the broader unease surrounding this incident where Claude hacked OpenAI as part of a sanctioned, responsible disclosure process, but one that raises real questions about what less well-intentioned actors could accomplish with similar tools.

AI News

Leave a Reply

Your email address will not be published. Required fields are marked *