In a recent security breach, independent researchers from Hacktron AI used Anthropic’s Claude to penetrate OpenAI’s defenses, exposing critical vulnerabilities in the company’s software. The hack, carried out as part of an OpenAI bug-bounty program, highlights the rapidly evolving landscape of AI security, where even the most advanced models can be exploited with relative ease.
The security team managed to gain access to multiple OpenAI employee ChatGPT accounts, enabling them to explore the company’s infrastructure. This incident comes at a time when top AI companies are under increasing pressure to ensure the safety and reliability of their models. Matt Fredrikson, CEO of Gray Swan, warns that anyone can use these tools for hacking, potentially leading to state-sponsored cyber threats.
The entry point was a flaw in Discourse, the third-party software powering OpenAI’s community forum. Hacktron discovered a memory bug in libheif, a library used to handle image uploads, which allowed for an exploit. Despite the bug having been fixed months earlier, it went unreported, underscoring the importance of proper vulnerability tracking.
The success of the hack was attributed to the release of a newer version of the Claude model, Opus 5, which could generate a working exploit within hours. This incident also raises questions about the balance between model capabilities and security restrictions, as newer versions of AI models are increasingly catching up to the cutting edge of cyber capabilities.







