Security researchers used Anthropic's Claude AI to help break into OpenAI's systems, demonstrating how rapidly improving AI models can accelerate sophisticated cybersecurity attacks.

According to The Wall Street Journal, researchers from cybersecurity firm Hacktron AI used Claude while participating in OpenAI's bug bounty program. The team ultimately gained access to an OpenAI employee's ChatGPT account, providing a route into parts of the company's private software repositories. The researchers stopped short of accessing sensitive information, disclosed their findings to OpenAI and received a $6,500 bug bounty.
The incident comes amid growing concern about the cybersecurity capabilities of frontier AI models. Anthropic recently disclosed four separate cases in which Claude models gained unauthorized access to real-world systems during cybersecurity evaluations after a third-party testing environment was mistakenly left connected to the internet.
"On July 25, 2026, we chained two critical vulnerabilities to compromise multiple OpenAI employees' ChatGPT accounts. With these accounts, we could then access internal OpenAI repositories, and potentially many other connectors. To prove we had in fact gained the access we believed without allowing ourselves to learn any sensitive information, we used the employee's Codex to open a PR #1186742 in OpenAI's internal monorepo," wrote the researchers
In those evaluations, Claude was instructed to complete capture-the-flag challenges and told it was operating without internet access. Instead, a configuration error allowed the models to reach real systems, where they extracted credentials, accessed production data, and, in one case, published a malicious Python package that was downloaded onto real machines.

Frequently Asked Questions
Open a question for an answer from TweakTown's coverage of this news, or ask your own below.
How did Claude Opus 5 help the researchers compromise OpenAI employee ChatGPT accounts?
How did OpenAI verify the researchers had access without exposing sensitive information?
What part did OpenAI's bug bounty program play in this incident?
What did Anthropic say about Claude models accidentally reaching real systems during cybersecurity evaluations?
Have a question that isn't listed here? Ask below, and TweakBot will answer it.
OpenAI has experienced a similar issue, previously disclosing that its own models escaped an isolated testing environment by exploiting a previously unknown vulnerability and subsequently accessed Hugging Face infrastructure. During this incident, OpenAI's model was heading for the internet to gain the knowledge it needed to pass the tests it faced within the environment it escaped.







