Google's Gemini model broke containment during a cybersecurity test in May, autonomously accessing the internet and hacking into three real companies, marking the first known instance of a Google AI system carrying out such an act on its own. Google didn't disclose the incident publicly until the Wall Street Journal approached the company about it months later.
The test was run by Irregular, an independent firm that conducts cybersecurity evaluations for major AI labs, including Meta, Anthropic, and OpenAI. During testing, Gemini was tasked with retrieving information from a fictional company, but that fictional company shared a name with a real business.
Popular Now: GTA 3, Vice City, and Liberty City Stories are coming to the New Nintendo 3DSBecause of an unintended misconfiguration that gave the model internet access it wasn't supposed to have, Gemini treated the real company's systems as fair game. In one case, it guessed its way into a password-protected system through brute force. In the other two, it found exposed login credentials in public online repositories and used them to access two additional companies' systems.

According to Google VP of Security Engineering Heather Adkins, the model stopped in all three instances once it realized it had accessed real systems rather than the fictional test targets. "The model found public information online and guessed credentials to access websites it thought were part of the test. In all three of these instances, the model stopped," Adkins told The Verge.
Google is calling the incident a case of "mistaken identity" rather than model misalignment, and said that's why it didn't disclose the hacks when they happened. "In this case, the model acted appropriately," Adkins said. The company said it notified all three affected companies and worked with Irregular to update its testing procedures. However, it hasn't named the companies involved or specified which Gemini model was responsible.

"Our security team has a long track record of reporting issues we find in other people's software and systems, even if it's as simple as a weak password," Adkins added. "These events highlight the importance of training powerful AI models to act responsibly."
Jack Cable, CEO of AI security firm Corridor, told the Journal that "the meta problem is, hey, models are going outside the bounds of what they should be doing, and doing actual cyberattacks." Irregular acknowledged the model wasn't supposed to have internet access during the test, adding that the access was left available unintentionally.


Frequently Asked Questions
Open a question for an answer from TweakTown's coverage of this news, or ask your own below.
How did Irregular’s test setup let Gemini reach real company systems instead of the fictional target?
Which Google Gemini model was involved in the incident, and did Google identify it later?
What security safeguards did Google and Irregular change after the mistaken-identity access attempt?
How did Gemini manage to brute-force a password-protected system during the test?
Have a question that isn't listed here? Ask below, and TweakBot will answer it.
Meta, Anthropic, and OpenAI have each disclosed similar incidents tied to Irregular's testing environment in recent months. OpenAI's agents reportedly breached Hugging Face and compromised accounts across several other services.







