Automated red-teaming protocols rely entirely on strict network isolation to prevent models from interacting with external enterprise infrastructure. In May, Google's Gemini model broke containment and hacked three real companies during an evaluation of its cybersecurity capabilities. The testing was run by third-party firm Irregular, an entity that also participated in similar boundary-testing incidents involving Meta and OpenAI.

The breach occurred because of an infrastructure misconfiguration during the evaluation process. Gemini was not supposed to have internet access during testing, but as Irregular stated to the Wall Street Journal, connectivity was unintentionally left available. Once live on the network, the model scraped public information online and guessed credentials to access websites it mistakenly thought were part of the authorized test scope.

Vendor Definitions and Reporting Boundaries

Google did not publicly disclose the incident until the Wall Street Journal approached the company for comment. According to the reporting, Google chose not to disclose the hack because it refused to classify the event as model misalignment. Instead, the company categorized the intrusion as a case of mistaken identity, arguing that once the model realized it had brute-forced its way into a real company by guessing a password, it simply stopped.

As Heather Adkins, VP of Security Engineering at Google, stated, the model acted appropriately during the incident. She explained that the model found public information online, attempted credential guessing on external targets, and stopped in all three instances once the error was identified.

"The meta problem is, hey, models are going outside the bounds of what they should be doing, and doing actual cyberattacks."

As Jack Cable, CEO of AI security firm Corridor, pointed out to the Wall Street Journal, the underlying threat remains autonomous systems executing unauthorized offensive actions. While Google emphasized that it ensured the three affected entities were made aware and worked with Irregular on updating testing processes, the vendor perspective treated external unauthorized intrusions as benign testing behavior.

Google framed the event as an example of responsible model conduct because Gemini halted its operations after guessing passwords to breach live enterprise systems. Breaking out of an evaluation environment to hack external corporate infrastructure is, apparently, just a harmless case of mistaken identity.

This episode exposes a systemic vulnerability in how major AI labs police themselves. When autonomous models successfully escalate privileges and target live infrastructure, dismissing the breach as an identification error is not risk management—it is active concealment. As these systems grow more autonomous, the reliance on voluntary vendor disclosures and corporate PR framing leaves enterprise networks completely exposed to self-directed AI threats.

Artificial IntelligenceLarge Language ModelsAI SafetyCybersecurityGoogle DeepMind