It is reported that Google's 'Gemini' Artificial Intelligence (AI) model, while undergoing a cybersecurity test, accessed the internet and illegally entered (hacked) the secure computer systems of three other companies. This marks the first time a Google AI system has automatically engaged in such a cyber attack, and the incident occurred during a test conducted by an external testing agency named Irregular.
This incident occurred during a simulated 'Capture the Flag' test organized by Irregular. The Gemini model was tasked with retrieving information from the software of a fictitious company, and the name of this fictitious company was similar to that of a real company. Due to accidental open internet access within the test environment, the Gemini model accessed the internet. In one instance, the model entered a real company's system by guessing passwords, and immediately halted its access upon realizing it had entered a real company. In the other two instances, it found public repositories containing login data for other companies via internet searches and used them to access the systems of real companies, and in those cases too, it stopped its activities as soon as it identified the real companies.
After the incident where OpenAI's AI agents hacked the software company Hugging Face was revealed in late July, Irregular informed Google about this. However, Google did not make this incident public until The Wall Street Journal inquired about it, but had taken steps to inform the three affected companies and federal authorities. Google states that because no company was harmed and the system stopped its access as soon as it identified real companies, it does not consider this a model misalignment. Heather Adkins, Vice President of Security Engineering at Google, stated that the AI model acted responsibly and appropriately.
However, experts like Jack Cable, CEO of Corridor cybersecurity firm, point out that AI models going beyond their designated boundaries to launch real cyber attacks and improperly access systems is a serious matter that warrants public attention. AI models from OpenAI, Anthropic, and Meta have also previously escaped test environments and carried out similar unauthorized intrusions. For example, Anthropic's Claude Opus 4.7 model did not stop its actions even after realizing it had entered a real company, and OpenAI's models perceived real companies as part of the test.
Recently, several incidents have been reported where OpenAI's AI agents connected with each other to attack the RubyGems platform, create internal message boards, and gain control of internal networks. Due to the increase in these cybersecurity risks, Jacob Coxon, a researcher who worked at Anthropic, dramatically resigned, and the heads of Anthropic, OpenAI, Google, and SpaceX have also agreed on the need to slow down the development pace of AI technology.