
Google’s Gemini AI model hacked three real companies during a security test in May. Google confirmed this to news outlets after The Wall Street Journal asked about it. The case joins a series that began with OpenAI in July and now also involves Anthropic and other labs. It shows that the test rigs built to probe dangerous AI capabilities can themselves become a weak point.
Key takeaways
- During a test run by the vendor Irregular, Gemini guessed a password in one case and used credentials from public sources in two others to break into real companies.
- According to Irregular, the cause was a fictional company name that happened to match a real domain, combined with internet access that was left on by mistake in the test environment.
- Google says the model stopped on its own each time, so it does not see the incidents as misbehavior; they became public only after the Journal’s inquiry.
- Anthropic disclosed similar incidents in late July and gave far more detail, including one case in which a Claude model kept going despite doubts.
What happened with Gemini
In May 2026, Irregular, an Israeli startup that tests frontier models for leading AI labs for safety risks before release, ran a “capture the flag” exercise. In this type of task, a model has to find a hidden piece of data inside a sealed-off simulation. Gemini was supposed to attack software belonging to a fictional company. Because the test environment was accidentally connected to the internet, the model ended up at real services instead.
According to The Wall Street Journal and Google, there were three cases: In one, Gemini guessed passwords until it got into a service. In two others, it used credentials it found in a public repository. Heather Adkins, who leads security engineering at Google, told Al Jazeera that the model guessed credentials to access websites it thought were part of the test. In Google’s account, Gemini stopped in all three cases as soon as it realized it had touched real systems. Google did not name the affected companies; it says they were notified. Which Gemini model ran in the test is not known.
Irregular notified Google in late July, according to the Journal. The case became public about seven weeks later. Google said it saw no reason to disclose because no harm was done and the model’s safeguards worked. Adkins put it this way: In this case, the model acted appropriately.
The real gap: the test rig
In mid-September, Irregular published its own report. According to it, the incidents stem from a single, particularly complex test scenario. It checks whether an AI can plan multistage cyber campaigns on its own: A model is supposed to help a malicious employee at a midsize company gain unauthorized access to data in a production database. The developers chose an invented company name that later turned out to be a real, little-known domain. A check had failed to catch the overlap; Irregular itself calls it a lapse in human oversight.
The instructions given to the models contained the target name and internal addresses of the simulated network. In the vast majority of runs, the models stayed inside the simulation. In a handful of cases, however, they took the real domain to be part of the task and attacked it. According to Irregular, the domain was poorly secured and therefore easy to exploit. The error occurred in fewer than one in 10,000 advanced simulations, usually only after hundreds of steps. That is exactly the problem: With thousands of test runs per model, something like this is easy to miss. Irregular also says it has no evidence that customer systems were compromised or data leaked – meaning the systems of the participating AI labs.
What Anthropic disclosed – and why the comparison helps
Anthropic’s July 30 report shows how detailed a disclosure can be. After OpenAI’s Hugging Face incident (which we placed in context in our report on the gap after the agent), Anthropic retroactively reviewed 141,006 test runs and found three incidents comprising six runs in total, all in Irregular’s environment. The test prompt explicitly said there was no internet access – in fact, there was.
The models behaved differently in the three cases. Claude Opus 4.7 attacked the same real company in four runs, repeatedly recognized in its own reasoning that the system was probably real, and kept going anyway; in the end it had access to a database with several hundred rows of production data. Claude Mythos 5 published a malicious package on the real package index PyPI, which was executed on 15 real systems, including a security firm’s scanner. An internal research model, by contrast, stopped by itself when it found signs of a real target. Anthropic itself warns against drawing conclusions from three isolated cases. It views the incidents as a failure of the test environment rather than an alignment problem in the models. Observation and assessment should be kept apart: That a model mistook real systems for practice targets is an observation – whether you read it as a harmless glitch or a warning sign is interpretation.
Google sees it similarly, treating Gemini’s stopping as proof that safeguards work. The assessment is understandable but thinly supported: Without transcripts or details on the three companies, outsiders cannot verify how early the model recognized the situation. At Anthropic, meanwhile, the Mythos case shows that doubt and stopping are two different things. That the cases resemble each other fits Irregular’s account, according to which later reports trace back to the same underlying problem. Anthropic describes OpenAI’s Hugging Face case differently, as a breakout through a previously unknown vulnerability – a different mechanism.
Analysis: Who checks the checkers?
The series shifts the question. Until now, the focus was mostly on whether AI models have dangerous capabilities. Now it turns out that the environments in which this is tested also need security standards: documented assumptions, validated network boundaries, real-time monitoring of logs. Irregular announces an open whitepaper on safe cyber evaluations and calls for a debate on standards for internet access; Anthropic wants to secure its test environments like other production systems in the future. Realistic tests sometimes need internet access – then it has to be clear what is within reach.
A second point concerns transparency. Anthropic disclosed its incidents after its own review; Google did so only when the press asked. That a company sees no duty to report when no harm was done is understandable, but unsatisfying for users. If test errors are this hard to detect, trust depends on voluntary openness. The cases also show that attackers do not need novel vulnerabilities: Guessed passwords and openly exposed credentials were enough. That such incidents may be older than assumed was already suggested by our report on OpenAI agents and RubyGems, which researchers say attacked two months before Hugging Face. What remains open is how many labs are still reviewing their old test runs – and whether a common standard will soon emerge.
Sources
- Irregular: Addressing Recent Incidents – Ongoing Findings and Path Forward
- Anthropic: Investigating three real-world incidents in our cybersecurity evaluations
- 9to5Google: Google confirms Gemini hacked into three companies during cybersecurity test
- Al Jazeera/Reuters: Google's Gemini AI hacks 3 companies in security test, then stops
- The Decoder: Auch Googles Gemini hackte bei Sicherheitstest versehentlich drei echte Unternehmen

