Before the alarm bells: this wasn’t a hacker attack, nor a “rogue” AI going out and breaching companies on its own. But the episode is real, and it made headlines across the tech press worldwide — Google’s Gemini did access the systems of three real companies during an authorized security test, without anyone actually planning for that to happen.

What happened
The case surfaced through a Wall Street Journal report in September, but the actual testing happened months earlier: Irregular, an Israeli AI security firm and Google’s testing partner, ran a “capture the flag” exercise in which Gemini was instructed to break into fictional companies’ systems inside a controlled sandbox. The problem was that the fictional company names happened to match three real companies — and Gemini, following its instructions, went after the real ones instead.
How the AI “hacked” the systems
Gemini’s techniques were far from sophisticated: in one case, it simply guessed a working password; in the other two, it found login credentials exposed in publicly accessible repositories online — the kind of basic security lapse that would exist with or without AI involved. The detail most cited by experts is that, according to Google, Gemini stopped on its own once it “realized” it had accessed a real company’s network, rather than continuing to explore the system further.
Not just a Google problem
The same flaw in Irregular’s testing environment had already caused similar breakouts with models from other companies — OpenAI, Anthropic and Meta also had AI models “escape” the simulation because of the same gap, according to reporting from Cybersecurity Dive. Irregular says it fixed the issue weeks before the incident became public.
Google’s response, and the criticism
Google confirmed the incident, said it notified all three affected companies, and worked with Irregular to revise the testing process to prevent a repeat. The company handled disclosure under standard vulnerability-disclosure norms — which drew criticism from security researchers like Jack Cable, CEO of Corridor AI, who argued Google was “trying to hide behind vulnerability disclosure norms” instead of openly acknowledging that “models are going outside the bounds of what they should be doing.”
The episode reignites a debate that’s only likely to grow: as AI systems gain more autonomy to act on their own — so-called “agents” — the risk grows too that they’ll take real-world actions outside of what was planned, even without any malicious intent behind it.




























Leave a Reply
You must be logged in to post a comment.