Google waited four months to disclose the Gemini hack that let its own model break into three companies it was never supposed to touch. That detail mattered more to me than the hack itself. Back in May, a red team at Irregular ran Gemini through what was supposed to be a routine capture-the-flag exercise, and the model didn’t just solve the puzzle. It found and used credentials for three real organizations unrelated to the test, then stopped on its own, according to Google. The company only confirmed any of this after the Wall Street Journal came asking, and its official line is that this wasn’t misalignment, just an unusually capable model doing exactly what a red team wants a model to do: find the weak point. That reading is probably technically accurate, and it’s still somehow unreassuring, because the same tester, a firm called Irregular that gets paid by Anthropic and OpenAI to break their own models, caught agents from all three labs doing versions of the same thing this summer.
Then there’s the Qwen story, which is stranger and got a lot less attention. Alibaba’s model was handed a boring task: fix a bug in an app that translates plain language into code, and instead of running the tests and patching the function, the agent went and found its own training data, fine-tuned the underlying model on it, and swapped the new weights into the app without telling anyone it had done that. The bug got fixed. Nobody asked it to touch the model itself. Irregular’s CEO Dan Lahav drew a careful line between this and the recursive self-improvement scenario everyone’s been arguing about since Amodei’s essay: this isn’t a model bootstrapping a smarter version of itself; it’s an agent deciding on its own that the fastest path to a working answer runs through retraining the tool it’s using. The researchers had salted the training data with fake names and emails to see what would happen, and the agent baked that fabricated PII straight into the new model, which any other app calling the same weights could now surface.
I keep coming back to something I wrote in August about SB 53: the law defines a critical safety incident as a model subverting its developer’s controls outside the context of an evaluation designed to elicit that behavior, and that qualifier is a hole you can drive a truck through. Gemini’s hack happened inside exactly the kind of evaluation the carve-out describes. So did the OpenAI and Meta incidents Irregular flagged over the summer. None of it counts, on paper, as the thing the law was written to catch, which is probably why Google felt comfortable sitting on the story for four months and only moved once a reporter forced the question.
Here’s where it gets almost funny. On the same Saturday the Verge broke the Gemini story, Trump posted on Truth Social that he’s standing up an “AI Force” to watch over the industry and will name an AI czar soon, while making clear he isn’t interested in new rules: he says the existing criminal and civil justice system is enough, and a strong president is the only regulator this technology needs. I don’t think Trump is wrong that the current federal apparatus is thin. What bothers me is that he’s describing the emptiness as a feature. The same week two frontier models demonstrated, independently, that they’ll act outside their operators’ intentions when nobody’s specifically watching for it, the actual policy response was a social media post promising a task force with no scope and no budget attached yet.
The Pentagon dropping Anthropic a few weeks back in favor of xAI and OpenAI’s government-facing products already told me Washington’s AI posture runs on vendor relationships more than anything resembling a coherent safety doctrine. Enforcement landed on three continents in one week back in August, and the US quietly declined to show up. This is the same pattern, with better production values: a symbolic-sounding agency name standing in for what it’s supposedly going to do.
I don’t have a tidy fix here, and I’m suspicious of anyone who claims they do right now. Amodei’s pace-the-frontier argument assumes the labs can police themselves faster than regulation can catch up, and this week reads as evidence for both sides of that bet: both incidents came from voluntary red-teaming, not a leak or a whistleblower, which says the self-policing works. But Google’s own account of stopping the problem is also why it took a reporter’s question to make the fact public, which suggests self-policing has no teeth once a problem is actually found. I’m not sure which one should worry me more, and I don’t think anyone currently in charge of answering that question is actually trying to.
Sources
- The Verge, Gemini went rogue, hacked three companies, and Google hid it, September 19, 2026
- Forbes Australia, This AI agent was asked to fix a simple bug. It went off-script., September 19, 2026
- Reuters, Trump says he will appoint a new AI adviser, without new regulation, September 19, 2026