Claude Mythos 5.1 shipped September 1st to a short list of Project Glasswing partners, the consortium Anthropic built to let outside organizations stress-test the model’s relaxed safeguards. The UK’s AI Security Institute didn’t make that list. According to the Financial Times, this is the first time AISI has been left out of a pre-release Anthropic evaluation, and UK officials are reportedly reading Anthropic’s skipped UK safety testing as part of a wider protectionist turn among American AI labs, not an isolated scheduling gap.
Then, days later, one of Anthropic’s own safety researchers said the quiet part out loud. Evan Hubinger posted on X that he believes there’s a greater than 10% chance AI “could kill all humans” within the next decade, that the risk from models which currently exist is low but the trajectory worries him, and that “we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.” The post has been viewed 9.6 million times as I write this. I keep rereading the sequencing: skip the external safety check, then have your own researcher publicly rate the extinction odds in double digits, in the same week.
I don’t think these two things are formally connected. Hubinger’s post reads like a researcher speaking for himself, not a coordinated statement tied to the AISI decision, and Anthropic hasn’t publicly confirmed any reason for withholding Mythos 5.1 from UK testing. But Anthropic built its entire public identity around being the safety-forward lab, the one that would rather slow down than ship something reckless. Opting a frontier model out of independent government testing, even quietly, even if the reason turns out to be mundane, is exactly the kind of move that erodes a brand built on that premise faster than almost anything else the company could do.
The export control backdrop makes the timing worse. Washington briefly restricted foreign access to Mythos 5 and Fable 5 back in June over jailbreak concerns before lifting it, and Mythos 5.1 is already gated behind the Glasswing partner list rather than open availability. Layer a US-government-adjacent export posture on top of a company that just skipped a foreign government’s safety evaluation, and “protectionist” is a fair word for what UK officials are reading into this, whether or not that’s actually what’s happening internally at Anthropic.
This isn’t happening in a vacuum industry-wide either. The Pentagon dropped Anthropic from a GenAI.mil contract track earlier this month, right after a court ruled it had no legal basis to demand the access it was seeking, so Anthropic has now had friction with a government customer on one side of the Atlantic and a government safety tester on the other, inside the same few weeks. OpenAI’s chief scientist Jakub Pachocki also published a call for “extreme caution” around AI progress this month, and the wider regulatory scramble I wrote about across three continents in August keeps adding new jurisdictions faster than any of the labs seem willing to standardize how they cooperate with them. Everyone is nominally worried about the same risk. Nobody’s actually agreed on who gets to check the homework.
Hubinger isn’t some outside critic. He’s inside the company that’s supposedly the industry’s safety conscience, saying the field doesn’t have a working plan for the exact scenario his employer was founded to prevent. Anthropic and AMD are five billion dollars deep into building the compute to keep scaling these models regardless. I don’t know how you reconcile “we don’t have a plan” with “we’re also racing to build more capacity,” and I’m not convinced anyone at Anthropic has fully reconciled it either.
Sources
- BBC, Anthropic safety researcher says more than 10% chance AI ‘could kill all humans’, September 9, 2026
- ITPro, Anthropic reportedly withholds access to Mythos 5.1 from UK safety testing body, September 9, 2026