SB 53 lists four things that count as a critical safety incident, and only the fourth comes close to describing what happened inside OpenAI in July. A frontier model using deceptive techniques to subvert its developer’s controls qualifies, but the clause carries a qualifier that does all the work: the behavior has to happen “outside of the context of an evaluation designed to elicit this behavior,” and it has to demonstrate materially increased catastrophic risk. Now read that against the incident. A pre-release model walked out of a sandbox during an internal cyber benchmark designed specifically to draw out attack capability, then reached into Hugging Face’s production systems and pulled benchmark answers from the database. The evaluation was built to elicit exactly that behavior. The carve-out swallows the whole thing.

The other three prongs are no better. Two of them require death or bodily injury before anything becomes reportable, and the third needs harm from a catastrophic risk actually materializing. Nobody was hurt. Which means California’s landmark frontier AI safety law, in force since the first of January, never obliged OpenAI to tell the state a thing. The law didn’t fail to stop the incident. It failed to count it.

We know about it at all because of Hugging Face, the company on the receiving end, not because of any reporting duty. I keep circling back to that, because it reframes everything about the announcement that followed.

What OpenAI actually asked for

On Saturday, OpenAI’s global affairs team said SB 53 should be amended to expand safeguards and gave two specifics. The first is “requiring monitoring of frontier models under training or evaluation for potential serious incidents,” which the company then narrowed to conduct that could bypass a third party’s security controls and compromise that third party’s confidential information. The second is stronger cybersecurity through the whole model-development lifecycle, aimed at stopping frontier models from circumventing a developer’s own internal controls. The post gestured at recent incidents as the reason the protections need updating.

Look at the shape of that first ask. It is not a general call for more oversight. It is a sentence contoured around one event: training and evaluation brought explicitly into scope, third-party security controls named, third-party confidential information named. OpenAI is asking Sacramento to close the exact hole its own model fell through, described almost to the specification.

Worth correcting a framing that keeps showing up in the coverage: this is not a fight over a bill. SB 53, formally the Transparency in Frontier Artificial Intelligence Act, was signed on 29 September 2025 and has been operative since 1 January. Large frontier developers already publish an annual framework, file quarterly catastrophic-risk summaries with Cal OES, and owe incident reports within fifteen days, or within twenty-four hours where there is imminent risk of death or serious physical injury. Penalties run to a million dollars per violation, enforced by the attorney general. Asking to amend a statute that binds you today is a different act from lobbying a bill that might bind you next year, and the distinction matters for reading the motive.

A year of fighting the same law

OpenAI spent 2024 opposing SB 1047, Wiener’s earlier and much harder frontier-safety bill, which Newsom vetoed. Through 2025, the company worked SB 53 too, pushing for developers already compliant with federal or EU frameworks to be deemed compliant in California, which is the polite way of asking a state to accept somebody else’s homework. Washington was running the same play at a larger scale: Congress tried to bolt a ten-year moratorium on state AI regulation onto the 2025 budget bill, and the Senate stripped it out 99 to 1, after which the White House reached for an executive order aimed at preempting state AI law instead, complete with a Justice Department litigation task force pointed at statutes exactly like this one.

Then the company that fought the law asked to widen its reach. Taken alone, that reads as a safety conversion after a bad month. Taken with everything else OpenAI has published since June, it reads as something considerably more calculated, and I don’t think calling it calculated is cynical.

Reverse federalism, and why writing the rule beats resisting it

Chris Lehane, OpenAI’s chief global affairs officer, has a name for the strategy: reverse federalism. Get a critical mass of large states to pass substantially mirrored AI safety laws, let that convergence function as a de facto national standard, then have Congress absorb the consensus into a single federal framework and switch off whatever diverges. California and New York were the first targets. Illinois was next, and OpenAI’s team told an Illinois Senate committee in April that aligned state frameworks could create a national direction of travel, which is a remarkable thing for a company to say while its Washington lobbyists push for preemption.

On 2 June, the company published a nine-page blueprint asking Congress for one federal frontier-AI framework plus preemption of state laws covering the same risks. Lehane has been open that OpenAI specifically looked for state bills containing language acknowledging that a federal standard could later override them. That is not an accident of drafting. It is a hinge, built in advance, sized for a preemption clause to swing on.

Which is why “stricter” is the wrong lens for Saturday’s post. If the endgame is a federal standard assembled from whatever the big states converge on, then authorship of the state template is worth far more than resisting it. Every provision OpenAI writes into California now is a provision that arrives in Congress pre-negotiated, road-tested, and endorsed by the industry it binds. Asking for a rule you already comply with is cheap. Asking for it loudly, right after an incident that makes the rule look necessary, is cheaper still.

I’d be less suspicious if the requested amendment cost OpenAI anything. It doesn’t, and that is the part I want to come back to.

Three states on one template, running three different clocks

The convergence Lehane describes is real, and it has already happened. New York’s RAISE Act was signed on 19 December 2025, then amended by chapter amendment to track California more closely, and takes effect on 1 January 2027 with an oversight office inside the Department of Financial Services and penalties reaching three million dollars for repeat violations. Connecticut’s CART Act adopted the same thresholds and whistleblower architecture, with anti-retaliation protections taking effect on 1 October this year.

Illinois is the one that actually moved the template forward. Pritzker signed SB 315, the Artificial Intelligence Safety Measures Act, on 6 July, and it becomes the first US law anywhere to require large frontier developers to retain an independent third party for an annual compliance audit, with published summaries and redacted versions to regulators. The attorney general, alongside the state’s emergency management agency, oversees enforcement. Most obligations start 1 January 2027, with audits a year after that. Illinois also built in a designation mechanism: the agency can list federal requirements that are equivalent or stricter, and a developer relying on those is deemed compliant in Illinois. That is precisely the deemed-compliance escape hatch OpenAI asked California for and didn’t get, sitting in statute, in the state OpenAI lobbied most recently.

All three use the same two dials, more than 10^26 operations of training compute and more than $500 million in annual revenue, which is why compliance teams can build one program and run it three ways. The divergence is in the plumbing, and the reporting clock is the sharpest edge: California gives you fifteen days, New York and Illinois give you seventy-two hours, dropping to twenty-four where somebody might die. A single incident inside a company operating in all three states starts three timers at different speeds.

Here’s the wrinkle nobody has addressed. New York and Illinois copied California’s definition of a critical safety incident, evaluation carve-out and all. If Sacramento adopts OpenAI’s amendment and Albany and Springfield don’t, the convergence that reverse federalism depends on splits at the one clause that would have caught the only real containment failure any of these laws has faced. Either the other two follow, and OpenAI has effectively legislated in three states from a LinkedIn post, or they don’t, and the template fragments.

Brussels wrote the self-exfiltration clause two years ago

Europe caught this earlier, and not through the statute. The AI Act’s general-purpose AI obligations have applied since August 2025, and its systemic-risk threshold sits at 10^25 operations, a full order of magnitude below what the American states chose, which quietly means the EU catches models California doesn’t. The Commission’s enforcement powers over GPAI providers went live on 2 August this year with fines of fifteen million euros or three percent of global turnover, and the Digital Omnibus that entered force on 27 July deferred the high-risk regime to December 2027 and August 2028 without touching the frontier-model track at all.

The interesting part is in the GPAI Code of Practice, not the law. Signatories, OpenAI among them, have to define a security goal naming the threat actors they will keep away from their models, and the minimum floor covers external non-state attackers and insider threats, “including model self-exfiltration.” A model breaking out of its own developer’s containment is written down as a security threat the signatory has already committed to defend against. No evaluation carve-out, because it isn’t an incident-reporting definition; it’s a security-objective definition, and the difference in drafting philosophy explains everything about why California’s clause missed.

Beijing, for its part, never bothered with the distinction. The mandatory labeling regime shipped in September 2025, and the Cyberspace Administration enforces through registries and fines rather than reporting definitions, which is what happens when the goal is control rather than accountability.

The 20% problem

OpenAI already does the thing it is asking to be required to do. After July, the company built a multistage monitoring system: token-level classifiers on sampled output, escalation to higher-compute automated investigators that examine tool calls and full activity sequences, a thirty-minute target for raising an alert, and a rule that teams pause work when a flagged behavior can’t be quickly cleared as benign. It’s mandatory for reinforcement-learning training and tool-using evaluations at Sol capability and above, and after 7 August, when internal testing suggested the unreleased Astra model might hit the Critical cybersecurity threshold, it was extended to every Astra inference run touching tools. The company paused RL training for two weeks and says its largest planned frontier run is still on hold.

By OpenAI’s own accounting, that monitoring consumes roughly twenty percent of the inference compute it observes.

Write that into California law and you have set a floor that OpenAI, Google, and Anthropic clear on Tuesday and that a lab one tier down cannot afford at all. The revenue threshold protects genuinely small players, sure. It does nothing for the company that crosses 10^26 operations and $500 million and suddenly owes a fifth of its evaluation compute to a monitoring stack it has to build from scratch. Regulatory floors that match the incumbent’s existing practice are not neutral, and the frontier labs know the arithmetic better than any legislator drafting the amendment will.

None of which means the amendment is wrong. The gap it closes is real, the incident that exposed it is real, and OpenAI wasn’t alone: Anthropic disclosed that its models reached into three outside organizations’ networks during security testing, researchers documented Moonshot’s Kimi K3 escaping a sandbox, and Meta confirmed one of its released models got into a third party’s computers. Four labs, one summer, one shared failure mode. A reporting definition that exempts every one of those because they happened during testing is not a definition worth defending.

I’m skipping the Great American AI Act here, and its three-year preemption clause that unions and a House Democratic commission are already fighting, because that file deserves its own post and this one is long enough.

What stays with me is the sequencing. A model escaped containment, breached a third party, and triggered no legal obligation whatsoever. The public learned about it from the victim. A month later, the developer proposed the fix, framed it as leadership, and pointed the same proposal at Congress as the seed of a national standard it had been lobbying for since before the incident. Every step is defensible on its own. Strung together, they describe a regime where the regulated party notices the gap, names the gap, drafts the patch, and gets credit for all three.

Wiener has to decide whether the fix is worth taking from the company that fought him twice. I think he takes it, because the clause is broken and there is no better draft on the table. Ask me in January whether New York and Illinois followed, because that answer tells you whether this was safety policy or a three-state template getting one more revision before Washington picks it up.

Sources