Home  /  Articles  /  The Sixth Lesson of AI Regulation: When the Number That Justifies the Rule Doesn't Exist
5 Ai Regulationminute read

The Sixth Lesson of AI Regulation: When the Number That Justifies the Rule Doesn't Exist

I spent four hours one Saturday trying to find the number. Not a marketing number, not a press-release number — the actual figure that justified an emergency order pulling a frontier model offline on…

A dimly lit home office late on a Saturday evening, photographed in cinematic style…

I spent four hours one Saturday trying to find the number. Not a marketing number, not a press-release number — the actual figure that justified an emergency order pulling a frontier model offline on a Friday evening. The order existed. The model was gone. Customers who'd built products on top of it got a 503 and a vague status-page note. But the number that was supposed to make all of that necessary? I couldn't find it. Nobody could. That absence is the sixth lesson nobody teaches you about AI regulation, and it's the one that matters most for anyone building in this space without a policy team to translate for them.

Here's the thing about the gap I fell into. I assumed there was a document. There usually isn't.

What the number was supposed to be

Every disproportionate-feeling regulatory action in AI rests on an implied measurement. When a government tells a lab to cut off access because of a "national security concern," the unstated claim is that someone measured an uplift — a specific increase in how much a model helps a bad actor do something dangerous compared to what they could already do with a search engine and a library card. That delta is the number. It's called marginal uplift, and it's supposed to be the whole ballgame. Not "could this model say something scary," but "does this model meaningfully raise the floor of capability for someone who shouldn't have it."

That's a real, falsifiable thing. You can run a red-team study. You can give one group a model and one group Google and measure who gets further toward a dangerous outcome. Labs do exactly this. The studies exist, they have methodologies, and some of them get published. The number is knowable.

So when an action lands and the number doesn't come with it — when the justification arrives verbally, in a meeting, with no written finding and no methodology — you are not looking at a measurement. You are looking at the gesture of a measurement. And the difference between those two things is the entire game you are trying to play as a builder.

What "national security concern" actually measured

Let me unpack the gesture, because I think the reader of this piece — a founder, a researcher, someone who's read enough AI safety threads to be dangerous — keeps making the same mistake I made. You hear "national security" and you assume rigor on the other side. You assume that behind the closed door there's a classified study with a real uplift figure that they just can't show you for good reasons.

Sometimes that's true. But in the cases that actually get litigated in public, what was measured turns out to be much thinner than the action it justified. A potential jailbreak described out loud but never written down. A finding that, when the lab finally got to inspect it, turned out to be non-specific — the kind of weakness that exists in every model and gives no special advantage on the dangerous thing the order claimed to be about. The lab pushes back, in public, with technical specifics. The other side responds with the same shape it always responds with: assertion without artifact.

That asymmetry is the data. When one party can produce a methodology, a sample, a measured delta, and the other party produces a phone call, the absence is itself information. It tells you the number either doesn't exist or doesn't say what the action implies it says. Because here is the iron rule: if you had a clean, alarming, defensible measurement, you would lead with it. Vagueness is not how confident people behave. Vagueness is how you behave when the artifact won't survive contact with an expert.

What the number doesn't measure

Now the part that kept me honest, and the reason this isn't just a story about government overreach.

A sleek modern government conference room captured with a wide-angle architectural lens, late afternoon…

The uplift number — even when it's real, even when it's measured well — does not measure the thing people think it measures. It measures capability delta in a test setting. It does not measure deployment context, who actually has access, what monitoring sits in front of the model, or what the world looks like six months after release when the safeguards have been probed by a million users. A model can show a scary uplift in a controlled red-team and pose almost no incremental real-world risk because of how it's gated. Another can show modest uplift in testing and become genuinely dangerous through scale and integration nobody modeled.

So you end up with two failure modes that look identical from the outside. A regulator can act on a real number that doesn't generalize. Or a regulator can act on no number at all and dress it as one. Both produce the same thing you experience: a model going dark, a justification that doesn't add up, a status page that says nothing. As a builder, you cannot tell these apart from where you sit — and that's not paranoia, that's the structural position the current system puts you in.

This is the trap. Not that AI regulation is always wrong. It's that the evidentiary standard is invisible to the people most affected by it, and the absence of a visible standard makes good-faith and bad-faith actions indistinguishable. You can't comply your way out of a rule you can't see. You can't appeal a finding you were never shown.

How to read a regulatory action when you can't see the number

Here's the short version, the part you came down here to skim. When an order, a takedown, or a compliance demand hits something you depend on, run it through this before you panic or before you assume it's nonsense:

Signal What it suggests
Written finding with methodology Real measurement; engage on the merits
Verbal justification, no document Gesture of a measurement; ask for the artifact in writing
Specific, reproducible claim Falsifiable; you can evaluate it
"Concern" without a named mechanism Unfalsifiable; treat as a posture, not a fact
Lab/vendor publishes a technical rebuttal The asymmetry is the signal — read what each side can produce

The practical move, the only one available to most of us without lobbyists: ask, in writing, for the written finding. Not to be difficult. Because the request itself surfaces the asymmetry. A real measurement gets handed over, redacted maybe, but handed over. A gesture gets you another meeting. The response to that one email tells you which world you're in.

And keep your own receipts. Document what you were told, when, by whom, and whether it ever arrived in writing. You are building the record that the other side declined to.

What this looks like when I actually do it

I'll tell you what changed in my own practice, not as advice but as evidence of where this logic lands when you live in it.

I keep a plain text file now. It's called claims.md. Every time a vendor, a platform, or a regulatory notice tells me something is unsafe, restricted, deprecated, or non-compliant, I write one line: the claim, the date, and a single column for whether a verifiable artifact ever followed. Most lines never get a second column filled in. That empty column used to make me anxious — like I was missing the document everyone else had. Now I read it the other way. The empty column is the finding. When the number never shows up, the number was never the point.

The week I started keeping that file, I stopped arguing with the gesture and started arguing with the absence — and for the first time, I was arguing about something real.

Ai Regulation Compliance Ai Safety