AI labs grade new models against a tiered capability framework. The top rating, "Critical," is assigned when a system shows potential to conduct cyberattacks against sophisticated cyber defenses. OpenAI disclosed that it could not rule out a new model had reached that level and tightened controls on the model in response.

The constraint and what it measures

Sophisticated cyber defenses, by design, require expert human attackers to probe them. Defeating a hardened target demands domain-specific knowledge and the capacity to chain exploits across layered systems. These are skills that take years to develop and are in short supply; that scarcity is one reason sophisticated defenses hold against most threats most of the time. The "Critical" designation marks the point where an AI model may replicate enough of that expertise to operate as an effective offensive tool without a skilled human directing each step.

That distinction shifts the threat model. Below the threshold, an AI system assists a human attacker. At or above it, the model may itself become the scarce resource that previously only expert labor could supply.

Reading "could not rule out"

OpenAI's phrasing is exact and worth parsing carefully. "Could not rule out" is not a confirmed finding. In clinical terms, it describes a result that fails to exclude the worst outcome rather than one that positively identifies it. The evaluation returned insufficient resolution to exclude the possibility the "Critical" ceiling had been breached, not confirmation that it had.

That matters because the response to a confirmed finding and the response to an ambiguous one are not equivalent. OpenAI chose restriction rather than full withdrawal, a posture consistent with treating the ambiguity as a warning signal rather than a verdict. A model confirmed above the "Critical" threshold would, by the lab's own definition, be capable of launching cyberattacks against sophisticated defenses.

The AI security debate

OpenAI's disclosure enters a debate that has been intensifying around this question: how should AI labs respond when capability evaluations return ambiguous results near a dangerous threshold, and what does the public need to know when they do? The argument has no settled framework, which means each lab's handling of a near-threshold finding becomes its own precedent.

The lab's answer in this case was restriction without full withdrawal. OpenAI assessed a new model, found the evaluation could not confirm the "Critical" capability was absent, and tightened access controls rather than pulling the model from operation.

Related reading