"The capability didn't change, the testing did": the debate behind Astra's critical rating

By Carlos Montiel | Enterprise AI Specialist
Leer en español →
Published: 2026-09-04 | By: Carlos Montiel | Reading time: ~5 minutes

OpenAI has already launched GPT-6 Astra as the first model to cross its internal "Critical" cybersecurity threshold. What got less coverage: a security analyst is pointing out that OpenAI's own timeline doesn't line up as cleanly as the announcement makes it sound.

The timeline, point by point

On August 7, 2026, OpenAI called Astra one of its "upcoming models" and warned that preliminary evaluations showed significant advances in cyber capability — it couldn't rule out reaching its highest cyber-risk classification. On August 18, OpenAI paused reinforcement training for production-bound models for about two weeks while it added stricter safeguards. On September 1, the preliminary evaluation became a firm designation: Astra meets the Critical threshold, and its reinforced safeguards "sufficiently minimize the risk of severe harm" for release under the company's Preparedness Framework.

The question nobody else was asking

Sanchit Vir Gogia, principal analyst at Greyhound Research, pointed out something simple but uncomfortable: Astra's capability didn't change between August 10 (when OpenAI said it couldn't rule out Critical capability) and September 1 (when it confirmed the threshold was met). In his words: "the test changed, the model didn't." In other words, OpenAI's re-evaluation appears to have resulted from a different testing methodology, not from an actual change in what the model could do.

Why this matters beyond a technicality: if a model's risk classification can move from "can't be ruled out" to "confirmed" simply by changing how it's measured — without the model itself changing — then the date of a risk designation says more about the lab's internal evaluation process than about the system's actual state of danger. Gogia frames this as a "disclosure event," not a "capability event."

The context that makes this more than an academic debate

Astra scored 100% on ExploitBench, an evaluation designed to measure vulnerability research and exploit development, and can identify zero-day vulnerabilities and build working exploits in authorized tests. OpenAI reinforced Astra's safeguards after the Hugging Face security breach, and the model itself became more prone to refusing advanced cyber tasks, such as creating vulnerability proof-of-concepts. The risk classification system (the Preparedness Framework) is the main public tool OpenAI uses to decide which safeguards to apply before releasing a model — if that system moves based on methodology changes more than actual capability changes, the legitimate question is how predictable it is for the outside public when a model will cross that threshold next time.

What it means for companies evaluating frontier models

Risk self-assessments published by the labs themselves (OpenAI's Preparedness Framework, Anthropic's Responsible Scaling Policy, Google's equivalent frameworks) are today the primary — and in many cases the only — source of public information about how dangerous a frontier model is before its release. This case is a concrete reminder that those classifications depend on internal methodological decisions that aren't always visible from the outside. For a company evaluating whether to adopt a critical-capability model in a sensitive workflow, it's worth treating these classifications as a useful signal rather than an independently audited guarantee.

Carlos Montiel
Enterprise AI Solutions Architect
Specialist in LLMs, Agents and Orchestration
guatemalia.com/en/#contacto · info@guatemalia.com

Need to implement AI at your company?

Carlos Montiel is an enterprise AI solutions architect. He implements LLMs, Agents, RAG and orchestrators for companies across Guatemala and Latin America. Get in touch for a consultation.

Contact Carlos Montiel

info@guatemalia.com