GPT-6 Astra: OpenAI's First 'Critical Risk' AI Model

閱讀中文版 →

GPT-6 Astra: OpenAI's First 'Critical Risk' AI Model

One Model, Two Identities: OpenAI’s Biggest Training Run Ever — and Its Riskiest

OpenAI launched GPT-6 Astra in limited preview on September 3, 2026, following up with a stable release for paid users the very next day. OpenAI calls it “the world’s most intelligent and aligned model,” describing it as a “generational leap” for cybersecurity, professional work, software engineering, and science. The examples OpenAI cites range from serious to mundane: debugging code, proving math theorems, and navigating a browser through multi-step tasks, but also filing tax returns, building video-game scenes, ordering food, and searching for jobs — the point being that OpenAI isn’t just touting higher benchmark scores, but a model that can take over a whole string of everyday chores directly. VP of Research Aidan Clark said this was, “by far,” OpenAI’s largest training run yet, using more than 100,000 GPUs at the company’s Stargate site in Texas.

But that same day, Astra picked up another label: it’s the first model to reach the “Critical” level of cybersecurity capability under OpenAI’s own Preparedness Framework. OpenAI’s definition is blunt — with the right tools and access, Astra can find previously unknown security flaws and develop entirely new ways to exploit well-defended systems, without a person guiding it step by step.

What OpenAI Did to Keep It Contained

Once a model hits “Critical,” internal deployment rules escalate across the board: model checkpoints get encrypted, chain-of-thought reasoning is placed under universal monitoring, models can’t be used internally until they clear alignment evaluations, and any tool-using external inference has to go through misalignment monitoring.

OpenAI also published specific numbers. In an internal Codex simulation, Astra triggered only 34 severity-level-3 misalignment flags, versus 73 for its predecessor — a 53% drop. Against indirect prompt-injection attacks, defender success rates improved from 96.23% to 99.79%.

Brockman: This Might Be Where the AGI Era Begins

What actually sparked debate was a single line from OpenAI President Greg Brockman: he suggested that, looking back, people might come to see Astra as the model that marked the start of “the AGI era” — OpenAI defines AGI as “an automated system that can perform all economically valuable work as well as or better than humans.”

That’s not an idle comment. Astra runs on a new reasoning technique called “recurrent depth,” which OpenAI says makes it “faster and capable of performing more tasks than any prior iteration.” But the technique comes with a tradeoff: it makes Astra’s decision-making harder to inspect from the outside than earlier models.

Even OpenAI’s Own Chief Scientist Admits It’s Getting Harder

OpenAI Chief Scientist Jakub Pachocki made a rare public admission: preventing unintended harm is “increasingly difficult,” he said, and could become a bottleneck on future progress. That line sits inside the same announcement that calls Astra OpenAI’s “most aligned model yet” — an odd juxtaposition to say the least.

AI safety researchers have raised similar concerns, and they zero in on that same recurrent-depth architecture: if the reasoning technique itself makes the model’s decision-making harder for outsiders to inspect, then no matter how many monitoring numbers OpenAI publishes, independent researchers have a much harder time verifying what’s actually happening behind them. In other words, the question isn’t just whether Astra is smart enough — it’s that the bar for figuring out whether it’s safe enough just got higher too.

What Happens Next

For now, the full-capability version of Astra is limited to a small set of trusted testing organizations. ChatGPT Plus, Pro, Business, and Enterprise users — along with developers using the API, Azure, and AWS Bedrock — will get a more restricted version over the coming days, one that declines to answer certain cybersecurity and other sensitive-domain prompts.

On one hand, a declaration that AGI might already be here. On the other, an admission that even OpenAI isn’t fully sure it can keep the thing under control. This launch may be the most unguarded moment yet where both of those feelings show up in the same press release.

Related tools

RecommendedChatGPT

About the author

I’m Ryan, and I run RyanOps. My day job is software development and automation; here I track what changes in AI models, developer tools and software engineering, and write up hands-on notes from problems I have debugged and built myself.

About this site and the editorial process →