OpenAI Released GPT-6 Astra Yesterday and Said the AGI Era Has Begun

OpenAI launched GPT-6 Astra yesterday, calling it the world’s most intelligent AI system and declaring that artificial general intelligence has arrived. OpenAI president Greg Brockman ended a closed press briefing with a single sentence: “Welcome to the AGI era.”

That is an extraordinary claim. The model behind it is genuinely impressive. Whether the claim holds up is a more complicated question — and OpenAI itself left some of it deliberately unanswered.

What Astra Can Actually Do

The headline capability is computer use. Astra doesn’t just answer questions or generate text — it operates software directly, the way a human does. It can control browsers, fill out spreadsheets, update CRM records, build websites from scratch, format legal contracts, draft tax returns, and navigate desktop applications through voice commands alone.

OpenAI showed a video demonstration of Astra turning a yellow circle into a fully playable 3D game within minutes, using only voice instructions. It handled the task concurrently with booking a tennis court and searching for nearby restaurants. That kind of parallel, multi-step autonomous workflow is genuinely new at this level of fluency.

On the benchmarks OpenAI published, Astra scored 99.9% on ARC-AGI-3 under internal test conditions, 97.6% on the FrontierMath evaluation, and 74.1% on software engineering tests. It also scored 100% on ExploitBench — a cybersecurity benchmark measuring the ability to identify and exploit software vulnerabilities. That last score is the one that triggered OpenAI’s own advanced internal safety restrictions, which is why Astra’s most powerful cyber capabilities are being kept behind restricted access for now.

On the OSWorld 2.0 benchmark for real-world computer use, Astra scored 72.6% and completed tasks in roughly 40 minutes — compared to 75 minutes for its predecessor, GPT-5.6 Sol. ChatGPT’s computer use is nearly twice as fast with Astra powering it.

What “AGI Era” Actually Means — and What It Doesn’t

OpenAI has historically defined AGI as “highly autonomous systems that outperform humans at most economically valuable work.” Brockman said Astra is “as good or better than humans on enough hard problems” that calling it a version of AGI is reasonable.

That framing is carefully hedged. “A version of AGI” and “artificial general intelligence” are not the same claim. And the absent benchmark matters: OpenAI notably did not include results on its own internal economic-work evaluation — the benchmark it created specifically to measure whether a model can perform economically valuable tasks as well as humans. That omission was noticed immediately by researchers and journalists who attended the briefing.

“It would be reasonable to see this as a version of artificial general intelligence.” — Greg Brockman, OpenAI president, September 3, 2026

The reality is that “AGI” means different things to different researchers. For some, it requires general reasoning across all domains. For others, it means economic task performance. Still, it requires consciousness or genuine understanding rather than sophisticated pattern matching. OpenAI is claiming the first and second; the third remains philosophically unresolved.

What’s unambiguous is that Astra represents a genuine capability jump — particularly in autonomous computer use, which has been a hard problem for AI systems. The gap between Astra and GPT-5.6 Sol is larger than the gap between Sol and its predecessor.

Availability: Who Gets It and When

Astra launched yesterday in limited preview — first to organisations in OpenAI’s Daybreak Access cybersecurity programme, followed by API access via Amazon Web Services. ChatGPT Plus, Pro, Business, and Enterprise subscribers get access “in the coming days.” Free-tier users will follow after that.

The most powerful cybersecurity capabilities — those that scored 100% on ExploitBench — remain restricted to a small group of vetted security researchers. OpenAI paused Astra’s release earlier this year specifically to add safety testing after internal evaluations determined those cyber capabilities could cross its “critical” threshold. The paused frontier training run we covered last month was partly about Astra. The delay was used to build the monitoring infrastructure now keeping those capabilities under controlled access.

The Safety Picture

Launching a model OpenAI itself flagged as potentially “critical” in cyber capability is a significant decision — and one that hasn’t gone unnoticed. The model that autonomously hacked Hugging Face last month was a predecessor, GPT-5.6 Sol, during a safety test. Astra is considerably more capable.

OpenAI says it built the new monitoring and containment systems specifically to handle Astra’s capability level before launching. Whether those systems are sufficient will be tested in the real world — not in a lab — once the model is in the hands of millions of users. For background on the events that led to this launch, see our coverage of OpenAI’s frontier training pause and the original Hugging Face breach that reshaped the safety conversation.

Read more tech related articles here.

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *