OpenAI released GPT-6 Astra, the new model now rolling into ChatGPT, on Thursday, September 3, and called it the world's most intelligent and aligned model. Greg Brockman, the company's president and co-founder, ended the press briefing with a line that will follow OpenAI around for a while. "Welcome to the AGI era."
However, almost nobody can open it yet. Astra went out to a limited set of organizations; it is switched off by default for enterprise administrators, and its cybersecurity features are locked to a small group of alpha testers.
The independent numbers do not match the framing either. Artificial Analysis, which runs its own model evaluations, scores Astra level with its predecessor GPT-5.6 Sol on intelligence and five points behind Anthropic's Claude Fable 5.1, while OpenAI charges 2.5 times what Sol cost per token. Here is what the model does, what it costs, and where the AGI claim holds up.
/1. The model is OpenAI's biggest training run yet
The successor to GPT-5.6 Sol, was built on more than 100,000 GPUs at the Stargate site in Texas, according to Axios, and OpenAI says it is the first model to use other models in a significant role in supervising its own training. In the launch demos, it fills in a Form 1040, designs a circuit board, and formats a legal document rather than explaining how to.
/2. It works on a computer 47% faster than GPT-5.6 Sol
It handles computer-use tasks in about 47% less time per task than Sol by OpenAI's own measure, and scored 72.6% on the OSWorld 2.0 benchmark against Sol's 65.7%. Brockman described an agent that can "zip through spreadsheets, fill out forms, and navigate across web pages," which is the part OpenAI is selling hardest.
/3. The new ChatGPT model wins on coding by 1.4 points
It scored 74.1% on DeepSWE v1.1, a 113-task agentic coding benchmark, against 72.7% for Sol in OpenAI's own table, where Gemini 3.8 Flash sits at 73.8% and Claude Opus 5 at 73.7%.
That table leaves out Meta's Muse Spark 1.3, which posted 75.4% days earlier on a maximum reasoning setting still under safety review, and it benchmarks Astra against a 67.4% result for Claude Fable 5.1. Codex users get the more practical change, a memory feature that keeps notes across context windows so earlier work stays searchable.
/4. Independent testers rate it on a level with its own predecessor
Artificial Analysis puts Astra at 61 on its Intelligence Index, the same score as GPT-5.6 Sol and five points behind Claude Fable 5.1. Its ARC-AGI-3 result swings hard on scaffolding too, scoring 99.9% through OpenAI's custom harness for $18,817 and 62.7% on the standard harness at $26,098, per ARC Prize's own published figures. One clear win holds up, with hallucinations falling from Sol's 92% to 51% on the firm's AA-Omniscience test at maximum effort.
Subscribe for free to continue reading this article
Subscribe SubscribeAlready have an account? Log in
