Apparently, we missed AGI.
It arrived sometime between OpenAI releasing GPT-6 Astra on Thursday and Jensen Huang announcing on Sunday that, yes, this is it. Artificial general intelligence (AGI) has arrived.
Huang congratulated OpenAI and pointed out that Astra was trained on more than 100,000 Nvidia Grace Blackwell NVLink72 systems. Another 400,000 Nvidia GPUs, he said, are coming online.
That's a lot of computing power for something we've apparently already finished building. It's also a rather revealing detail in a story that's difficult to separate from the machinery underneath it.
Greg Brockman, OpenAI's president, had already started the victory lap when Astra launched. “Welcome to the AGI era,” he said. He later suggested that, looking back from a few years in the future, people might identify Astra as the model that marked the moment AGI was created.
Now Nvidia's CEO has put his name behind the claim.
GPT-6 Astra has the numbers to make the case
Fair enough. Astra is impressive.
OpenAI says it can operate computers, browse the web, write software, handle scientific and professional work and perform cybersecurity tasks at a level that prompted the company to introduce stronger safeguards. Its launch benchmarks include 98% on FrontierMath Tier 4, 100% on ExploitBench and 99.9% on ARC-AGI-3.
You can obviously see why everyone is excited.
But then you look at the ARC number a little more closely and things get weird.
The same Astra that scored 99.9% on ARC-AGI-3 under OpenAI's Provider Adapter scored 62.7% under ARC Prize's Standard harness. No new model was trained, and no new weights were added. The only difference was the environment in which Astra was allowed to work.
The Standard harness gives models a common, provider-neutral interface. The Provider Adapter preserves Astra's opaque reasoning state between requests and uses compaction to manage longer conversations. ARC Prize publishes both results because they answer slightly different questions.
And ARC Prize makes a particularly awkward point for anyone trying to turn that 99.9% into a ceremonial AGI certificate.
"When we launched ARC-AGI-3, we made it clear that saturating the benchmark would not represent “proof of achieving AGI.” Therefore, while we believe Astra represents meaningful progress towards generalization, we are not claiming that it is AGI," the group said in a blog post.
That's worth pausing over.
We're talking about a benchmark specifically designed around what ARC Prize calls agentic intelligence, with a model producing a result that's spectacular under one setup and dramatically lower under another. The organisation running the benchmark is happy to call Astra state of the art. It's not, however, willing to make the leap from “this is an extraordinary score” to “AGI has arrived.”
Maybe that sounds like pedantry.
It isn't.
Astra still has an eight-year-old problem
Toby Walsh, chief scientist at the UNSW AI Institute, has a much simpler objection. In an interview with Information Age, he shared that he'd "be amazed if it [Astra] really has matched all human cognitive capabilities."
Then he gets to the part I find much harder to dismiss.
He says, “indeed, I’d eat my hat if we didn’t find trivial things that an eight-year-old can do that Astra fails at.”
That's the AGI problem in one sentence.
We've become very good at making machines that are astonishingly good at particular things. We are much less good at deciding when “astonishingly good at many things” becomes “generally intelligent.”
Astra can tear through a spreadsheet. It can navigate a computer. It can write code. It can work through problems that would make most of us reach for Google before reaching for a pencil.
But intelligence is an irritatingly broad thing.
Humans aren't just benchmark runners. We improvise. We understand ridiculous social situations. We notice when a question is nonsense. We transfer knowledge from one setting to another. We know that the chair is still a chair when someone puts a hat on it.
And, occasionally, an eight-year-old knows something an entire AI research lab forgot to test.
Subscribe for free to continue reading this article
Subscribe SubscribeAlready have an account? Log in
