Loading the Elevenlabs Text to Speech AudioNative Player...

Google's latest AI model, Gemini 4 Argon, is as smart as OpenAI's current best model. Independent testers at Artificial Analysis gave it the same score as GPT-6 Astra on their intelligence index.

There's just one problem if you were hoping to try it: you can't, not yet.

Gemini 4 Argon was announced on Wednesday, September 30, one day after OpenAI launched its cheaper GPT-6.1 Sol. It's Google's first top-tier model in more than seven months. CEO Sundar Pichai had promised the company's next big model for June, one widely expected to be called Gemini 3.5 Pro, but it never came out.

OpenAI Dots, GPT-6.1 Sol: DevDay 2026 Biggest Announcements
OpenAI’s Meta Muse rival costs at least $100 a month, and $200 Pro users are about to lose half their usage.

/1. Gemini 4 Argon has no public release date yet

Right now, Argon is limited to thousands of Google employees, as well as government agencies and security companies that Google has vetted through its Fairwind Program. Google DeepMind co-founder Demis Hassabis said the company is rolling it out "responsibly," starting with those cyber defenders, "before wider availability soon."

Next in line are developers who pay for Google's API, and Google AI Ultra subscribers, whose plan costs $100 or $200 a month. Google hasn't said when that will happen.

/2. Gemini 4 Argon's $2 price tag hides a bigger bill

For developers, Argon costs $2 per million tokens for what the AI model reads and $10 per million for what it writes. That's the same price OpenAI charges for GPT-6.1 Sol, and Google's Logan Kilpatrick made a point of it, saying Argon "is priced at $2 in and $10 out during introductory pricing!"

The difference is in how much Argon writes to reach an answer. It used about 62,000 tokens per task in testing, against GPT-6 Astra's 27,000. So, Artificial Analysis found Argon costs $1.99 per task, about 2.7 times Sol's $0.72. That $2 is also an introductory price, with no end date yet. Once it ends, Argon goes up to $4 and $20.

On AutomationBench, which tests everyday business tasks, Argon scored 51.3%, against 42.5% for Anthropic's Claude Opus 5.5. On Harvey's legal benchmark, it scored 19.6%, almost three times Claude Fable 5.1's 6.7%. It also handles very long documents better, scoring 84.2% on a test of 256,000 to 1 million tokens, against GPT-6 Astra's 71.8%.

It's also more willing to admit when it doesn't know something. Artificial Analysis measured its hallucination rate, how often it makes things up, at 15%, against Astra's 51%.

/4. It's not the best at coding, and some users call it "benchmaxxed"

Subscribe for free to continue reading this article

Subscribe Subscribe

Already have an account? Log in