AivexaNewsSearch
AI news for builders and product teamsChecked every hour

Google Gemini 4 Argon closes the gap with OpenAI and Anthropic but doesn't take a clear lead

Collected Sep 30, 2026

Google has unveiled Gemini 4 Argon, its first frontier model in more than seven months, following Gemini 3.1 Pro. According to The Decoder, Argon closes the gap with frontier models from OpenAI and Anthropic and beats some of them on key benchmarks, though it does not take a clear lead.

On Artificial Analysis's Intelligence Index at its "High" reasoning level, Argon scores 53 points, tying OpenAI's GPT-6 Astra (max) and Claude Fable 5.1, and one point ahead of GPT-6.1 Sol (max). Anthropic's Claude Opus 5.5 (58 points) and Claude Sonnet 5.5 (56) remain ahead. Argon is 23 points above Gemini 3.1 Pro Preview.

At the introductory price of $2 per million input tokens and $10 per million output tokens, one Intelligence Index task costs $1.99, which is 60 percent of GPT-6 Astra's $3.26 but 2.7 times more than GPT-6.1 Sol. Regular pricing later rises to $4 and $20. Cached input tokens cost 95 percent less. Argon averages 62,000 output tokens per task versus GPT-6 Astra's 27,000.

Google raised the output limit from 64,000 to one million tokens, which it calls an industry first. A new "Long Decode Continuation" feature in the Gemini API pauses long responses and resumes them through follow-up requests. The input context window stays at one million tokens; Argon accepts text, images, video, and audio but outputs text only.

Argon first goes to a group of "trusted cyber defenders" in the Fairwind program, who with Google's internal teams get the model without cyber guardrails. Google cites a "phased approach" and is taking part in a US government voluntary program giving agencies pre-release access. Developer, business, and consumer access follows, starting with paying API customers and Google AI Ultra subscribers, with no date given beyond "as soon as possible."

On AutomationBench-AA, Argon takes first place at 77.5 percent, six points ahead of Claude Sonnet 5.5 (max). On Terminal Bench 4 it hits 57 percent, behind Claude Sonnet 5.5 (64), Claude Opus 5.5 (60), and GPT-6 Astra (59). Its hallucination rate on AA-Omniscience is 15 percent, versus 51 percent for GPT-6 Astra (max) and 54 percent for GPT-6.1 Sol (max); accuracy reaches 50 percent. Argon tops Arena.ai's Text Arena with 1,525 points and leads the Vals Index at 68.9 percent.

Read at The Decoder

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

Gemini 4 Argon is Google's first frontier model in over seven months. It matches GPT-6 Astra in independent testing but can't keep up with Anthropic's Claude Opus 5.5. The per-token price is low, but Argon burns through more than twice as many tokens per task as Astra. Select testers get access first, with the API and paid tiers following later. The article Google Gemini 4 Argon closes the gap with OpenAI and Anthropic but doesn't take a clear lead appeared first on The Decoder .