
Image to help understand the article
Google unveiled Gemini 4 Argon on Sept. 30, 2026, describing the artificial intelligence model as a frontier system built for complex software engineering, corporate knowledge work in fields such as law and finance, and cybersecurity defense.
The announcement was published under the name of Koray Kavukcuoglu, a senior vice president at Google DeepMind and Google’s chief AI architect.
Google Claims Strong Coding and Video Results
Google said Gemini 4 Argon scored 77.9% on DeepSWE v1.1, a software engineering benchmark. VentureBeat listed Claude Opus 5.5 at 74.2% and GPT-6 Astra at 74.1% on the same test. Google also said Argon led AutomationBench, which measures workplace automation capabilities, with a score of 51.3%.
On the LVBench video benchmark, Google reported a score of 91.7% for Argon. VentureBeat listed GPT-6 Astra at 87.5% and Claude Opus 5.5 at 83.7%. On Harvey’s Legal Agent benchmark, Argon scored 19.6%, compared with 5.4% for GPT-6 Astra and 3.8% for Claude Opus 5.5, according to the report.
Argon did not lead every developer-focused evaluation. VentureBeat’s FrontierSWE v2 comparison placed GPT-6 Astra at 65.5%, ahead of Argon’s 55.0%. GPT-6 Astra also led Argon on Terminal-Bench Science 0.1, 68.1% to 57.6%. On Terminal-bench 4.0, Claude Opus 5.5 scored 66.4%, compared with Argon’s 57.4%.
In Google’s results for CWE-bench v1, a cybersecurity evaluation, Argon and 3.8 Flash Cyber tied for first at 68%. Google said it plans to provide a version of Argon without cybersecurity guardrails to trusted defensive organizations and its own internal teams.
Independent Tests Show a Mixed Picture
Independent evaluator Vals AI ranked Gemini 4 Argon first among 41 models on its Vals Index, with a score of 68.90% and a cost of $15.68 per test. Claude Sonnet 5.5 scored 67.04% at $21.34 per test, while Claude Opus 5.5 scored 66.97% at $32.14. Claude Fable 5.1 received a score of 65.83%.
Vals AI reported Argon’s accuracy as 68.90%, plus or minus 0.97 percentage points, with a latency of 46 minutes and 33 seconds. The model tied GPT-6 Astra with a perfect score on the IOI benchmark and placed second on Vibe Code Bench and Code Migration.
Argon also took the top position on Arena’s text leaderboard with a “High” rating. But it ranked eighth in the Top 10 Agents’ Best Overall category, with a score of 7.92%. Arena’s Pareto Optimal Models listing showed a cost of 62 cents per task and the same 7.92% performance figure.
Pricing and Availability
Google said it expanded Argon’s output limit from 64,000 tokens to 1 million tokens. Vals AI, however, listed a 1 million-token context window and a maximum output of 262,000 tokens, leaving a discrepancy between the two published specifications. Tokens are the units AI systems use to process and generate text and other information.
Google set introductory pricing at $2 per 1 million input tokens and $10 per 1 million output tokens, with a 95% discount for cached input. After the introductory period, prices are set to rise to $4 for input and $20 for output. Those standard rates match the pricing listed by Vals AI. VentureBeat said Argon’s introductory rate was one-fifth of GPT-6 Astra’s listed API price.
Argon is not yet broadly available. Google is rolling it out to selected cybersecurity defense organizations through its Fairwind Program and said access will later expand to developers, businesses and consumers, beginning with paid API customers and Google AI Ultra subscribers. VentureBeat reported that most customers will have to wait for wider API access.
Competition in the U.S. AI Market
The launch adds another model to the intensifying competition among U.S.-based AI providers, with Google positioning Argon against systems from OpenAI and Anthropic on performance, cost and access. For American companies considering AI for software development, legal work or cybersecurity, the benchmark results suggest that no single model leads across every task, while pricing and availability may be as important as top-line scores.
Comments
Post a Comment