Google's Gemini 4 Argon leads in knowledge work but limits coding access
Google released Gemini 4 Argon, a new flagship model that outperforms rivals in knowledge tasks and long-context reasoning, though coding results vary and availability remains restricted.
Google has officially unveiled Gemini 4 Argon, its latest flagship artificial intelligence model, positioning it as a superior alternative to current offerings from OpenAI and Anthropic. Announced on September 30, 2026, the model demonstrates significant advantages in knowledge work and long-context processing, although its performance in coding benchmarks is inconsistent. Access to the system is currently limited to a select group of testers and government partners as Google navigates a phased rollout strategy.
What happened
The release marks a shift in Google’s recent AI timeline, effectively replacing the previously announced Gemini 3.5 Pro plan with this more advanced iteration. While the company initially intended to launch a new Pro model in June 2026, it instead deployed a series of Flash models before arriving at Argon. The announcement follows a meeting between Google CEO Sundar Pichai and President Donald Trump, where Pichai joined leaders from Anthropic, Meta, Nvidia, OpenAI, and SpaceX in signing a voluntary commitment to self-police AI development. This agreement lacks formal enforcement mechanisms but signals increased industry coordination on safety standards.
Gemini 4 Argon achieves top or tied scores in 13 out of 18 benchmarks when compared against OpenAI’s GPT-6 Astra and Anthropic’s Claude Opus 5.5 and Fable 5.1. The model shows particular strength in tasks involving complex information synthesis and automation. For instance, it scores 51.3% on Zapier’s AutomationBench, nearly nine points higher than Claude Opus 5.5. In legal contexts, Argon achieves 19.6% on Harvey’s Legal Agent Benchmark, which is almost triple the score of Claude Fable 5.1, although this still indicates that it fully completes only about one in five tasks.
Despite these wins, the model trails competitors in specific technical areas. On the FrontierSWE v2 coding benchmark, Argon scores 55.0%, falling behind GPT-6 Astra by 10.5 points and Claude Opus 5.5 by nine points. It also places last on Terminal-Bench 4.0 among the compared models. These mixed results suggest that while Argon is optimized for general knowledge and agentic workflows, it may not yet be the definitive choice for all software engineering tasks. Google plans to gather feedback from early testers in its Fairwind Program before expanding access to paid API customers and AI Ultra subscribers.
How it works
A distinguishing feature of Gemini 4 Argon is its expanded output capacity. While one million input tokens has become standard for frontier models, Argon can generate up to one million output tokens in a single response. Previous Gemini models were limited to 64,000 output tokens. This increase allows the model to engage in deeper reasoning processes, generating hundreds of thousands of tokens to solve complex problems in a single trajectory without breaking them into smaller steps.
The model also incorporates specialized training for cybersecurity applications. Google trained Argon to autonomously identify, validate, and patch software vulnerabilities. For participants in the Fairwind Program and internal teams, the model is released without certain cyber guardrails to facilitate this work. Wiz, a security firm acquired by Google for $32 billion in March 2026, is already using Argon in its Scan for Good initiative. In this capacity, the model reportedly discovered a critical vulnerability in healthcare software used globally, which earlier frontier models had failed to detect.
Key details
- Gemini 4 Argon scores 68.9% on the Vals Index for knowledge work, surpassing GPT-6 Astra (63.1%) and Claude Opus 5.5 (67.0%).
- The model supports up to one million output tokens, a significant increase from the 64,000 token limit in previous versions.
- Introductory pricing is set at $2 per million input tokens and $10 per million output tokens, rising to $4 and $20 respectively later.
- Argon achieves 84.2% on the GraphWalks benchmark for inputs between 256K and 1M tokens, leading GPT-6 Astra by more than 12 points.
- In cybersecurity tests, Argon scores 85.8% on Google’s internal vulnerability discovery benchmark and 70.9% on Wiz’s penetration testing benchmark.
- Access is currently restricted to the Fairwind Program and U.S. government partners, with broader availability planned for later dates.
Why it matters
For developers and technical founders, the mixed coding performance of Gemini 4 Argon suggests that model selection must remain task-specific. While the high score on DeepSWE v1.1 (77.9%) indicates strong capabilities in certain software engineering workflows, the lower scores on FrontierSWE v2 and Terminal-Bench 4.0 imply that it may not replace specialized coding models for all use cases. Teams building agentic workflows for knowledge work, legal analysis, or automation will likely find immediate value, but those focused purely on code generation should continue to evaluate alternatives like GPT-6 Astra or Claude Opus 5.5.
The expansion of output tokens to one million changes how engineers can design prompts and agent loops. Instead of chaining multiple API calls to maintain context or break down large tasks, developers can potentially rely on a single, lengthy generation pass. This could simplify architecture for complex reasoning tasks but may increase latency and cost per request. The introductory pricing structure offers a temporary cost advantage, but the planned price increase to $20 per million output tokens aligns it with premium competitors, requiring careful cost modeling for production applications.
What you can do
- Evaluate your current workflows to identify tasks that benefit from long-context reasoning rather than rapid code generation.
- Apply for access to the Fairwind Program if your organization qualifies for early testing opportunities.
- Benchmark your existing agents against the reported scores on AutomationBench and Vals Index to estimate potential performance gains.
- Review your token usage patterns to prepare for the shift toward higher output limits and associated costs.
- Monitor the rollout schedule for paid API customers and AI Ultra subscribers to plan integration timelines.
- Consider the security implications of using models without cyber guardrails if you participate in vulnerability research programs.

