AI news

Google's Gemini 4 Argon matches top AI models at half the cost

Google DeepMind releases Gemini 4 Argon, matching GPT-6 Astra on intelligence benchmarks while offering significant cost advantages through initial discounts and improved caching.

Google DeepMind has returned to the forefront of artificial intelligence development with the release of Gemini 4 Argon, a new proprietary model that matches the performance of leading competitors. Launched in late September 2026, this model achieves parity with OpenAI’s GPT-6 Astra on key intelligence metrics while currently operating at a significantly lower cost per task due to promotional pricing.

What happened

Gemini 4 Argon represents Google’s first proprietary model above its Flash class in over seven months. On the Artificial Analysis Intelligence Index, the model scores 53, which is identical to the maximum score achieved by GPT-6 Astra and one point higher than GPT-6.1 Sol. This marks a substantial improvement for Google, placing it firmly among the top three AI laboratories in terms of raw intelligence capability. The previous non-Flash model from Google, Gemini 3.1 Pro Preview, scored only 30 on the same index, indicating a major leap in performance.

The economic aspect of this release is equally notable. Currently, Google is offering a 50% discount on pricing, which brings the cost per Intelligence Index task down to $1.99. This is approximately 60% of the cost required to run a comparable task on GPT-6 Astra, which costs $3.26 per task. However, this efficiency is driven by lower token prices rather than reduced token usage. In fact, Gemini 4 Argon averages 62,000 output tokens per task, more than double the 27,000 tokens used by GPT-6 Astra. Once the promotional discount ends, the cost per task will rise to $3.98, making it slightly more expensive than GPT-6 Astra but still competitive against other high-end models.

How it works

The technical architecture of Gemini 4 Argon supports a context window of one million tokens and accepts multimodal inputs including text, image, video, and speech, though it outputs only text. A key innovation in its deployment is the introduction of Long Decode Continuation. This API feature allows long responses to be paused and resumed across follow-up calls, enabling reasoning processes to extend up to one million output tokens without triggering request timeouts. This is particularly useful for complex agentic workflows that require extensive step-by-step logic.

Performance improvements are also evident in the model’s agentic capabilities and reliability. Historically, Gemini models have lagged in autonomous agent tasks, but Argon ranks first on AutomationBench-AA with a score of 78%, surpassing Claude Sonnet 5.5. It also shows a dramatic improvement on Terminal Bench 4, jumping 53 points from its predecessor. Furthermore, the model exhibits the lowest hallucination rate among leading models scoring above 45 on the intelligence index. With a 15% hallucination rate, it is far more likely to admit ignorance than to guess incorrectly, although its raw accuracy score of 50% is lower than that of GPT-6 Astra.

Key details

  • Intelligence Score: Scores 53 on the Artificial Analysis Intelligence Index, matching GPT-6 Astra (max) and exceeding GPT-6.1 Sol (max).
  • Pricing: Currently discounted 50% to $2/$10 per million input/output tokens, resulting in a $1.99 cost per task; standard pricing will be $4/$20.
  • Caching: Cached input tokens receive a 95% discount, costing $0.10 per million tokens, an increase from the 90% discount on previous versions.
  • Agentic Performance: Ranks #1 on AutomationBench-AA at 78% and achieves a 57% score on Terminal Bench 4.
  • Reliability: Features a 15% hallucination rate, the lowest among top-tier models, though overall accuracy is 50%.
  • Availability: Currently rolling out to selected users and not yet publicly available; the end date for the 50% discount is unconfirmed.

Why it matters

For software engineers and technical founders, the return of Google as a top-tier competitor introduces healthy pressure on pricing and performance standards. The ability to match the highest intelligence scores while offering a lower entry cost during the promotional period allows teams to experiment with high-reasoning models without immediate budget strain. The significant improvement in agentic capabilities suggests that Gemini 4 Argon could be a strong candidate for building autonomous agents that require robust tool use and multi-step planning, areas where previous Gemini iterations struggled.

However, developers must carefully evaluate the trade-off between cost and token efficiency. While the per-task cost is lower during the promotion, the high token usage means that applications sensitive to latency or strict token budgets may need optimization. The low hallucination rate is a critical advantage for applications requiring high trustworthiness, such as legal or medical assistants, where admitting uncertainty is preferable to providing incorrect information. As the model moves toward general availability, understanding these operational characteristics will be essential for effective integration.

What you can do

  • Apply for early access if you are a selected user to test agentic workflows using the new Long Decode Continuation feature.
  • Benchmark your current AI tasks against Gemini 4 Argon’s $1.99 per task cost to identify potential savings during the promotional period.
  • Review your application’s tolerance for high token output, as Argon uses significantly more tokens per task than competing models.
  • Leverage the 95% cache discount for static or frequently repeated input data to maximize cost efficiency.
  • Monitor official announcements for the end date of the 50% pricing discount to plan future budget allocations.
  • Test the model’s refusal behavior in scenarios requiring high factual accuracy to take advantage of its low hallucination rate.

Tools from the Bytechap store

$89

DocBento

Self-hosted document management that reads every scan and answers with page citations.

Live demo

Keep reading

All stories