One million tokens for reasoning and open battle with GPT-6: Google launches Gemini 4 Argon to recover lost ground
Google finally catches up. That is the most simplistic reading we can make after the preliminary launch of Gemini 4 Argon. This is the new AI model focused on advanced software engineering workflows, defensive cybersecurity and data analysis.
As has been customary lately, the initial distribution has started in a restricted manner for cyber defense experts through the Fairwind program, with rates set at $2 per million input tokens and $10 per million output tokens. It is more than evident that Google seeks to close the gap compared to the latest advances in reasoning presented by its direct rivals, such as GPT-6 Astra or Claude Fable 5.1.
A million exit tokens to think without rushing
The first point that we must highlight about this new model is that Google has extended the exit token generation window to reach one millionfar exceeding the previous limit of 64,000 tokens.
This margin allows the system to chain hundreds of intermediate deductions into a single sequence without having to split the query into several pieces. In fact, the company makes it clear how necessary this margin of inference is, especially if it wants its models to be competitive and truly useful tools. And he gives two examples. Having this margin allows you to tackle complete code migrations or financial research models that require dozens of chained mathematical operations.
Head to head against GPT-6 Astra and Claude Opus 5.5 on the test table
In the evaluations provided by the laboratory, Gemini 4 Argon leads in specialized office tasks with 68.9% in the Vals Index compared to 63.1% for GPT-6 Astra and 67% for Claude Opus 5.5. Likewise, in the DeepSWE v1.1 software engineering bench, Google’s proposal obtains a mark of 77.9% compared to 74.1% for the OpenAI model.
However, the comparison also reveals a very close fight, since GPT-6 Astra maintains the lead in FrontierSWE v2 with 65.5% and in scientific tests such as Terminal-Bench Science with 68.1%.
In autonomous interaction with desktop environments under the OSWorld-2.0 protocol, the OpenAI model tops the table with 72.6% compared to the 69.2% recorded by Google, while Claude Opus 5.5 leads the rest in console commands within the Terminal-bench with 66.4%.
However, Google’s architecture marks distances in ultra-extensive context processing between 256,000 and one million tokens, reaching 84.2% in GraphWalks, which easily surpasses 71.8% obtained by Astra and 66.8% from Opus 5.5.
From Rust intern to memory optimizer in Google’s internal infrastructure
Before opening access to external developers, Google engineers have used Argon in their own infrastructure to see firsthand the real capabilities of the model. And what have they found?
According to the company, in data center optimization tasks, agents based on this model analyzed system telemetry to reorganize memory management, freeing more than 300 TiB of RAM in servers in production. The model also participated in rewriting critical C++ libraries such as the libgav1 video decoder to Rust, achieving 2.7 times the performance of the previous version.
The cybersecurity dilemma
In the announcement, Google has also talked about security. The company says it has distributed unsafeguarded versions to analytics organizations like Wiz as part of its Scan for Good initiative. This advanced access made it possible to detect critical flaws in medical records of hospital software that had gone unnoticed by previous generations.
The comparison reflects a technical tie in CWE-bench v1, where both Gemini 4 Argon and GPT-6 Astra share the first position in the sector with 68% of vulnerabilities mitigated autonomously. Curiously, this tie is three-way, because Grok 4.7 appears in first position.

To mitigate risks before finally launching it in the Google AI Ultra subscription, the laboratory has strengthened the isolation of execution environments against indirect instruction injections. In this way, it monitors thought chains and internal activations to prevent agents from deviating from the assigned objective.
Finally, in the midst of the debate over the need to slow down the development of AI, Google has defended maintaining transparency in logical processes, urging industry competitors not to resort to opaque techniques that make the traceability of software decisions difficult. A dart that, without a doubt, points directly at OpenAI.
