
Google is back at the top of the AI model race with Gemini 4 Argon. The new model can produce up to one million tokens in a single response, leads several rankings for knowledge work and, at launch, costs half as much as Anthropic’s Claude Opus 5.5. For now, though, only selected security teams can use it. It still pays to look closely, because the numbers show where Argon truly shines and where the competition stays ahead.
Key takeaways
- Gemini 4 Argon is Google’s first frontier model in more than seven months; Google skipped a version 3.5 entirely.
- Security teams in the “Fairwind” program get access first. Paying API customers and Google AI Ultra subscribers come next, with no firm date.
- The introductory price is $2 per million input tokens and $10 per million output tokens, rising later to $4 and $20.
- On Artificial Analysis’ independent index, Argon ties GPT-6 Astra but trails Claude Opus 5.5 and Sonnet 5.5.
- Notably strong: a low hallucination rate and high resistance to prompt injection. Notably high: token consumption per task.
What Argon can do that is new
The most striking change is output length. Instead of 64,000 tokens, Argon can now generate up to one million tokens in one go. According to an overview by MarkTechPost, rival models Claude Opus 5.5, Claude Fable 5.1 and GPT-6 Astra top out at 128,000 tokens each. A million tokens is roughly several novels’ worth of text; in practice it matters for large code migrations, long reports or complete translations in a single pass. To keep such responses from hitting time limits, Google is introducing an API feature called “Long Decode Continuation,” The Decoder reports. Argon accepts text, images, video and speech as input and produces text as output.
Google illustrates what this looks like in practice with internal projects. Thousands of employees already use the model. It is currently helping migrate more than 800,000 lines of C and C++ code in the kernel of the Fuchsia operating system to Rust, a programming language that rules out entire classes of memory bugs; the result is still undergoing automated and manual review before it goes into production. In the libgav1 video decoder, Argon replaced 32,000 lines of SIMD code, meaning highly optimized processor instructions. According to Google, the result runs 2.7 times faster than the existing Rust port.
The second focus is cybersecurity. Argon is designed to find, validate and fix critical software vulnerabilities on its own. That is why the model goes first to vetted defenders in the Fairwind program, explicitly in a version without the usual cyber guardrails. According to Google, the cloud security company Wiz has already used it to uncover a critical flaw in hospital software used worldwide that exposed sensitive personal data. In parallel, Google is taking part in the U.S. government’s voluntary process that gives officials access to new models before release.
The benchmarks: strong at knowledge work, not ahead everywhere
Google’s own comparison tables are emphatic. According to an analysis by VentureBeat, Argon leads outright on 12 of 18 published tests and ties for first on one more. On the DeepSWE v1.1 coding benchmark it scores 77.9 percent, compared with 74.1 percent for GPT-6 Astra and 74.2 percent for Opus 5.5. The gap is widest in legal agent work: on the legal agent benchmark from law firm software maker Harvey, Argon reaches 19.6 percent, Astra 5.4 and Opus 5.5 3.8 percent. Argon also leads clearly on finance tasks (Vals Finance Agent v2: 65.4 versus 53.5 and 58.6 percent).
The tables also reveal the gaps. On Terminal-Bench 4.0, which tests work on the command line, Opus 5.5 scores 66.4 percent, well ahead of Argon’s 57.4 percent. On FrontierSWE v2, a demanding software test, GPT-6 Astra leads 65.5 to 55.0 percent. Anyone who mainly wants Argon as a coding assistant in the terminal should not rely on the overall tally alone.
Independent measurements confirm this mixed picture. On the Artificial Analysis Intelligence Index, Argon at its “High” reasoning setting scores 53 points, tying GPT-6 Astra and Claude Fable 5.1. Claude Opus 5.5 leads with 58 points, followed by Sonnet 5.5 with 56. Compared with the preview version of Gemini 3.1 Pro, that is a jump of 23 points. On the Vals Index from Vals AI, by contrast, Argon takes first place with 68.9 percent, the first time a Gemini model has topped that list.
The catch is token consumption
On paper, Argon is cheap. The introductory price is one fifth of GPT-6 Astra ($10 and $50 per million tokens) and half of Opus 5.5 ($4 and $20); cached input costs 95 percent less. Google has not said how long the discount will last. After that, Argon will sit at the same price level as Opus 5.5. So Google is opening with an aggressive price, much like the recently launched Gemini 3.8 Live, which undercut OpenAI on price.
What really matters, however, is the cost of a completed task. Argon thinks at length: Artificial Analysis measures an average of 62,000 output tokens per task, while GPT-6 Astra gets by with 27,000. At the introductory price, an index task on Argon therefore costs $1.99, or 60 percent of what Astra charges ($3.26). Once the discount ends, that rises to $3.98, about 20 percent above Astra. GPT-6.1 Sol, unveiled at OpenAI’s DevDay in September, still handles the same tasks 2.7 times more cheaply than Argon at its introductory price.
Two measurements clearly favor Argon, though. On Artificial Analysis’ AA-Omniscience knowledge test, it makes up an answer it does not know in only 15 percent of cases; for GPT-6 Astra the figure is 51 percent. And on Gray Swan’s prompt injection test, in which hidden instructions in web pages or documents try to hijack a model, only 0.7 percent of attacks succeed against Argon, compared with 1.0 percent for Opus 5.5 and 8.5 percent for GPT-6 Astra. For agents that read outside content on their own, that is an important property.
Outlook: a strong model with no launch date
With Argon, Google is once again a serious contender for demanding knowledge work after a long dry spell, especially where long outputs, legal and financial analysis or reliable facts matter. For teams choosing a model today, however, nothing changes yet: there is neither a date for the API nor any word on availability in Europe. Anyone planning to switch later should not compare list prices but run their own tasks with real token counts. Only then will it become clear whether the low entry price makes up for the heavy reasoning effort.

