Google is not treating this as a routine model update. Gemini 4 Argon, which the company announced on September 30, is starting its rollout through a restricted program for trusted cybersecurity professionals before it reaches developers and enterprises more broadly. That phased approach tells you something: this is a model Google believes needs careful handling before it goes wide.
Argon is positioned as a frontier reasoning model built for long, complex workflows. Not chatbot tasks. Not quick summarization. The target use cases are software engineering at scale, legal and financial research, and cybersecurity defense — workloads where a model needs to sustain coherent reasoning across hours of output, not seconds.
A 1 million output token limit changes what’s possible
The headline spec is the output token limit. Google has expanded it from 64K to 1 million tokens. That’s not just a number to put in a spec sheet. It means a model can generate the equivalent of a full novel’s worth of structured output in a single run, which matters enormously for tasks like large codebase migrations or multi-step legal analysis. No competitor currently matches that at this scale on output specifically. Anthropic’s Claude 3.7 Sonnet and OpenAI’s o3 both have strong reasoning credentials, but neither offers this kind of output ceiling in production.
On pricing, Argon launches at $2 per million input tokens and $10 per million output tokens, with cached input tokens at 95% off the standard input price. That’s not cheap, but it’s competitive with comparable frontier-tier models. For context, OpenAI’s o3 runs at $10 per million input and $40 per million output, so Google is undercutting on both ends.
What it’s actually doing inside Google
Google says Argon is already running in internal production, and the examples they share are worth taking seriously. A team used Argon agents to analyze memory usage across Google’s data centers and apply optimizations that freed over 300 TiB of memory, with an estimated 500 TiB to 1 PiB in total savings expected. That’s not a benchmark — that’s fleet-level infrastructure work.
On quantum computing, Argon beat a published baseline by 40% when optimizing qubit and gate efficiency for research subroutines. And on codebase migration, Argon agents are rewriting C and C++ into Rust across Google’s infrastructure, including the Fuchsia Zircon kernel at over 800,000 lines of code. For the libgav1 video decoding library, Argon rewrote 32,000 lines of SIMD code into safe Rust that the compiler vectorizes automatically, producing a memory-safe decoder that runs 2.7x faster than the existing Rust port.
Benchmark performance across coding, legal, and finance
Argon’s benchmark numbers are strong across the board. Key results include:
- 77.9% on DeepSWE v1.1, measuring real-world long-horizon software engineering
- 51.3% on AutomationBench from Zapier, ranking first across business automation tasks
- 91.7% on LVBench for long video understanding, currently state of the art
- First place on Vals Index, which weights performance across finance, coding, legal, and tax by U.S. GDP contribution
- 68% on CWE-bench v1 for security vulnerability remediation, tied for first
The Vals Index result is particularly notable. It’s not a synthetic coding test — it measures economic output across the kinds of knowledge work that enterprises actually pay for. Leading there matters more than winning on another math benchmark.
Cybersecurity mode and the risks that come with it
Google is releasing a version of Argon without cybersecurity guardrails for trusted defenders and internal teams. The intent is to give cyber professionals access to the model’s full capability for finding and patching vulnerabilities. Wiz is already using it through its Scan for Good program and reportedly used Argon to find a critical vulnerability in healthcare software deployed in hospitals globally — one that previous frontier models had missed.
But releasing a frontier model with no cyber guardrails, even to a vetted group, is a real policy decision. Google says it’s using activation monitoring and red team testing to prevent misuse for CBRN or cyberattack purposes. So the broad release is still gated. Still, the dual-use tension here is genuine, and it’s the reason the rollout is staged rather than immediate.
Who should pay attention
If you’re building enterprise software, running a security team, or doing anything that involves sustained, multi-step AI workflows, Argon is worth evaluating when access opens up. It’s not the right model for lightweight tasks — the pricing and complexity overhead don’t make sense there. But for the workloads it targets, Google has built something that looks genuinely ahead of what’s currently available.



