Google introduced Gemini 4 Argon on Wednesday, but the company is not opening its new flagship AI model to ordinary developers or subscribers yet. The first users are a limited group of vetted cyber defenders in Google’s Fairwind Program, who will test Argon’s strongest security capabilities before a wider rollout.
The staged release makes Argon an unusual flagship launch. Google has already disclosed API prices, a one-million-token output limit and a wide set of benchmark results, yet it has given no date when most customers can use the model. Paid Gemini API customers and Google AI Ultra subscribers will be first in line after the security testing phase, followed by broader enterprise and consumer access.
For developers and technology buyers, that split matters more than the headline benchmark rankings. Argon may be Google’s strongest model on paper, but its real value will depend on access, sustained-task reliability and whether Google’s internal results hold up in independent production tests.
Gemini 4 Argon access starts with cyber defenders
Google is distributing an unrestricted version of Argon’s cyber capabilities to trusted defenders through the Fairwind Program. The company is also taking part in the U.S. government’s voluntary pre-release model-access process.
That version is designed to find, validate and patch software vulnerabilities autonomously. Wiz is already using it in Scan for Good, a free program that looks for high-risk exposures in critical public infrastructure. Google reports that Argon found a critical flaw exposing personal information in healthcare software used by hospitals worldwide after earlier frontier models missed it.
The disclosure stops short of naming the affected product, vulnerability class, attack path or remediation. That protects organizations before a fix is complete, but it also means outsiders cannot yet evaluate the most striking security claim. Google similarly cites an internal benchmark spanning complex codebases in 20 programming languages and a Wiz black-box penetration test, neither of which currently provides a public task set or reproducible result.
One public result is easier to compare: Argon scored 68% on CWE-bench v1, which tests whether models can repair real security vulnerabilities. That tied GPT-6 Astra and edged Claude Opus 5.5 by one percentage point in Google’s comparison.
A one-million-token output limit changes the job size
Argon can produce as many as one million output tokens in a single run, up from 64,000 for earlier Gemini models. This is an output ceiling, not merely a large input context window. It gives an agent room to continue a long code migration, audit or research task without splitting the work across dozens of separate responses.
The bigger number does not eliminate the engineering problems around long-running agents. A longer trajectory creates more opportunities for an early mistake to propagate, for tool state to drift or for a generated change to escape review. Teams evaluating Argon should measure completion rate, correction cost and checkpoint quality, not just whether a task fits in one call.
Google’s own deployments show the kind of work it has in mind. The company says Argon agents analyzed data-center profiling telemetry and identified memory optimizations expected to free more than 300 TiB once deployed, with total potential savings of 500 TiB to 1 PiB. Other agents are helping migrate C and C++ code to Rust, including more than 800,000 lines in the Fuchsia Zircon kernel.
For the open-source libgav1 video decoder, Google reports that agents replaced 32,000 lines of SIMD code in an existing Rust port. The resulting decoder ran 2.7 times faster while producing identical video output, according to the company. Google also notes that its large migrations remain subject to automated and manual audits, emulation tests and code review. Those controls are a useful reminder that the model is contributing to the work, not independently approving production changes.
Argon leads several benchmarks, but not all of them
Google’s comparison table gives Argon the top score or a tie on 13 of 18 disclosed evaluations. Its clearest results center on long-horizon coding, business workflows and document-heavy professional work.
| Evaluation | Gemini 4 Argon | What it tests |
|---|---|---|
| DeepSWE v1.1 | 77.9% | Long-running software engineering tasks |
| AutomationBench | 51.3% | End-to-end business workflows |
| Vals Finance Agent v2 | 65.4% | Multi-step financial research |
| LVBench | 91.7% | Understanding long videos |
| CWE-bench v1 | 68% | Repairing software vulnerabilities |
The results do not amount to a clean sweep. In the benchmark table reviewed by VentureBeat, GPT-6 Astra remained ahead on FrontierSWE v2 and Terminal-Bench Science 0.1 by 10.5 percentage points. Claude Opus 5.5 led Argon 66.4% to 57.4% on Terminal-bench 4.0 and also posted the better PostTrainBench score.
That pattern argues against choosing a model from an overall leaderboard. Argon appears strongest in Google’s selected long-context and professional-work evaluations, while rival models retain advantages on some terminal, scientific and computer-use tasks. Buyers should build a test set from their own repositories, documents, tools and approval processes before moving production work.
Gemini 4 Argon pricing starts low, then doubles
Google plans to charge $2 per million input tokens and $10 per million output tokens during an unspecified introductory period. Cached input receives a 95% discount, reducing the introductory cache-read price to $0.10 per million tokens.
After the introductory offer ends, rates rise to $4 per million input tokens and $20 per million output tokens. At list price, a fully used one-million-token output allowance would therefore cost $10 initially or $20 later, before input and tool costs. Few tasks should need an output that large, but the calculation shows why teams must pair long trajectories with hard budget limits and stop conditions.
The introductory rate matches GPT-6 Sol and Claude Sonnet 5.5 on their published input and output prices, while the later rate matches Claude Opus 5.5. Argon is substantially cheaper than GPT-6 Astra at either rate. The comparison remains provisional because Google has not announced when the introductory price ends, and developers cannot yet run representative workloads against the public API.
Google is using the rollout to test new safety controls
Argon’s phased access follows a summer in which frontier agents drew scrutiny for taking unintended actions and crossing test boundaries. Google says it is strengthening protections in four areas: cyber and chemical or biological misuse, indirect prompt injection, model misalignment and the security of agent sandboxes.
The company is monitoring internal model activations for signs of misuse and using controls that can inspect Argon’s reasoning and actions, then halt execution when behavior moves beyond a user’s intent. It says training runs used a similar monitoring system connected to a dedicated incident-response team.
On Gray Swan’s indirect prompt-injection evaluation, Argon recorded a 0.7% attack success rate in Google’s disclosed results, compared with 8.5% for GPT-6 Astra and much higher rates for several other models. A low score on one adversarial set does not make an agent immune to instructions hidden in websites, documents or messages. Production deployments still need isolated credentials, narrow tool permissions, network controls, logging and human approval for consequential actions.
What developers should prepare before access opens
Teams interested in Argon can use the waiting period to define a focused evaluation rather than planning a wholesale model switch. A useful test should include tasks that take hours rather than minutes, repositories large enough to expose navigation problems, adversarial documents that contain hidden instructions and explicit checks for whether the agent respects permission boundaries.
Cost tests should separate input, cached input, output and external tool usage. Reliability tests should record how often Argon finishes without intervention, how much reviewer time it saves and whether repeated runs converge on the same safe result. Security teams should test the guarded public model they will actually receive; results from the unrestricted Fairwind version may not predict standard API behavior.
Gemini 4 Argon gives Google a credible new claim at the frontier, backed by unusually concrete internal engineering examples and competitive pricing. Its launch is still a preview, however. Until customers can run the model independently, the most consequential details are the ones Google has not supplied: a general-release date, public evidence for its headline vulnerability discovery and production reliability across million-token trajectories.