On September 30, Google DeepMind announced its frontier model, Gemini 4 Argon. It first goes to a set of trusted cyber defenders through the Fairwind Program. The company says it is taking part in the U.S. government’s voluntary process for pre-release model access, and that it will keep gathering feedback from early testers and adjusting safeguards before access widens. Developers, enterprises, and consumers cannot use it yet. A later public release is planned to start with paid API customers and Google AI Ultra subscribers. The post does not give a date for that.

[1]
Woodblock: a blank scroll rolls up at the right, with a vermilion seal just outside the roll.
The blank scroll ends in a roll. The red seal sits just outside it. That is Argon lengthening a single output while access stays sealed and not yet public. An illustration, not a launch photo., AI-generated illustration, not a news photograph

The introductory price is $2 per million input tokens and $10 per million output tokens. Cached input tokens are priced at 95 percent off the input price. After the introductory period, the price becomes $4 per million input tokens and $20 per million output tokens. For longer tasks, the output limit rises from the previous 64,000 tokens to 1 million. The company says that room lets the model write hundreds of thousands of tokens in one trajectory and finish a hard problem in one pass. That is a description of the model, not a quota the public can use today.

[1]

The external scores have to be read with their tests. DeepSWE v1.1 measures long-horizon software engineering in real settings. Argon scores 77.9 percent, which the company calls a new best result. On the Vals Index, which weights finance, coding, legal, and tax work by contribution to U.S. GDP, the company says Argon leads. The same passage cites Vals Finance Agent v2 and Harvey’s Legal Agent Benchmark. On Zapier’s AutomationBench, which measures end-to-end execution across core business functions, Argon ranks first at 51.3 percent. On LVBench, which measures long-video understanding, the score is 91.7 percent.

The internal examples are also the company’s own account. In quantum computing, it cut the spacetime resources (qubits times gates) of subroutines that bottleneck important applications by 40 percent versus a published baseline, in a matter of minutes. A team of agents read fleet-wide profiling telemetry, found memory optimizations, and applied them across Google data centers, freeing more than 300 TiB once rolled out. The company estimates total savings of 500 TiB to 1 PiB. Work migrating C and C++ to Rust runs from tens of thousands of lines in libraries such as re2 and libgav1 up to more than 800,000 lines in the Fuchsia Zircon kernel. Because many of these systems are critical, the rewrites go through automated and manual auditing, emulation testing, and review before production. libgav1 is Google’s open-source video decoder. Agents replaced 32,000 lines of SIMD in an existing Rust port. The resulting memory-safe decoder runs 2.7 times faster than that Rust port, with identical video output, and closer to the optimized C++.

[1]

The company describes Argon as a defensive cybersecurity model that can find, validate, and patch critical software vulnerabilities on its own. For trusted defenders and Google’s own internal teams, this release does not add cyber guardrails, so they can use the full defensive capability. Wiz is already using it in Scan for Good, a program that protects critical public infrastructure at no charge. In one early demonstration in the post, the model found a severe risk in healthcare software used by hospitals worldwide: sensitive personal information was exposed, and earlier frontier models had missed it. On CWE-bench v1, which measures remediation of security vulnerabilities, Argon ties for first at 68 percent.

Before a wider release, the company lists four kinds of safeguards. The model is designed to refuse harmful requests while keeping legitimate scientific research, under its Frontier Safety Framework. On indirect prompt injection, the company says Argon leads Gray Swan’s IPI benchmark. To stop the model from going past a user’s intent in order to finish a task, the deployment monitors the chain of thought and the actions and can halt execution. A similar monitor was used in training, with alerts sent to an incident-response team. Findings are not fed back into training, so the model is not shaped to evade the monitor. Before high-risk training or evaluations, sandboxes are isolated and sealed. All of this is the company’s account of the launch.

[1]

要点

  • Argon first reaches trusted cyber defenders through Fairwind, under a U.S. voluntary pre-release process. The public cannot use it yet.
  • Introductory price: $2 per million input tokens and $10 per million output tokens, with cached input at 95 percent off. Later, $4 and $20. Output limit rises from 64,000 tokens to 1 million.
  • Company figures: 77.9 percent on DeepSWE v1.1, first place at 51.3 percent on AutomationBench, and 91.7 percent on LVBench.
  • Defenders and internal teams get the model without cyber guardrails. Wiz’s early demo involved sensitive personal information in hospital software. Monitoring findings are not fed back into training.