Google holds back its new frontier model
Gemini 4 Argon goes to trusted cyber defenders first, without cyber guardrails, while broad access waits on safeguards.
2 minMain AI HubFresh · 30 Sept
Google has announced Gemini 4 Argon, its new frontier model, and is rolling it out first to a set of trusted cyber defenders through what it calls the Fairwind Program. Developers, enterprises and consumers come later: Koray Kavukcuoglu, who signs the announcement, writes that releasing capabilities at this level requires a phased approach, and that Google is taking part in the US government's voluntary process for pre-release model access while it widens availability.
The headline change for builders is room to think. Argon's output limit rises to 1M tokens, up from 64K, so a single trajectory can carry hundreds of thousands of tokens. Pricing starts at $2 per million input tokens and $10 per million output tokens, with cached input at 95% off the input price.
- DeepSWE v1.1, long-horizon software engineering: 77.9%, which Google calls a new state of the art.
- AutomationBench, Zapier's end-to-end business benchmark: first place with 51.3%.
- LVBench, long video understanding: 91.7%.
- CWE-bench v1, remediating security flaws: tied first at 68%.
Google also reports results from its own use. Argon beat a published baseline by 40% on a quantum subroutine, agents freed more than 300 TiB of memory across its data centres with 500 TiB to 1 PiB estimated in total, and agents working on a Rust port of libgav1 replaced 32K lines of hand-written SIMD code with safe Rust that runs 2.7 times faster than the port while producing identical video.
The cybersecurity framing explains the staged release. Argon was trained to find, validate and patch vulnerabilities, and for trusted defenders Google is shipping it without cyber guardrails. That is also why the safeguards get their own section: refusal training for cyber and CBRN misuse, resistance to indirect prompt injection, and monitoring of the model's chain of thought that halts execution when it strays. Google says findings from that monitoring are deliberately kept out of training, so the model is not shaped to evade it.
Retold from Google. This is a summary in our own words; follow the link for the original reporting.