Google Gemini 4 Argon: 1M-Token Output, DeepSWE 77.9%, and the Fairwind Cybersecurity Rollout
Google DeepMind ships Gemini 4 Argon: 1M-token output (up from 64K), DeepSWE v1.1 SOTA at 77.9%, CWE-bench v1 at 68% tying the leader, Fairwind cybersecurity rollout to 650+ partners, and Google's own internal use cases (300 TiB-1 PiB memory savings, 800K LOC Rust migration, 2.7x libgav1 speedup).
Google DeepMind announced Gemini 4 Argon on September 30 — the same day as OpenAI’s DevDay keynote, which is part of why the launch got less coverage than it would in a quieter week. Argon is Google’s next frontier model, targeting long-horizon software engineering, financial research, legal drafting, and cybersecurity remediation. The headline numbers: 1 million-token output (up from 64K), DeepSWE v1.1 SOTA at 77.9%, and CWE-bench v1 at 68% tying the leader. The rollout is the unusual part — Argon ships to a “Fairwind Program” of vetted cybersecurity organizations first, not the public API.
The headline capability: 1M-token output
Gemini 4 Argon’s maximum output is 1 million tokens, up from the 64,000-token cap on previous Gemini generations. Per their announcement, the larger output window lets the model sustain deeper reasoning and execute multi-stage workflows without breaking tasks into smaller interactions.
This is the meaningful capability shift, not the raw token count. Models with high context windows but small output caps (64K) hit a wall on long-horizon work: you can read 1M tokens of codebase but you can only emit 64K tokens of code change before the model needs to be re-prompted. Argon’s 1M output window eliminates that bottleneck for migrations, refactors, and multi-file edits. It’s also the output cap GPT-6.1 Sol (released Sep 22) doesn’t match — Sol’s focus is per-token cost, not output length.
The competitive question for the rest of 2026 is whether Anthropic, OpenAI, and the open-weight community match the 1M output cap. Right now Gemini 4 Argon has a clear capability lead on long-output tasks.
DeepSWE 77.9% — new state-of-the-art
On DeepSWE v1.1, a benchmark focused on real-world long-horizon software engineering, Google reports Gemini 4 Argon at 77.9% — a new state-of-the-art on the benchmark. The “DeepSWE” family tests tasks that look like real engineering work (multi-file refactors, dependency upgrades, test-driven debugging) rather than toy code-completion. Argon’s score is the kind of number that matters for teams evaluating model choice on migration workflows.
For context: the prior frontier-model SOTA on DeepSWE-class benchmarks sat in the 65-72% range earlier in 2026. Argon’s 77.9% is a meaningful step up.
CWE-bench v1 68% — ties the leader
On CWE-bench v1, which evaluates vulnerability remediation, Argon scored 68% — tying for the highest reported score on the benchmark alongside other leading frontier models. The benchmark tests whether a model can autonomously discover, validate, and patch a known CVE-class vulnerability in a target codebase.
A 68% tie means Argon is competitive on the defensive-security side. The Fairwind tie-on isn’t just about gated access — it’s about a model that’s actually good enough at the task to warrant the program.
The Fairwind Program — gated cybersecurity rollout
Argon ships first through Google’s Fairwind Program: vetted cybersecurity organizations can access versions of Argon with advanced defensive capabilities. Google says the program works with more than 650 partners globally, prioritizing governments, critical infrastructure operators, and major technology platforms. Access is controlled through organizational vetting and authentication requirements.
This is a different release strategy from OpenAI’s Dots (Pro/Business Premium/Enterprise self-serve). Google’s gating reflects two things: (a) the model’s capability profile (autonomous vulnerability remediation is high-impact in the wrong hands), and (b) the user base that needs it most (governments, critical infrastructure) wants a relationship, not an API key.
For builders not in the cybersecurity vertical, this gating means Argon is effectively unavailable in early October 2026. The broader rollout — to developers, enterprises, and consumers — will follow on Google’s phased timeline, but no specific date is published.
Google’s own internal use cases
Google’s announcement includes concrete numbers from internal use of Argon across Google engineering:
- Memory optimization at fleet level. Google teams used Argon agents to analyze fleet-wide profiling data and identify memory optimizations across its data centers. The expected result: more than 300 TiB of memory freed, with potential total savings between 500 TiB and 1 PiB.
- C/C++ → Rust migrations. Argon has been applied to codebase migrations ranging from “tens of thousands of lines of code in core libraries” to more than 800,000 lines within the Zircon kernel used by Fuchsia OS. Migrations are subject to automated testing, emulation, and human review before deployment — Argon proposes, humans approve.
- libgav1 SIMD rewrite. Argon agents reworked 32,000 lines of SIMD code in the Rust port of the libgav1 video decoder. The resulting implementation ran 2.7 times faster than the previous Rust port while producing identical video output.
These are real numbers from real internal Google workloads. They tell you what Argon is good at today: large-scale code transformation under human review, not “run a chatbot.”
What this means for builders
Three takeaways:
- The output-window arms race is on. 1M-token output changes the calculus for migrations, refactors, and multi-step agents. Watch for OpenAI and Anthropic response in the next 90 days — if they don’t match the cap, long-horizon workflows will have a Google-shaped advantage.
- Gated rollouts are the new frontier-model pattern. OpenAI’s DevDay put Dots on Pro plans; Google put Argon behind a 650-partner Fairwind program. Both reflect the same reality: frontier models are now valuable enough that distribution is a strategic choice, not a default.
- Internal-use numbers are the most honest benchmark. Google’s published DeepSWE 77.9% and CWE-bench 68% are paper claims; the 300 TiB-1 PiB memory, 800K-LOC Zircon migration, and 2.7x libgav1 speedup are operational evidence. The latter is what tells you whether the model actually delivers.
What we still don’t know: when Argon leaves Fairwind, what the API pricing looks like, whether the open-weight community gets a 1M-output Argon variant, and how Argon’s cybersecurity defense gate interacts with Apex SDK’s Marketplace (the OpenAI enterprise launch from DevDay 2026).
Sources
- IntelligentHQ — Google Launches Gemini 4 Argon With 1M-Token Output — primary; 1M-token output, DeepSWE 77.9%, CWE-bench 68%, Fairwind program, Google internal use cases
- Magai — Gemini 4 Argon: What It Is and Who Can Use It — Sep 30 announcement context, “you can’t use it” framing
- AnotherCodingBlog — Another Daily AI Newsletter October 1, 2026 — Sep 30 Gemini 4 Argon announcement, Fairwind trusted-tester program, cyber-defender first cohort
- DemandSphere — AI Frontier Model Tracker — frontier model benchmark baselines for context
- LLM Stats — AI News Today (October 2026) — GPT-6.1 Sol benchmark comparison, model release timeline