Google Gemini 4 Argon Released: 1M Output Tokens, DeepSWE SOTA — How It Compares for Coding and Cyber Defense
On Sept 30, Google DeepMind launched Gemini 4 Argon with 1M max output tokens per response (vs 128K for rivals), a new DeepSWE v1.1 SOTA at 77.9%, and introductory pricing of $2/$10 — half of Claude Opus 5.5. The model is purpose-trained for long-horizon software engineering, enterprise knowledge work, and cybersecurity defense, available through the Fairwind Program. This article compares Argon vs GPT-6 Astra and Claude Opus 5.5 across benchmarks.
Public-source compilation
Synthesized from public posts/docs. Prefer the original source for primary claims.
How we review content · Primary source · TechCrunch / MarkTechPost / Google Blog
TL;DR
Google DeepMind launched Gemini 4 Argon on Sept 30 — the first model of the Gemini 4 generation. Three breakthroughs: 1M max output tokens per response (rivals cap at 128K), a new DeepSWE v1.1 SOTA at 77.9%, and introductory pricing of $2/$10 — half of Claude Opus 5.5. But access is restricted: currently available only to cybersecurity partners through the Fairwind Program.
Why 1M Output Tokens Matters
Current frontier model output limits compared:
| Model | Max Output Tokens |
|---|---|
| Gemini 4 Argon | 1,000,000 |
| Claude Opus 5.5 | 128,000 |
| Claude Fable 5.1 | 128,000 |
| GPT-6 Astra | 128,000 |
This means developers can complete large-scale refactors or generate long reports in a single response without splitting across turns. The cost: a full 1M output tokens costs $10 at introductory pricing, $20 after.
Google has not disclosed Argon's input context window size.
Benchmarks: Where Argon Leads and Trails
Google compared Argon against GPT-6 Astra, Claude Opus 5.5, and Claude Fable 5.1 across 18 benchmarks. Argon leads outright on 12 and ties for first on 1.
Where Argon leads:
| Benchmark | Argon | Opus 5.5 | GPT-6 Astra | Notes |
|---|---|---|---|---|
| DeepSWE v1.1 | 77.9% | 74.2% | 74.1% | Long-horizon SE (new SOTA) |
| Vals Index | 68.9% | 67.0% | 63.1% | Economic impact (finance/coding/legal/tax) |
| AutomationBench | 51.3% | 42.5% | — | Zapier end-to-end business execution |
| Harvey Legal Agent | 19.6% | — | 5.4% | Legal agent |
| LVBench | 91.7% | — | — | Long video understanding (new SOTA) |
| CWE-bench v1 | 68% (tie) | 67% | — | Vulnerability remediation |
Where Argon trails:
| Benchmark | Argon | Leader | Gap |
|---|---|---|---|
| FrontierSWE v2 | 55.0% | GPT-6 Astra 65.5% | -10.5pp |
| Terminal-Bench 4.0 | 57.4% | Opus 5.5 66.4% | -9.0pp |
| OSWorld-2.0 | 69.2% | GPT-6 Astra 72.6% | -3.4pp |
Artificial Analysis reports that using discounted prices, Argon matches GPT-6 Astra on the Intelligence Index at 60% of the cost per task.
Pricing Comparison
| Model | Input / Output (per 1M tokens) | Cached Input |
|---|---|---|
| Gemini 4 Argon (intro) | $2 / $10 | $0.10 |
| Gemini 4 Argon (regular) | $4 / $20 | — |
| Claude Opus 5.5 | $4 / $20 | $0.20 |
| Claude Fable 5.1 | $10 / $50 | $0.25 |
| GPT-6 Astra | $10 / $50 | $1.00 |
At introductory pricing, Argon's output cost is 1/5 of GPT-6 Astra's.
Cyber Defense: Find, Validate, Patch
Argon is trained to autonomously find, validate, and patch critical software vulnerabilities. Trusted defenders and internal Google teams receive it without cyber guardrails.
On CWE-bench v1 (vulnerability remediation), Argon ties for first at 68%.
Real-world case: Wiz is already using Argon through its Scan for Good initiative. The model found a critical vulnerability in healthcare software used by hospitals worldwide — one that previous frontier models had missed.
Google is strengthening safeguards in 4 areas before broad release:
- Misuse defenses for cyber and CBRN risks (including activation monitoring under its Frontier Safety Framework)
- Indirect prompt injection resistance (Argon leads Gray Swan's IPI benchmark)
- Misalignment monitoring of chain-of-thought and actions (with ability to stop execution)
- Sealed, isolated sandboxes for high-risk training and evaluations
Google's Internal Results
Thousands of Googlers already use Argon. Google shared 4 internal results:
- Memory optimization: Agents freed over 300 TiB across data centers, with 500 TiB to 1 PiB projected
- SIMD code replacement: Replaced 32K lines of SIMD code in libgav1 Rust port; decoder runs 2.7x faster with identical output
- C/C++ to Rust migration: Up to 800K+ lines in the Fuchsia Zircon kernel
- Quantum algorithm: Beat a published quantum algorithm baseline by 40% in minutes
Access and Availability
Argon is currently only available through the Fairwind Program to cybersecurity partners — no public release date. Google is participating in the U.S. government's voluntary pre-release model access process, gathering feedback from early testers and iterating on guardrails before wider release.
This means most developers and enterprises can't access Argon directly in the short term — a stark contrast to OpenAI's GPT-6.1 Sol, announced the same day at DevDay 2026 and already available to Pro/Business users.
Developer Selection Guide
| Use Case | Recommended Model | Rationale |
|---|---|---|
| Long-horizon SE (large refactors) | Gemini 4 Argon (if accessible) | 1M output + DeepSWE SOTA |
| Daily Agentic Coding | GPT-6.1 Sol | 92% Astra capability, 1/5 price, available now |
| Terminal/CLI tasks | Claude Opus 5.5 | Terminal-Bench 4.0 leader |
| Computer use (GUI automation) | GPT-6 Astra | OSWorld-2.0 leader |
| Cybersecurity defense | Gemini 4 Argon | Purpose-trained for vuln discovery/patching |
| Legal/compliance agent | Gemini 4 Argon | Harvey Legal Agent significantly ahead |
| Cost-sensitive long workflows | GPT-6.1 Sol / Argon (intro price) | Both at 1/5~1/2 of competitors' pricing |
Sources
- Google Blog: https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/
- TechCrunch: https://techcrunch.com/2026/09/30/google-releases-gemini-4-argon-called-its-most-powerful-model-yet/
- MarkTechPost: https://www.marktechpost.com/2026/09/30/google-deepmind-unveils-gemini-4-argon-with-1m-output-tokens-for-coding-knowledge-work-and-cyber-defense/
- CNBC: https://www.cnbc.com/2026/09/30/google-gemini-4-argon-ai.html
Topic hub
AI Coding Tools Hub (2026)
From Copilot pricing changes to Claude Code + DeepSeek cost-saving setups—one place to compare tools, read explainers, and follow tutorials.
Explore AI Coding Tools Hub (2026) →Monetization angle
How can you make money from this trend?
WayToClawEarn focuses on verified earn playbooks—not just news. Start from these cases.
AI code review & spec-driven agency
Offer migration consulting as Copilot pricing shifts
Claude Code 48h Micro SaaS
Validate products fast with a low-cost agent stack
Related tutorials
- Claude Code Too Expensive? Cut 90%+ Cost with DeepSeek V4 in 10 Minutes
- GitHub Copilot Pricing 2026: Plan Comparison, Monthly Cost, and 3 Ways to Save
- How to turn off signature in VS Code Copilot AI: remove Co-Authored-by with one line of command
- Copilot vs Cursor vs Claude Code (2026): Which Should You Pick?
Related news
- OpenAI DevDay 2026: GPT-6.1 Sol at 1/5 Astra Price, Dots Autonomous Agent, Codex Cloud — What Developers Need to Know
- Plugin4Shell: Zero-click RCE hits four major AI coding agents — how to fix
- Claude Opus 5.5 launched: 40% cheaper, new coding SOTA — which model to pick?
- GPT-6 Sol and Luna launch: API prices permanently cut 50%, which model should developers pick?