WayToClawEarn
High impactTechCrunch / MarkTechPost / Google Blog

Google Gemini 4 Argon Released: 1M Output Tokens, DeepSWE SOTA — How It Compares for Coding and Cyber Defense

On Sept 30, Google DeepMind launched Gemini 4 Argon with 1M max output tokens per response (vs 128K for rivals), a new DeepSWE v1.1 SOTA at 77.9%, and introductory pricing of $2/$10 — half of Claude Opus 5.5. The model is purpose-trained for long-horizon software engineering, enterprise knowledge work, and cybersecurity defense, available through the Fairwind Program. This article compares Argon vs GPT-6 Astra and Claude Opus 5.5 across benchmarks.

Edisen Lu · WayToClawEarnVia TechCrunch / MarkTechPost / Google BlogPublished Oct 2, 2026

Reviewed from public sources · AI-assisted drafting under editorial oversight. How we work · Original source

Public-source compilation

Synthesized from public posts/docs. Prefer the original source for primary claims.

How we review content · Primary source · TechCrunch / MarkTechPost / Google Blog

TL;DR

Google DeepMind launched Gemini 4 Argon on Sept 30 — the first model of the Gemini 4 generation. Three breakthroughs: 1M max output tokens per response (rivals cap at 128K), a new DeepSWE v1.1 SOTA at 77.9%, and introductory pricing of $2/$10 — half of Claude Opus 5.5. But access is restricted: currently available only to cybersecurity partners through the Fairwind Program.

Why 1M Output Tokens Matters

Current frontier model output limits compared:

ModelMax Output Tokens
Gemini 4 Argon1,000,000
Claude Opus 5.5128,000
Claude Fable 5.1128,000
GPT-6 Astra128,000

This means developers can complete large-scale refactors or generate long reports in a single response without splitting across turns. The cost: a full 1M output tokens costs $10 at introductory pricing, $20 after.

Google has not disclosed Argon's input context window size.

Benchmarks: Where Argon Leads and Trails

Google compared Argon against GPT-6 Astra, Claude Opus 5.5, and Claude Fable 5.1 across 18 benchmarks. Argon leads outright on 12 and ties for first on 1.

Where Argon leads:

BenchmarkArgonOpus 5.5GPT-6 AstraNotes
DeepSWE v1.177.9%74.2%74.1%Long-horizon SE (new SOTA)
Vals Index68.9%67.0%63.1%Economic impact (finance/coding/legal/tax)
AutomationBench51.3%42.5%—Zapier end-to-end business execution
Harvey Legal Agent19.6%—5.4%Legal agent
LVBench91.7%——Long video understanding (new SOTA)
CWE-bench v168% (tie)67%—Vulnerability remediation

Where Argon trails:

BenchmarkArgonLeaderGap
FrontierSWE v255.0%GPT-6 Astra 65.5%-10.5pp
Terminal-Bench 4.057.4%Opus 5.5 66.4%-9.0pp
OSWorld-2.069.2%GPT-6 Astra 72.6%-3.4pp

Artificial Analysis reports that using discounted prices, Argon matches GPT-6 Astra on the Intelligence Index at 60% of the cost per task.

Pricing Comparison

ModelInput / Output (per 1M tokens)Cached Input
Gemini 4 Argon (intro)$2 / $10$0.10
Gemini 4 Argon (regular)$4 / $20—
Claude Opus 5.5$4 / $20$0.20
Claude Fable 5.1$10 / $50$0.25
GPT-6 Astra$10 / $50$1.00

At introductory pricing, Argon's output cost is 1/5 of GPT-6 Astra's.

Cyber Defense: Find, Validate, Patch

Argon is trained to autonomously find, validate, and patch critical software vulnerabilities. Trusted defenders and internal Google teams receive it without cyber guardrails.

On CWE-bench v1 (vulnerability remediation), Argon ties for first at 68%.

Real-world case: Wiz is already using Argon through its Scan for Good initiative. The model found a critical vulnerability in healthcare software used by hospitals worldwide — one that previous frontier models had missed.

Google is strengthening safeguards in 4 areas before broad release:

  1. Misuse defenses for cyber and CBRN risks (including activation monitoring under its Frontier Safety Framework)
  2. Indirect prompt injection resistance (Argon leads Gray Swan's IPI benchmark)
  3. Misalignment monitoring of chain-of-thought and actions (with ability to stop execution)
  4. Sealed, isolated sandboxes for high-risk training and evaluations

Google's Internal Results

Thousands of Googlers already use Argon. Google shared 4 internal results:

  1. Memory optimization: Agents freed over 300 TiB across data centers, with 500 TiB to 1 PiB projected
  2. SIMD code replacement: Replaced 32K lines of SIMD code in libgav1 Rust port; decoder runs 2.7x faster with identical output
  3. C/C++ to Rust migration: Up to 800K+ lines in the Fuchsia Zircon kernel
  4. Quantum algorithm: Beat a published quantum algorithm baseline by 40% in minutes

Access and Availability

Argon is currently only available through the Fairwind Program to cybersecurity partners — no public release date. Google is participating in the U.S. government's voluntary pre-release model access process, gathering feedback from early testers and iterating on guardrails before wider release.

This means most developers and enterprises can't access Argon directly in the short term — a stark contrast to OpenAI's GPT-6.1 Sol, announced the same day at DevDay 2026 and already available to Pro/Business users.

Developer Selection Guide

Use CaseRecommended ModelRationale
Long-horizon SE (large refactors)Gemini 4 Argon (if accessible)1M output + DeepSWE SOTA
Daily Agentic CodingGPT-6.1 Sol92% Astra capability, 1/5 price, available now
Terminal/CLI tasksClaude Opus 5.5Terminal-Bench 4.0 leader
Computer use (GUI automation)GPT-6 AstraOSWorld-2.0 leader
Cybersecurity defenseGemini 4 ArgonPurpose-trained for vuln discovery/patching
Legal/compliance agentGemini 4 ArgonHarvey Legal Agent significantly ahead
Cost-sensitive long workflowsGPT-6.1 Sol / Argon (intro price)Both at 1/5~1/2 of competitors' pricing

Sources

googlegemini-4-argondeepmindcodingcybersecurityllmbenchmark

View source →

Educational reference only: cases summarize public sources and may use AI-assisted drafting under editorial review. Not financial advice; outcomes are not guaranteed.