Its headline result is 77.9% on DeepSWE v1.1, ahead of GPT-6 Astra at 74.1% and Claude Opus 5.5 at 74.2%.
Gemini 4 Argon delivers frontier performance in complex workflows across real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense Google said.
Technical specifications
Benchmark comparison
| Benchmark | Gemini 4 Argon | GPT-6 Astra | Claude Fable 5.1 | Claude Opus 5.5 |
| DeepSWE v1.1 | 77.9% | 74.1% | 67.4% | 74.2% |
| Vals Index | 68.9% | 63.1% | 65.8% | 67.0% |
| AutomationBench | 51.3% | 41.4% | 31.4% | 42.5% |
| Finance Agent v2 | 65.4% | 53.5% | 58.9% | 58.6% |
| Legal Agent | 19.6% | 5.4% | 6.7% | 3.8% |
| Vibe Code Bench | 91.9% | 89.6% | 90.3% | 90.3% |
| GraphWalks 256K–1M | 84.2% | 71.8% | 65.0% | 66.8% |
| LVBench | 91.7% | 87.5% | 79.7% | 83.7% |
| CWE-bench v1 | 68.0% | 68.0% | 58.0% | 67.0% |
Argon does not lead every benchmark. GPT-6 Astra scores 65.5% vs. 55.0% on FrontierSWE v2 and 68.1% vs. 57.6% on Terminal-Bench Science 0.1. Claude Opus 5.5 leads Terminal-Bench 4.0 with 66.4% vs. 57.4%.
Gemini 4 Argon vs. previous generation
Gemini 4 Argon follows Gemini 3.1 Pro in Google's frontier model lineup.
The largest documented technical change is maximum output length:
| Previous limit | Gemini 4 Argon | |
| Maximum output | 64K | 1M tokens |
| Increase | — | 15.6× |
| Long-running coding agents | Supported | Expanded |
| Cybersecurity | General | Dedicated defensive capabilities |
Google says the larger output capacity is intended for long-running agentic workflows rather than conventional chatbot responses.
Coding and cybersecurity
Google reports using Argon for large internal software-engineering tasks, including code migrations ranging from tens of thousands of lines to more than 800,000 lines for the Fuchsia Zircon kernel.
In another project, Argon replaced 32,000 lines of SIMD code in a Rust port of Google's libgav1 video decoder. Google says the resulting implementation was 2.7× faster than the previous Rust version while producing identical output.
For vulnerability detection and patching, Argon scores 68% on CWE-bench v1, matching GPT-6 Astra in Google's tests.
Availability
Gemini 4 Argon is initially being provided to selected cybersecurity researchers through Google's Fairwind Program.
Google plans a broader rollout starting with paid Gemini API customers and Google AI Ultra subscribers. A general-release date has not been announced.