How a new 1-million-token frontier model is revolutionizing software engineering, enterprise workflows, and proactive cybersecurity defense.
- Unprecedented Cognitive Endurance: Gemini 4 Argon expands its output capacity to an industry-leading 1 million tokens, enabling the model to sustain complex, multi-step reasoning to solve massive, long-horizon problems in a single trajectory.
- Transforming Enterprise and Engineering: Already driving unprecedented efficiency inside Google, Argon excels at immense codebase migrations, advanced quantum algorithmic optimization, and complex multimodal enterprise workflows across finance and law.
- Proactive Security and Rigorous Safeguards: Released initially to trusted cyber defenders without restrictive cyber guardrails, the model autonomously hunts and patches vulnerabilities while operating under a robust framework of alignment monitoring and prompt injection defenses.

The landscape of artificial intelligence is shifting from rapid, short-turn interactions to sustained, long-horizon problem solving. Leading this transition is Gemini 4 Argon, Google’s newest frontier model designed explicitly to execute complex workflows that require deep reasoning. By expanding the model’s output token limit to an unprecedented 1 million tokens—a massive leap from the previous 64K limit—Argon is granted the cognitive headroom to think deeply and generate hundreds of thousands of tokens in a single, uninterrupted trajectory. This architectural shift transforms the AI from a simple assistant into an autonomous digital workforce capable of tackling enterprise knowledge work, software engineering, and defensive cybersecurity at a systemic level.
To support broad adoption among developers and enterprises, Argon is launching with an aggressive introductory pricing model. Access will cost $2 per million input tokens and $10 per million output tokens, with cached input tokens heavily discounted at 95% off the standard rate. This economic accessibility is paired with a phased rollout, prioritizing safety and feedback. Currently available to trusted cyber defenders through the Fairwind Program, Google is also actively engaging in the U.S. government’s voluntary pre-release model access process before a wider release to paid API customers and Google AI Ultra subscribers.

Inside Google, Gemini 4 Argon is already fundamentally altering engineering productivity and infrastructure management. A team of Argon agents recently analyzed fleet-wide profiling telemetry to autonomously identify and execute memory optimizations across Google’s global data centers. This initiative is actively freeing up over 300 TiB of memory, with projected total savings scaling between 500 TiB and 1 PiB. In the realm of theoretical physics, Argon assisted quantum computing researchers in optimizing the spacetime resources—measured in qubits multiplied by gates—of critical subroutines, beating the published baseline by 40% in just minutes.
Nowhere is Argon’s capability more apparent than in large-scale codebase migrations. The model is actively migrating legacy C and C++ codebases into memory-safe Rust across Google’s infrastructure. These projects scale from tens of thousands of lines in core libraries like re2 up to more than 800,000 lines for the Fuchsia Zircon kernel. In a standout example involving libgav1, Google’s open-source video decoding software, Argon agents autonomously replaced 32,000 lines of SIMD code. By running iterative, profile-guided experiments and analyzing compiler outputs, the model produced safe Rust code that allowed automatic compiler vectorization. The result is a memory-safe video decoder that matches the optimized C++ output while running 2.7 times faster than the original Rust port.

Beyond internal software engineering, Argon dominates a wide array of industry benchmarks. It sets a new state-of-the-art score of 77.9% on DeepSWE v1.1, a rigorous test of real-world, long-horizon software engineering tasks. Its proficiency extends deep into the enterprise, ranking as the leading model on the Vals Index, which measures economic impact across finance, coding, legal, and tax sectors weighted by their U.S. GDP contributions. Argon also captured the #1 spot on AutomationBench with a 51.3% score for end-to-end business execution, and demonstrated leading performance on both the Vals Finance Agent v2 and Harvey’s Legal Agent Benchmark. Furthermore, Argon excels in visual and multimodal tasks, achieving a state-of-the-art 91.7% on LVBench for long video understanding, proving its ability to drive professional chart analysis and extract actionable data from extensive media.
As cyber threats grow increasingly sophisticated, Argon serves as a critical new asset for proactive defense. Recognizing the need to equip security teams with advanced tools, Google is releasing Argon without restrictive cyber guardrails to trusted defenders and internal teams. The security firm Wiz is already deploying the model through its “Scan for Good” initiative to protect critical public infrastructure. In an early deployment, Argon identified a critical vulnerability exposing sensitive healthcare data in global hospital software—a severe risk that previous frontier models failed to detect. Argon also ties for first place on CWE-bench v1 with a 68% score for remediating vulnerabilities, and consistently outperforms its predecessor, 3.8 Flash Cyber, in attack surface discovery and black-box penetration testing across multiple programming languages.

Deploying capabilities of this magnitude requires a proportional investment in safety and alignment. Google is hardening its systems by sealing sandboxed environments prior to high-risk evaluations and sharing these agent control best practices across the ecosystem. To defend against misuse, Argon adheres to the Frontier Safety Framework, utilizing monitored internal activations to prevent the facilitation of CBRN (chemical, biological, radiological, and nuclear) attacks while preserving legitimate scientific research. The model also leads the Gray Swan Indirect Prompt Injection benchmark, proving highly resilient against malicious hijacking attempts. Crucially, Google has deployed advanced misalignment mitigations that actively monitor Argon’s chain-of-thought, pausing execution if the model attempts to circumvent user intentions. By preserving this reasoning transparency without feeding incident reports back into the training data, developers can safely diagnose alignment risks, ensuring that as AI grows more capable, it remains a secure and reliable partner in solving the world’s most complex problems.
