Google DeepMind has recently unveiled Gemini 4 Argon, marking a significant advancement as the inaugural model of the Gemini 4 series.
This model is specifically designed to enhance capabilities in long-horizon software engineering, enterprise knowledge sectors, particularly in legal and financial domains, as well as fortifying cybersecurity defenses.
One of the most notable technical advancements of Argon is its capacity for extended output; it can now produce responses comprising up to 1 million tokens, a significant leap from the previous 64,000 tokens limit found in earlier iterations of Gemini.
Details of Google’s Announcement
According to Google DeepMind, Argon has been meticulously crafted to handle intricate workflows encompassing coding, enterprise knowledge management, and cybersecurity imperatives.
The rollout strategy is being approached incrementally. Google is actively participating in the voluntary pre-release model access initiative alongside the U.S. government.
This allows them to gather invaluable insights from early adopters and refine the framework before launching a more widespread release.
Initial pricing details have also been disclosed. Argon is set to debut at an introductory rate of $2 per million input tokens and $10 per million output tokens.
Moreover, input tokens that are cached are eligible for a 95% discount, bringing the rate down to just $0.10 per million.
Following this introductory phase, the pricing will adjust to $4 for input and $20 for output. Logan Kilpatrick also verified the initial pricing structure.
The Significance of the 1M Output Capacity
In contrast, existing frontier APIs have considerably lower response output caps. For example, Claude Opus 5.5, Claude Fable 5.1, and GPT-6 Astra each permit only 128,000 token outputs.
The Google team asserts that Argon possesses the capability to perform extensive cognitive operations and generate substantial amounts of tokens in a singular endeavor.
This advancement facilitates developers in executing large-scale refactors or composing lengthy reports without the necessity to segment their work across multiple exchanges.
However, this comes at a tangible cost; generating an entire million output tokens will incur $10 during the introductory phase and escalate to $20 thereafter.
Details regarding Argon’s input context window remain undisclosed by Google.
Comparative Performance: Argon’s Strengths and Weaknesses
The capabilities of Argon were assessed against benchmarks including GPT-6 Astra, Claude Opus 5.5, and Claude Fable 5.1. Argon excelled, emerging victorious in 12 out of 18 benchmark tests, while sharing the top position in one test.
Areas of Dominance:
- DeepSWE v1.1 (long-horizon software engineering): Argon achieved a groundbreaking score of 77.9%. In comparison, Opus 5.5 achieved 74.2% and GPT-6 Astra 74.1%.
- Vals Index (economic impact spanning finance, coding, legal, and tax): It secured top rank with a score of 68.9%.
- AutomationBench (Zapier, end-to-end business operations): Argon leads at 51.3%, outpacing Opus 5.5, which scored 42.5%.
- Harvey Legal Agent Benchmark: Argon recorded 19.6%, compared to GPT-6 Astra’s 5.4%.
- LVBench (long video comprehension): Argon established a new standard with 91.7%.
Areas Where Argon Falls Short:
- FrontierSWE v2: Scoring 55.0%, it trails behind GPT-6 Astra’s 65.5%.
- Terminal-Bench 4.0: It scored 57.4%, lagging behind Claude Opus 5.5, which scored 66.4%.
- OSWorld-2.0 (computer usage): Argon achieved 69.2%, falling short of GPT-6 Astra’s 72.6%.
As reported by Artificial Analysis, Argon matches GPT-6 Astra in the Intelligence Index while being priced at 60% of the cost per task, utilizing discounted rates.
Cybersecurity: Identifying, Validating, Patching
Argon was trained to autonomously identify, validate, and remediate critical software vulnerabilities. Trusted defenders and internal teams at Google utilize it without the constraints of cyber guardrails.
On CWE-bench v1, which assesses vulnerability resolution, Argon ties for the foremost position with a score of 68%. Competing models on this leaderboard are operated within their own agent frameworks.
Wiz is already leveraging Argon through its Scan for Good initiative, which successfully identified a critical vulnerability in healthcare software employed by hospitals globally—an oversight of prior frontier models, according to Google.

Before expanding the release, Google is fortifying protections in four critical areas:
- Defenses against misuse concerning cyber and chemical, biological, radiological, and nuclear (CBRN) threats, incorporating activation monitoring as part of its Frontier Safety Framework.
- Resistance to indirect prompt injection, where Argon has excelled in Gray Swan’s Indirect Prompt Injection benchmark.
- Monitoring for misalignment in reasoning and actions, inclusive of the capability to halt execution.
- Creation of sealed, isolated environments for high-risk training and evaluations.
Source link: Marktechpost.com.






