Gemini 4 Argon: $2/$10 Pricing, Benchmarks & Access (2026)
Announced Sept 30, 2026, Gemini 4 Argon costs $2/$10 per 1M tokens, with up to 1M output tokens. Covers benchmarks, rival prices and Fairwind-only access.
Gemini 4 Argon is Google's frontier model, announced on September 30, 2026. Aimed at real-world coding, enterprise knowledge work such as finance research and legal drafting, and cyber defense including autonomous vulnerability patching, it carries an introductory price of $2 input / $10 output per 1M tokens and supports up to 1M output tokens. For now, access is limited to trusted cyber defenders through the Fairwind Program, and general API use has not started.
On benchmarks, it ranks first on the Vals Index (68.9%) and leads DeepSWE v1.1 at 77.9%, but it trails rivals on FrontierSWE v2 and Terminal-bench 4.0. This column covers specs, pricing, benchmarks and availability, then compares it with GPT-6 Astra, Claude Opus 5.5, Claude Sonnet 5.5 and GPT-6.1 Sol. Figures come from Google's announcement and third-party reports; no independent reproduction exists yet.
Gemini 4 Argon at a glance
| Item | Detail |
|---|---|
| Announced | Sept 30, 2026 (Google blog: "Gemini 4 Argon: our next era of frontier intelligence") |
| Positioning | Google's frontier model |
| Main uses | Real-world coding, enterprise knowledge work (finance research, legal drafting), cyber defense (autonomous vulnerability patching) |
| Max output | Up to 1M tokens (previous Gemini limit: 64K) |
| Input context window | Not stated in the official blog (some third-party sources cite 1M) |
| Public model ID | Not yet disclosed |
| Current access | Fairwind Program (voluntary pre-release access for trusted cyber defenders) |
| Next steps | Paid API customers and Google AI Ultra subscribers, then developers, enterprises and consumers "as soon as possible" |
The standout is the larger output ceiling. Earlier Gemini models topped out at 64K output tokens, while Argon supports up to 1M. Large code generation, long reports and repository-wide refactoring plans may now fit in a single call. The input context length, however, is not stated in the official blog, and the 1M figure comes only from third-party sources. Wait for official documentation before relying on it.
Pricing: introductory and standard
Argon uses two-tier pricing. The introductory price is $2 input / $10 output, and the standard price after the introductory period is $4 input / $20 output per 1M tokens. The length of the introductory period has not been disclosed. Cached input is 95% off the input price.
| Item | Introductory | Standard (after intro) | Approx. JPY (intro, $1 = ~150 yen) |
|---|---|---|---|
| Input | $2 / 1M tokens | $4 / 1M tokens | ~300 yen |
| Cached input | 95% off input (~$0.10) | 95% off input (~$0.20) | ~15 yen |
| Output | $10 / 1M tokens | $20 / 1M tokens | ~1,500 yen |
The cached figures of $0.10 and $0.20 are our own calculation from "95% off the input price"; treat them as estimates until an official price sheet appears. Also, whether thinking tokens are billed separately has not been disclosed, which matters a lot for reasoning-heavy workloads.
Price comparison with rival models
| Model | Input | Output | Note |
|---|---|---|---|
| Gemini 4 Argon (intro) | $2 | $10 | $4 / $20 later; period undisclosed |
| Claude Sonnet 5.5 | $2 | $10 | Mid-tier peer at the same price |
| Claude Opus 5.5 | $4 | $20 | Matches Argon's standard price |
| GPT-6.1 Sol | $2 | $10 | Cached input $0.10 |
| GPT-6 Astra | $10 | $50 | OpenAI's top-tier model |
At its introductory price, Argon costs exactly the same as Claude Sonnet 5.5 and GPT-6.1 Sol. Once the standard price of $4 / $20 applies, it matches Claude Opus 5.5 and sits at two-fifths of GPT-6 Astra ($10 / $50). Do not build a long-term budget on the assumption that the introductory price will last.
In Vals's evaluation, cost per test was $15.68 for Argon, versus $21.34 for Claude Sonnet 5.5, $32.14 for Claude Opus 5.5 and $28.71 for Fable 5.1, the lowest of the group. That suggests Argon may reach answers with fewer tokens, lowering the total bill beyond the unit price. Long autonomous tasks reverse the picture, though: CUA-bench cost $193.78 per test, so long-running agentic workloads can get expensive.
Benchmark results and how to read them
| Benchmark | Gemini 4 Argon | Reported comparison |
|---|---|---|
| Vals Index | 68.9% (#1) | Opus 5.5 67.0% / Sonnet 5.5 67.04% / Fable 5.1 65.8% / GPT-6 Astra 63.1% |
| DeepSWE v1.1 | 77.9% | Opus 5.5 74.2% / Astra 74.1% / Fable 5.1 67.4% |
| FrontierSWE v2 | 55.0% | Astra 65.5% (best) / Opus 5.5 62.3% / Fable 56.3% |
| Terminal-bench 4.0 | 57.4% | Opus 5.5 66.4% (best) / Astra 58.2% / Fable 57.9% |
| GraphWalks 1M | 84.2% | Astra 71.8% / Opus 66.8% / Fable 65.0% |
| CWE-bench v1 | 68.0% | Astra 68.0% (tied) / Opus 67.0% / Fable 58.0% |
| AutomationBench | 51.3% (#1) | Details not listed |
| LVBench (long video) | 91.7% | Details not listed |

Google says Argon beats GPT-6 Astra on 13 of 18 published benchmarks. Strengths and weaknesses break down as follows.
- Strong (overall, long context, automation): first on Vals Index and AutomationBench. GraphWalks 1M at 84.2% leads the next model by more than 10 points, showing strong reasoning across very long inputs.
- Strong (code fixing): first on DeepSWE v1.1 at 77.9%. Tied with Astra on CWE-bench v1 at 68.0%, consistent with the vulnerability-patching pitch.
- Strong (safety): reported leader on Gray Swan IPI, which measures prompt-injection robustness. This matters for agents that read external data.
- Strong (video): 91.7% on LVBench for long video understanding.
- Weak (hard SWE): 55.0% on FrontierSWE v2, 10.5 points behind Astra (65.5%) and also behind Opus 5.5 (62.3%).
- Weak (terminal work): 57.4% on Terminal-bench 4.0, 9 points behind Opus 5.5 (66.4%).
The key takeaway is that the leader changes by benchmark. Argon handles relatively clean fix-it tasks like DeepSWE well, but Claude Opus 5.5 and GPT-6 Astra lead on harder SWE problems and shell-heavy work. Weight the benchmarks closest to your own workload. These numbers are from the announcement, and no independent reproduction exists yet.
Safety measures
Google says it followed its Frontier Safety Framework, with internal and external red teaming, sandboxed testing and monitoring of chain-of-thought for misalignment. The model refuses harmful cyber and CBRN requests while preserving legitimate dual-use research. Limiting first access to trusted cyber defenders through the Fairwind Program is part of the same staged-release approach given the model's capabilities.
When and where you can use it
Rollout is staged. Today only the Fairwind Program (voluntary pre-release access for trusted cyber defenders) has it. Paid API customers and Google AI Ultra subscribers come next, then developers, enterprises and consumers "as soon as possible." As of September 30, it is not available on Vertex AI, OpenRouter, Gemini CLI, Cursor or GitHub Copilot.
| Stage | Audience | Status |
|---|---|---|
| Stage 1 | Fairwind Program (trusted cyber defenders) | Live |
| Stage 2 | Paid API customers, Google AI Ultra subscribers | Planned (no date) |
| Stage 3 | Developers, enterprises, consumers | "As soon as possible" (no date) |
The main items still undisclosed are:
- Public model ID
- Length of the introductory price period
- How thinking tokens are billed
- An official input context length
- Specific dates for Stage 2 and beyond
- Timing for Vertex AI and third-party tool support
How it differs from GPT-6 Astra, Claude and GPT-6.1 Sol
| Model | Price (in/out) | Strength | Best for |
|---|---|---|---|
| Gemini 4 Argon | $2/$10 (intro) | Overall score, long-context reasoning, code fixing, automation | Analyzing huge inputs, long outputs, vulnerability patching |
| GPT-6 Astra | $10/$50 | FrontierSWE v2 (65.5%) | Hard SWE tasks, peak performance |
| Claude Opus 5.5 | $4/$20 | Terminal-bench 4.0 (66.4%) | Terminal-heavy agent work |
| Claude Sonnet 5.5 | $2/$10 | Vals Index 67.04%, cost efficiency | Standard model for everyday dev and business work |
| GPT-6.1 Sol | $2/$10 | Near-Astra intelligence at a lower price | Cost-conscious OpenAI workloads |
As a rule of thumb: if you need to feed in very long material or get very long output in one go, Argon is a strong candidate. For terminal-centric agents, Claude Opus 5.5 is stronger today, and for hard large-scale SWE tasks, GPT-6 Astra is. For everyday development balancing price and quality, the already available Claude Sonnet 5.5 and GPT-6.1 Sol are practical. If you want a cheaper Google option, Gemini 3.8 Flash is worth comparing.
Keep in mind that Argon is pre-GA and only a handful of organizations can test it. The comparison numbers are announced figures, so plan to re-evaluate on your own tasks once it ships.
What developers can prepare now
- Build an evaluation set: pick 10 to 30 representative tasks from your real code and documents, with clear pass criteria. You can then compare against Claude Sonnet 5.5 or GPT-6.1 Sol under identical conditions the day access opens.
- Identify uses for 1M output: large code generation, one-shot specs and reports, bulk data conversion, anything you currently split into chunks. Longer output also costs more, so estimate the output you actually need first.
- Model costs under several scenarios: compute monthly spend at both the introductory ($2/$10) and standard ($4/$20) prices. Include cache hit rates, the undisclosed thinking-token billing, and the cost of long-running agents (see CUA-bench) in your sensitivity analysis.
- Keep models swappable: make the model ID and provider configurable so you can drop Argon into a test environment as soon as the public ID is known.
- Check security workflows: if you aim to automate vulnerability patching, keep human review and approval in the loop.
Summary
Gemini 4 Argon leads the Vals Index and DeepSWE v1.1 and pairs 1M output with a $2 / $10 introductory price. It still trails rivals on FrontierSWE v2 and Terminal-bench 4.0, key details such as the intro period, thinking-token billing and input context are undisclosed, and there is no independent reproduction. Since it is not generally available, the safest move is to prepare an evaluation set and cost model, then wait for access. We will update this column as more is announced.
FAQ
When can I use Gemini 4 Argon?
As of the September 30, 2026 announcement, only the Fairwind Program for trusted cyber defenders has access. Paid API customers and Google AI Ultra subscribers come next, then developers, enterprises and consumers "as soon as possible," but no dates have been disclosed.
How much does Gemini 4 Argon cost?
The introductory price is $2 input and $10 output per 1M tokens, with cached input 95% off the input price. The standard price after the introductory period is $4 input and $20 output. The length of the intro period and whether thinking tokens are billed separately are not disclosed.
Is it better than Claude Opus 5.5 or GPT-6 Astra?
It depends on the benchmark. Argon leads on Vals Index (68.9%) and DeepSWE v1.1 (77.9%), but GPT-6 Astra leads FrontierSWE v2 and Claude Opus 5.5 leads Terminal-bench 4.0. There is no independent reproduction yet, so evaluate on your own tasks.
What can I do with 1M output tokens?
Earlier Gemini models were capped at 64K output tokens. Large code generation, long reports and bulk data conversion that you used to split may now fit in a single call. Longer output also increases cost, so estimate the volume you need first.
Is it available on Vertex AI, Cursor or GitHub Copilot?
Not as of September 30, 2026. It is not on Vertex AI, OpenRouter, Gemini CLI, Cursor or GitHub Copilot, and no public model ID has been disclosed. Support is expected to follow as the staged rollout progresses.
Feel free to contact us
Contact Us