Gemini 4 Argon increases its single-output limit to 1 million tokens, but Google DeepMind said initial access is restricted to selected cybersecurity teams rather than general users.
Gemini 4 Argon: Quick take
- Argon can output up to 1 million tokens at once, which refers to generated model output, not the number of input words.
- Google DeepMind’s comparison table shows Argon placed first alone in 13 of 19 benchmark tasks, and tied for first in one other task, according to Google DeepMind.
- Google has not announced a date when Gemini 4 Argon will be available to ordinary users; early access goes to trusted teams under the Fairwind program, Google said.
One million tokens, not one million words
Argon’s headline specification is the raised single-output cap, from 64,000 tokens to 1,000,000 tokens, Google DeepMind said. In practical terms, a token is a unit the model uses to process text, and tokens do not map directly to words.
Google gave work-focused examples for Gemini 4 Argon, such as assisting with large code refactors and multi-step research tasks, and said partners used Argon to scan software for security issues. Those examples come from Google and partners including the security firm Wiz, and they illustrate intended use cases more than everyday consumer gains.

Benchmarks: 13 first places, 5 tests lost
According to a benchmark table published by Google DeepMind, Gemini 4 Argon scored 77.9 percent on the long-form software engineering test DeepSWE v1.1, ahead of GPT-6 Astra at 74.1 percent and Claude Opus 5.5 at 74.2 percent, Google DeepMind reported.
But scores vary by test: Argon scored 55.0 percent on FrontierSWE v2, behind GPT-6 Astra at 65.5 percent. Google DeepMind cautioned that some Argon results were computed internally, while other models’ numbers were taken from public leaderboards or vendor submissions, and not every model was re-run under identical conditions.
An independent evaluator, Artificial Analysis, gave Argon and GPT-6 Astra the same composite index score of 53 points, and noted Argon lagged on certain terminal and coding benchmarks. That independent assessment reinforces that Gemini 4 Argon has notable strengths, but is not dominant across every measure.
When will ordinary users get access?
Google said Gemini 4 Argon is initially available to trusted cybersecurity teams selected under the Fairwind program. Google plans to expand access to paying API customers and Google AI Ultra subscribers later, but gave no date for broad public availability.

That staged rollout means most consumers cannot yet try Gemini 4 Argon themselves, so early impressions will rely on published scores and partner case studies rather than hands-on comparisons, Google said.
API pricing and what it means for developers
Google listed introductory API prices: $2 per million input tokens and $10 per million output tokens during the promotional period, and $4 and $20 respectively after the promotion ends, Google said. The company did not specify when the promotional period ends.
Those rates mean a sustained heavy workload using Gemini 4 Argon could become costly for teams that generate very large outputs. Google positioned Argon for lengthy code tasks and multi-step research, and the pricing reflects that target use case, Google said.
Google DeepMind’s 19-test comparison
The following summary is based on a Google DeepMind chart comparing Gemini 4 Argon with peer models across 19 tasks. In that chart, green shading marks the top score in each task, and “not listed” means the table did not report a particular model’s number, Google DeepMind said.

| 範疇 | 測試項目 | Gemini 4 Argon | GPT-6 Astra | Claude Fable 5.1 | Claude Opus 5.5 |
|---|---|---|---|---|---|
| 專業工作 | Vals Index | 68.9% | 63.1% | 65.8% | 67.0% |
| 專業工作 | AutomationBench(分數) | 51.3% | 41.4% | 31.4% | 42.5% |
| 專業工作 | Vals Finance Agent v2 | 65.4% | 53.5% | 58.9% | 58.6% |
| 專業工作 | Harvey’s Legal Agent Benchmark | 19.6% | 5.4% | 6.7% | 3.8% |
| 程式開發 | DeepSWE v1.1 | 77.9% | 74.1% | 67.4% | 74.2% |
| 程式開發 | FrontierSWE v2 | 55.0% | 65.5% | 56.3% | 62.3% |
| 程式開發 | Vibe Code Bench | 91.9% | 89.6% | 90.3% | 90.3% |
| 程式開發 | Terminal-bench 4.0 | 57.4% | 58.2% | 57.9% | 66.4% |
| 機器學習工程 | PostTrainBench | 45.3% | 44.3% | 40.2% | 49.3% |
| 科學及數學 | Terminal-Bench Science 0.1 | 57.6% | 68.1% | 52.6% | 63.3% |
| 科學及數學 | LABBench 2 | 88.8% | 85.4% | 68.6% | 73.1% |
| 科學及數學 | RiemannBench | 76.0% | 72.0% | 65.6% | 69.6% |
| 長篇內容理解 | GraphWalks(最多 12.8 萬 Token) | 99.7% | 98.7% | 91.4% | 90.6% |
| 長篇內容理解 | GraphWalks(25.6 萬至 100 萬 Token) | 84.2% | 71.8% | 65.0% | 66.8% |
| 電腦操作 | Agent’s Last Exam(通過率) | 39.5% | 34.2% | 未列出 | 38.2% |
| 電腦操作 | OSWorld-2.0(離線測試部分得分) | 69.2% | 72.6% | 未列出 | 未列出 |
| 圖像及影片理解 | Chartography | 71.6% | 71.0% | 46.2% | 66.3% |
| 圖像及影片理解 | LVBench | 91.7% | 87.5% | 79.7% | 83.7% |
| 網絡保安 | CWE-bench v1 | 68.0% | 68.0% | 58.0% | 67.0% |
Note: The above summary follows Google DeepMind’s published comparisons for 19 tasks. Some Argon results were computed by Google DeepMind, and other models’ figures may come from public leaderboards or vendor reports. Test settings vary by task; see Google DeepMind’s official evaluation methodology for details: https://deepmind.google/models/evals-methodology/gemini-4-argon/.



