Gemini 4 Argon can generate up to one million output tokens, Google says

Gemini 4 Argon can generate up to one million output tokens, Google says

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
3. 10. 2026
3 minutes reading · 17 views
Listen to the article
Audio version of the article
Gemini 4 Argon can generate up to one million output tokens, Google says

Google introduced its new flagship model, Gemini 4 Argon. The company says it supports continuous reasoning and generation sequences of up to one million output tokens, compared with 64,000 in the previous generation. The model also took first place in the Vals Index, which evaluates tasks resembling professional workflows. Access begins with trusted cybersecurity partners.

Longer generation and evaluation settings

Google describes the million-token output limit as allowing extended, continuous reasoning and generation. That is a separate parameter from context-window size: the Vals test configuration gave Argon a one-million-token context window but capped its output at 262,000 tokens, below the ceiling Google advertises.

Vals tested the model through Google’s provider with temperature set to 1 and default top-p and top-k values. Reasoning effort was set to high.

A Vals Index lead with lower token consumption

In the Vals evaluation, Argon scored 68.9%, giving Google its first outright lead in the index. The tasks cover finance and law, as well as software development and tax, and are designed to approximate workflows used in those professions.

Alongside the score, Vals recorded lower token consumption. For comparable work on index tasks, Argon used roughly one-quarter of Claude Sonnet 5.5’s output tokens. Its input-token consumption was approximately one-third of Sonnet’s.

The Tax Agent evaluation also showed a difference in the number of turns. Vals reported that Argon averaged 19 turns against Sonnet 5.5’s 47, even though its answers were more than twice as long.

Finance and law: fewer errors, more completed tasks

Argon ranked first on Finance Agent v2 with 65.4%, according to Vals. Its strong results included work with earnings filings and financial disclosures. In the same test, it made approximately one-fifth as many tool errors as Sonnet 5.5.

In the Harvey Legal Agent evaluation, Vals recorded roughly seven times as many fully completed tasks as Sonnet 5.5. Argon also met nearly every grading criterion across 24 areas of legal practice.

Coding: ahead on applications, but not every test

In Vals’ Vibe Code Bench v1.1 results, Argon achieved perfect results on 30 applications. Claude Opus 5 did so on 25 applications and GPT-6 Astra on 24. Vals also recorded roughly one-third fewer tool calls for Argon.

On Terminal-Bench 4.0, Vals measured Argon at 57.6%, compared with Gemini 3.8 Flash’s 19.0%. Argon also completed the tasks in half the total elapsed time. Claude Opus 5.5 nevertheless retained the lead in this test with 66.4%.

In competitive programming, Argon matched GPT-6 Astra, according to Vals. It scored 100% on each of the IOI 2024, 2025 and 2026 sets.

Cybersecurity results and weaker areas

Vals ranked Argon first on CyberBench’s proof-of-concept binary exploitation tasks, where it scored 70.0%. In the benchmark’s overall standings, it finished second behind GPT-6 Sol and 18 points ahead of Sonnet 5.5.

Argon performed much worse on tasks involving graphical interfaces, according to Vals. Argon scored 4.83% on CUA-bench, finishing seventh out of eight models. Its MedScribe score was 87.4%, but that placed it 15th in the model rankings.

Formal code generation from specifications was another weak area. Vals measured Argon’s ProgramBench score at 2.5%.

Partner access first, with Gemini API availability expected later

Google is beginning a phased rollout of Argon to trusted cybersecurity partners while evaluating its safety with the U.S. government. Broader availability through the Gemini API is expected later, with no launch date announced.

Standard rates are $4 per million input tokens and $20 per million output tokens, while the introductory rates announced by Google’s Logan Kilpatrick are $2 per million input tokens and $10 per million output tokens.

Advertisement

Content created with help from UpTier.

SEO and GEO on autopilot. UpTier’s multi-agent systems write and optimize content for search engines and AI answers.

Discover UpTier ↗

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

Qwen-Image-2.1 combines image generation and editing with transparent outputQwen-Image-2.1 combines image generation and editing with transparent output
Alibaba has released a single checkpoint for image generation and editing. Qwen-Image-2.1 supports transparent PNGs and, according to the company, accepts up to 10 reference images. Commercial deployment requires a separate license.
3 min read
3. 10. 2026
CoreWeave launches Forge for training and continuously improving AI agentsCoreWeave launches Forge for training and continuously improving AI agents
Forge connects model training, evaluation and improvement with insights from production. It offers post-training without a dedicated cluster, experiment analysis and isolated environments for agents.
2 min read
3. 10. 2026
Microsoft releases MAI-Transcribe-2-Streaming for live speech transcriptionMicrosoft releases MAI-Transcribe-2-Streaming for live speech transcription
MAI-Transcribe-2-Streaming produces text as audio arrives. Artificial Analysis ranked it first for final-transcript word error rate among 38 models. It is available through Voice Live API in public preview.
2 min read
3. 10. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok