Google introduced its new flagship model, Gemini 4 Argon. The company says it supports continuous reasoning and generation sequences of up to one million output tokens, compared with 64,000 in the previous generation. The model also took first place in the Vals Index, which evaluates tasks resembling professional workflows. Access begins with trusted cybersecurity partners.
Longer generation and evaluation settings
Google describes the million-token output limit as allowing extended, continuous reasoning and generation. That is a separate parameter from context-window size: the Vals test configuration gave Argon a one-million-token context window but capped its output at 262,000 tokens, below the ceiling Google advertises.
Vals tested the model through Google’s provider with temperature set to 1 and default top-p and top-k values. Reasoning effort was set to high.
A Vals Index lead with lower token consumption
In the Vals evaluation, Argon scored 68.9%, giving Google its first outright lead in the index. The tasks cover finance and law, as well as software development and tax, and are designed to approximate workflows used in those professions.
Alongside the score, Vals recorded lower token consumption. For comparable work on index tasks, Argon used roughly one-quarter of Claude Sonnet 5.5’s output tokens. Its input-token consumption was approximately one-third of Sonnet’s.
The Tax Agent evaluation also showed a difference in the number of turns. Vals reported that Argon averaged 19 turns against Sonnet 5.5’s 47, even though its answers were more than twice as long.
Finance and law: fewer errors, more completed tasks
Argon ranked first on Finance Agent v2 with 65.4%, according to Vals. Its strong results included work with earnings filings and financial disclosures. In the same test, it made approximately one-fifth as many tool errors as Sonnet 5.5.
In the Harvey Legal Agent evaluation, Vals recorded roughly seven times as many fully completed tasks as Sonnet 5.5. Argon also met nearly every grading criterion across 24 areas of legal practice.
Coding: ahead on applications, but not every test
In Vals’ Vibe Code Bench v1.1 results, Argon achieved perfect results on 30 applications. Claude Opus 5 did so on 25 applications and GPT-6 Astra on 24. Vals also recorded roughly one-third fewer tool calls for Argon.
On Terminal-Bench 4.0, Vals measured Argon at 57.6%, compared with Gemini 3.8 Flash’s 19.0%. Argon also completed the tasks in half the total elapsed time. Claude Opus 5.5 nevertheless retained the lead in this test with 66.4%.
In competitive programming, Argon matched GPT-6 Astra, according to Vals. It scored 100% on each of the IOI 2024, 2025 and 2026 sets.
Cybersecurity results and weaker areas
Vals ranked Argon first on CyberBench’s proof-of-concept binary exploitation tasks, where it scored 70.0%. In the benchmark’s overall standings, it finished second behind GPT-6 Sol and 18 points ahead of Sonnet 5.5.
Argon performed much worse on tasks involving graphical interfaces, according to Vals. Argon scored 4.83% on CUA-bench, finishing seventh out of eight models. Its MedScribe score was 87.4%, but that placed it 15th in the model rankings.
Formal code generation from specifications was another weak area. Vals measured Argon’s ProgramBench score at 2.5%.
Partner access first, with Gemini API availability expected later
Google is beginning a phased rollout of Argon to trusted cybersecurity partners while evaluating its safety with the U.S. government. Broader availability through the Gemini API is expected later, with no launch date announced.
Standard rates are $4 per million input tokens and $20 per million output tokens, while the introductory rates announced by Google’s Logan Kilpatrick are $2 per million input tokens and $10 per million output tokens.



