Gemini 4 Argon

Gemini 4 Argon targets long-horizon AI work: what changes

•Make Better Editorial

Google?s Gemini 4 Argon pushes output to 1M tokens and targets long-running coding, enterprise and agent workflows. Here?s what matters before broad access.

Google announced Gemini 4 Argon on September 30, positioning it for complex, long-horizon work rather than short chat turns. Argon expands the output-token limit from 64K to 1 million tokens and is designed to sustain deep reasoning across long software, research and enterprise workflows.

What Google actually released

Argon is not broadly available yet. Google says it is first rolling out to trusted cybersecurity defenders through its Fairwind Program while participating in the U.S. government?s voluntary pre-release access process. Broader access is planned for developers, enterprises and consumers, starting with paid API customers and Google AI Ultra subscribers.

Key verified numbers

1M tokens
Maximum output
Up from 64K
$2 / 1M
Intro API input
Later $4 / 1M
$10 / 1M
Intro API output
Later $20 / 1M
51.3%
AutomationBench
Google reports #1 on Zapier?s benchmark

Why the 1M-token output limit matters

Make Better analysis

A larger output budget matters when an agent needs to keep working through a task rather than return a concise answer: migrating a large codebase, researching many documents, iterating through tool calls, or producing a long artifact. It does not mean every workflow should consume hundreds of thousands of tokens. Longer trajectories can also increase cost, latency and opportunities for drift. The practical question is whether extra trajectory length improves successful completion enough to justify those costs.

Google is testing Argon on work, not just benchmarks

Google says Argon agents are already used internally for code migrations, research and infrastructure optimization. One example involved agents analyzing data-center telemetry and applying memory optimizations that freed more than 300 TiB after rollout, with estimated total savings of 500 TiB to 1 PiB. Google also reports Argon agents working on C/C++ to Rust migrations ranging from core libraries to more than 800,000 lines in the Fuchsia Zircon kernel. These are Google-reported deployments, not independent proof that every organization will see similar results.

Important limitation

Argon?s strongest capabilities are not yet generally available. Early access is restricted while Google strengthens safeguards for misuse, prompt injection and agent misalignment. Do not build a production migration plan around Argon until your required access path, limits and final pricing are confirmed.

The introductory price is temporary

Google lists introductory API pricing of $2 per million input tokens and $10 per million output tokens, with cached input at a 95% discount. The announcement says pricing will later move to $4 input and $20 output per million tokens. Evaluate workflow economics using the expected long-term price as well as the launch rate.

How to evaluate Argon when access opens

  1. Choose a genuinely long-horizon task where your current model loses context, stops early or requires repeated restarts.
  2. Define a measurable success criterion before comparing models.
  3. Track total tokens, time, tool calls, retries and human corrections?not only list price.
  4. Test prompt-injection and permission boundaries when workflows read external content or take actions.
  5. Route short predictable work to cheaper models and reserve Argon for tasks where sustained reasoning improves completion.
Bottom line

Gemini 4 Argon is notable because Google is explicitly optimizing a frontier model for long trajectories across coding, enterprise work and agents. The 1M-token output ceiling creates room for tasks that previously needed repeated handoffs, but the model is still in controlled rollout and introductory pricing is temporary. Treat Argon as a candidate for difficult long-horizon work, not a default replacement for cheaper models.

Sources & useful resources