Gemini 4Google AILLM pricingAI benchmarksAI APIs

Gemini 4 Argon: Pricing, Benchmarks and When Builders Can Use It

person
LaunchBoosts Research Desk·AI-assisted research
8 min read

Researched and drafted with AI assistance from the 8 public sources listed at the end of this article, then published after automated editorial checks. Spotted an error? Tell us at support@launchboosts.com.

Gemini 4 Argon: Pricing, Benchmarks and When Builders Can Use It — LaunchBoosts

Gemini 4 Argon is Google's new flagship AI model, announced on September 30, 2026. It costs $2 per million input tokens and $10 per million output tokens at an introductory rate, which later doubles to $4 and $20. It also raises the output limit from 64K to 1 million tokens. The catch is access: right now only vetted cybersecurity defenders in Google's Fairwind Program can use it. Paid API customers and Google AI Ultra subscribers are promised access "as soon as possible", with no date.

This guide covers what has been confirmed, how the benchmarks compare, where the pricing looks better than it really is, and what to do before the API opens up.

What Google actually announced

Argon is the first new generation of Gemini since Gemini 3 in November 2025. It also arrives months late. According to Bloomberg's reporting, Google announced a Gemini 3.5 Pro in May 2026 and promised it for June, then dropped it without a public explanation.

Confirmed details from Google's announcement and launch coverage:

  • Focus: sustained deep reasoning for software engineering, legal and financial work, cybersecurity (finding and patching vulnerabilities), and multimodal analysis of charts, long videos and sets of documents.
  • Output limit: up to 1 million output tokens, up from 64K. Artificial Analysis also lists a 1 million token context window. Fello AI reports text and image input with text output.
  • Rollout: Fairwind Program partners first. These are governments, critical-infrastructure operators and platforms doing defensive security work, and they get a version "without cyber guardrails". Paid API customers and AI Ultra subscribers come next, then wider access after more testing, according to The Next Web.
  • Safety process: Google says it is taking part in the U.S. government's voluntary pre-release model access process (VentureBeat).
  • Internal use: 9to5Google reports Google's own examples of Argon agents at work. They freed more than 300 TiB of memory across Google's data-center fleet, migrated more than 800K lines of C/C++ to Rust, and made a video decoder 2.7x faster.

Google-owned security firm Wiz is already using Argon, and it says the model found a critical vulnerability in healthcare software that earlier frontier models had missed (The Next Web). That claim comes from Google's own side, so treat it as a demo, not independent proof.

Benchmarks: where Argon leads and where it trails

All the figures below are Google's own published results, as compiled by VentureBeat. VentureBeat counts Argon as leading or tying in 13 of 18 disclosed categories.

Benchmark Gemini 4 Argon GPT-6 Astra Claude Opus 5.5
DeepSWE v1.1 (long software engineering tasks) 77.9% 74.1% 74.2%
AutomationBench 51.3% 41.4% 42.5%
Vals Finance Agent v2 65.4% 53.5% 58.6%
Harvey Legal Agent 19.6% 5.4% 3.8%
GraphWalks 84.2% 71.8% 66.8%
LVBench (long video) 91.7% 87.5% 83.7%
CWE-bench (vulnerability fixing) 68% (tie) 68% (tie) 67%
FrontierSWE v2 55.0% 65.5% —
Terminal-Bench 4.0 57.4% — 66.4%
PostTrainBench 45.3% — 49.3%

Argon is also reported to resist prompt injection better than its rivals. On Gray Swan's indirect prompt injection test, attacks succeeded 0.7% of the time, against 1.0% for Opus 5.5 and 8.5% for Astra. That matters if you build agents that read untrusted web pages, emails or documents.

The independent view

Artificial Analysis puts Argon at 53 on its Intelligence Index (v4.3.2). That ties GPT-6 Astra and Claude Fable 5.1 and is one point above GPT-6.1 Sol (52). It is still behind Claude Opus 5.5 (58) and Claude Sonnet 5.5 (56). So "back at the frontier" holds up, but "clear leader" doesn't.

The most interesting independent result is on hallucination. On AA-Omniscience, Argon's hallucination rate is 15%, compared with 51% for Astra and 54% for GPT-6.1 Sol. The trade-off is that Argon's raw accuracy on that test is 50%, against 63% for Astra. In practice, Argon is more likely to say "I don't know" than to make something up. For research, legal or support tools where a confident wrong answer is costly, that is often the better failure mode.

Sources disagree on two small points. Engadget reports Astra's hallucination rate as 54%, while Artificial Analysis's own figure is 51%. Terminal-Bench 4.0 shows up as 57.4% in Google's table and 57.1% in Artificial Analysis's run. Neither changes the overall picture.

The internal skepticism

Bloomberg reported that some Google employees think Argon does better on benchmarks than in day-to-day work. They say it "struggles to handle certain coding tasks" and is "not particularly adept at front-end design". Some also believe Anthropic and OpenAI are improving faster. Google disputes this. DeepMind's Koray Kavukcuoglu said "it's a certainty that we are always gonna be at the frontier." Benchmarks back up some of the complaints: Argon trails on FrontierSWE v2 and Terminal-Bench. If your product is mostly UI generation or agentic terminal work, don't assume Argon wins. Test it.

Pricing: the launch rate is cheap, the list price is not

Here are the per-million-token list prices as reported by VentureBeat and Fello AI:

Model Input Output Cached input
Gemini 4 Argon (introductory) $2 $10 $0.10
Gemini 4 Argon (standard) $4 $20 ~$0.20 (95% off; VentureBeat estimate)
Claude Opus 5.5 $4 $20 —
Claude Sonnet 5.5 $2 $10 —
GPT-6 Astra $10 $50 —
Gemini 3.1 Pro $2 $12 —

At the launch rate, Argon costs a fifth of Astra's per-token price and half of Opus 5.5's. Google hasn't said how long the introductory period lasts. Once it ends, Argon costs exactly the same per token as Opus 5.5.

Why per-token price isn't per-task price

Argon uses a lot of tokens. Artificial Analysis measured an average of about 62,000 output tokens per Intelligence Index task for Argon, compared with 27,000 for Astra. That changes the comparison:

  1. At the introductory price, Argon cost about $1.99 per task, against $3.26 for Astra. That's roughly 60% of Astra's cost.
  2. At the standard price, Argon comes to about $3.98 per task, which is more than Astra.

So "a fifth of the price" is true per token but misleading per task, and it only applies during the promotion. If your workload is output-heavy (long code generation, multi-step agents, report writing), budget for the list price and measure real token usage on your own prompts. You can plug your own volumes into the LLM API Cost Calculator to compare monthly spend at both price levels. The Token Counter helps you estimate input size per call.

The 95% cache discount is the other big lever. If your app sends the same long system prompt, codebase snapshot or document set with every request, cached input at $0.10 per million (introductory) makes long-context workloads much cheaper than the headline numbers suggest.

What it changes for people building software

If you're on Gemini 3.x today: Argon's introductory output price ($10) is below the $12 that Fello AI lists for Gemini 3.1 Pro, with similar input pricing. When access opens, it may be a cheap upgrade for a while. Recheck your cost model before the promotion ends.

If you build long-output products: a 1 million token output limit removes the need for some chunk-and-stitch workarounds: full-repo refactors, large migrations, book-length drafts, big structured data exports. Long outputs still cost money and time, so cap max_tokens on purpose rather than by default.

If you build agents that touch untrusted content: the Gray Swan prompt injection result and the low hallucination rate are the most practically useful numbers in this launch. Both are Google-reported or single-benchmark results, so confirm them on your own adversarial test set.

If you build security tools: Fairwind is the only way in today. Google describes it as a program for vetted defenders, and The Next Web reports it gives access without cyber guardrails. Contact Google about eligibility rather than waiting for general availability.

If your product is front-end or UI generation: this is the weakest area according to both Bloomberg's sources and the benchmarks. Keep your current model as the default and run Argon side by side when you can.

How to prepare before the API opens

  1. Build a model-agnostic eval set now. Collect 50–200 real prompts from your product with expected outputs or grading rubrics. When Argon reaches paid API customers, you can score it within a day instead of going by benchmark headlines.
  2. Log output tokens per task, not just per-token prices. Argon's token use per task is more than twice Astra's on Artificial Analysis's suite. Your own ratio decides whether it actually saves money.
  3. Model two cost scenarios. Use $2/$10 for the launch period and $4/$20 for afterwards. If the product only works at the launch price, that's a margin risk.
  4. Put a routing layer in front of your models. The benchmark split (Argon for long reasoning, finance, legal and video; Opus 5.5 or Astra for terminal and some SWE tasks) suggests routing by task type instead of picking one model.
  5. Restructure prompts for caching. Put stable content (system prompt, docs, code context) first so the 95% cached-input discount applies.
  6. Watch for "abstains" in your evals. A model with a 15% hallucination rate will say "I don't know" more often. Decide whether your UX handles that well. For many tools it's better than a confident wrong answer, but users have to be told what happened.

If you're comparing coding models for your stack, the AI Coding Assistants category lists tools built on top of these models. Many of them will add Argon as soon as the API opens.

What to watch next

  • The general API date. Google has only said "as soon as possible" for paid API customers and AI Ultra subscribers. Until then, Argon isn't an option for most production apps.
  • How long the introductory price lasts. The per-task cost advantage over Astra depends on it.
  • Smaller Gemini 4 models. Argon is described as very large, which usually means pricing pressure. A cheaper Gemini 4 tier would matter more to most startups than the flagship.
  • Independent coding evals once access widens. The employee doubts about coding and front-end work will either hold up or fall apart as soon as developers outside the Fairwind Program can test the model.

For now, write your evals, set up your cost model and routing layer, and wait for real-world numbers before switching.

Frequently asked questions

How much does Gemini 4 Argon cost?

At launch, Gemini 4 Argon costs $2 per million input tokens and $10 per million output tokens, and cached input is 95% off. After an introductory period whose length Google hasn't announced, the price doubles to $4 input and $20 output per million tokens. That list price matches Claude Opus 5.5.

Can I use Gemini 4 Argon right now?

Only if your organization is in Google's Fairwind Program for vetted cyber defenders. Google says paid API customers and Google AI Ultra subscribers come next, "as soon as possible", but it hasn't given a date. Everyone else gets access after more testing.

Is Gemini 4 Argon better than GPT-6 Astra and Claude Opus 5.5?

It depends on the task. Google's own numbers show Argon leading on DeepSWE v1.1, AutomationBench, finance and legal agent tests, and long-video understanding. It trails Astra on FrontierSWE v2 and Opus 5.5 on Terminal-Bench 4.0. On Artificial Analysis's Intelligence Index it ties Astra at 53, and Claude Opus 5.5 scores higher at 58.

What is the Gemini 4 Argon context window and output limit?

Google says Argon can produce up to 1 million output tokens, up from the previous 64,000-token limit. Artificial Analysis lists a 1 million token context window. Google's announcement itself focuses on the output limit.

Is Gemini 4 Argon actually cheaper per task than GPT-6 Astra?

Only at the introductory price. Artificial Analysis measured about $1.99 per Intelligence Index task for Argon at the discounted rate, against $3.26 for Astra. At list price Argon comes to about $3.98, because on average it generated roughly 62,000 output tokens per task, compared with 27,000 for Astra.

Sources

  1. Google: Gemini 4 Argon — our next era of frontier intelligence— blog.google
  2. VentureBeat: Google unveils Gemini 4 Argon, retaking benchmark lead — but in limited release— venturebeat.com
  3. 9to5Google: Google announces Gemini 4 Argon as its new frontier model— 9to5google.com
  4. The Next Web: Gemini 4 Argon reaches cyber defenders first— thenextweb.com
  5. Engadget: Google's first Gemini 4 model is 'Argon'— engadget.com
  6. OfficeChai: Gemini 4 Argon ties with GPT-6 Astra on Artificial Analysis Intelligence Index— officechai.com
  7. Bloomberg via Yahoo Finance: Google grapples with employee skepticism about new Gemini model— finance.yahoo.com
  8. Fello AI: Gemini 4 Argon — benchmarks, price and who gets it— felloai.com
person

LaunchBoosts Research Desk

AI-assisted research

Explainers on software and AI industry trends, drafted with AI assistance from the public sources cited in each article and published after automated editorial checks for length, independent sourcing and originality.