Strong on Paper, Short on Access: What We Know About Google’s Gemini 4 Argon

Published:

Most new AI models arrive with a chart showing they beat everyone else. Gemini 4 Argon arrived with a chart and a velvet rope.

On 30 September 2026, Google announced Gemini 4 Argon, which it calls “our next era of frontier intelligence” [1]. The announcement came with record benchmark scores, a far bigger output limit and a low starting price. But the first people to get it are a small group of cybersecurity defenders, not the public. And a few days later, Bloomberg reported that some Google staff, speaking anonymously, doubt it codes as well in real work as it does on tests [2]. Google disputes that.

So there are two stories here: what Google says Argon can do, and how much of that anyone outside Google can check yet.

What did Google announce?

The announcement was a blog post by Koray Kavukcuoglu, SVP of Google DeepMind and Google’s Chief AI Architect [1]. In it, Google says:

  • Argon is a new frontier model trained with a focus on defensive cybersecurity.
  • It’s rolling out first to trusted cyber defenders through Google’s Fairwind Program.
  • Wider access will come “as soon as possible,” starting with paid API customers and Google AI Ultra subscribers, after US voluntary pre-release engagement and more guardrail work.

Google didn’t give a date for that wider rollout [1].

There’s one detail worth slowing down on. Google says Fairwind defenders and its own internal teams get Argon without its cyber guardrails [1]. The idea is that people defending systems need a model that will dig into security flaws without refusing. It also means the most capable version is going to a small, vetted group first.

What does it cost?

These are Google’s announced API prices [1]:

  • Introductory price: $2 per million input tokens and $10 per million output tokens.
  • Cached input: 95% off the input price.
  • After the introductory period: $4 per million input and $20 per million output.

A token is a small chunk of text, often part of a word. We didn’t find an end date for the introductory price.

For comparison, $2 in and $10 out is the same headline rate OpenAI lists for its new GPT-6.1 Sol model [3]. That’s a comparison of price lists, not of quality. Argon’s price is set to double once the introductory period ends.

What does Google say it can do?

Google’s post lists benchmark scores. All of these are Google’s own claims, and we didn’t open any third-party checks of them [1]:

  • DeepSWE v1.1: 77.9%, which Google calls state of the art
  • AutomationBench: 51.3%, which Google says ranks first
  • LVBench: 91.7%, which Google calls state of the art
  • CWE-bench v1: 68%, which Google says ties for first
  • Vals Index: Google says Argon is “leading”

Google also says Argon’s output limit has grown to 1 million tokens, up from 64,000 [1]. Put simply, the model can write a reply that’s book-length or longer in one go, rather than stopping after a long chapter.

The post also has internal success stories, which again are Google’s claims [1]. Google says Argon helped optimise a quantum computing subroutine that “beat published baseline by 40%,” found memory savings across Google’s computer fleet “freeing up over 300 TiB,” and helped move C and C++ code to Rust, including a rewrite of part of the libgav1 video decoder that Google says runs 2.7 times faster than an earlier Rust version. Google also names Wiz’s “Scan for Good” project as an early user and says it turned up a critical flaw in healthcare software.

Why are some people at Google skeptical?

On 4 October 2026, NWA Online ran a reprint of a Bloomberg story by Julia Love and Davey Alba [2]. Most of its critical claims come from people who asked not to be named. Here’s what they said, according to Bloomberg:

  • Gemini 4 has done well on benchmarks but does less well when employees put it to work, and struggles with certain coding tasks.
  • One person called its front-end design work uneven.
  • Two people said it seems affected by “benchmaxxing,” meaning it’s been tuned to score well on tests rather than to do real jobs well.
  • Google had planned to release Gemini 3.5 Pro in June, after announcing it at its I/O event in May, but dropped that effort.

The same report shows disagreement inside Google. Some employees think Anthropic’s Fable and OpenAI’s Astra are improving faster. Others think the coming version has caught up. One Google employee familiar with the work said there’s “large consensus” internally that Gemini 4 is at the frontier and denied that it struggles with messy real-world coding [2].

Google’s official answer, as reported: “Google said it would be inaccurate to say that Gemini 4 is underperforming in areas such as coding.” The company pointed Bloomberg to Kavukcuoglu’s earlier comments that he was encouraged by the model’s performance [2].

One more thing to be careful about: the report talks about “Gemini 4.” From the reprint, it isn’t clear which internal version the employees were describing, or whether it was exactly the Argon model announced on 30 September.

Why the stakes are high

A frontier model is a costly bet. A Bloomberg Intelligence analyst, Mandeep Singh, estimated that a single training run can cost as much as $400 million [2]. That’s an analyst’s estimate, not a Google figure.

Google also has a lot riding on Gemini across its products. According to Bloomberg, Google says its consumer Gemini app and AI Mode in Search have each passed 1 billion users [2]. That’s a company claim.

When this much rides on a model, benchmark scores turn into a marketing battle, and that’s why the gap between test results and day-to-day use matters.

What this does not prove

  • That Argon is the best model. The benchmark scores are Google’s own. We didn’t see independent tests.
  • That Argon is bad at coding. The doubts come from anonymous sources, Google denies them, and other insiders disagree.
  • That you can use it soon. Google gave no date for wider access beyond Fairwind defenders.
  • That the starting price will last. Google has already said it will rise to $4 in and $20 out.
  • The full Bloomberg story. We read a reprint, not the original on Bloomberg’s site, and we didn’t open the source of Kavukcuoglu’s earlier comments.

The Bottom Line

On paper, Gemini 4 Argon is Google’s strongest pitch yet: top-tier benchmark claims, a huge output limit and a low starting price. In practice, almost nobody outside Google and its vetted cyber defenders can use it, so almost nobody can check those claims. Bloomberg’s anonymous sources say the benchmarks flatter it on real coding, and Google says that’s wrong. The fair verdict for now is “unproven,” not “overhyped” and not “best in class.” The real test starts when paying developers get access.

Sources

  1. Koray Kavukcuoglu, “Gemini 4 Argon: our next era of frontier intelligence,” Google blog, 30 September 2026. https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/
  2. Julia Love and Davey Alba, Bloomberg, reprinted as “Google grapples with Gemini skepticism,” NWA Online, 4 October 2026. Critical claims are from unnamed sources. https://www.nwaonline.com/news/2026/oct/04/google-grapples-with-gemini-skepticism/
  3. OpenAI API documentation, GPT-6.1 Sol model page (pricing, used only for the price comparison). https://developers.openai.com/api/docs/models/gpt-6.1-sol
TSN
TSNhttps://tsnmedia.org/
Welcome to TSN. I'm a data analyst who spent two decades mastering traditional analytics—then went all-in on AI. Here you'll find practical implementation guides, career transition advice, and the news that actually matters for deploying AI in enterprise. No hype. Just what works.

Related articles

Recent articles