Google launches Gemini 4 Argon, its first frontier model in seven months, gated to cyber defenders first
Gemini 4 Argon leads on cost per task and output window. Independent scores sit well below Google's own. Access is rationed through a cyber-defender trust programme, and there is no public API date.
What happened
Summary of reporting by Google (blog.google)On 30 September Google announced Gemini 4 Argon, the anchor model of its Gemini 4 generation and its first new frontier model since Gemini 3.1 Pro in February. Google targets it at software engineering, enterprise knowledge work in legal and finance, creative work and cyber defence. The specs: output ceiling of 1M tokens (up from 64K), up to 128K thinking tokens. Introductory API pricing runs at $2 per 1M input and $10 per 1M output, stepping to $4/$20 later. Cached input is $0.10 per 1M.
The rollout is phased. Fairwind Program members get first access. More than 650 organisations qualify, including CrowdStrike, Palo Alto Networks, Wiz, Mandiant and the US Air Force. Safety restrictions normally applied to the model are lifted for these trusted defenders. Enterprise, Vertex AI and AI Ultra access is promised for "the coming weeks" with no firm date. Google is also taking part in the US government's voluntary pre-release access process, as covered by TechCrunch and Ars Technica.
“Built to sustain deep reasoning across complex, long-horizon workflows, Argon is fundamentally changing the way we work and build at Google.”
Koray Kavukcuoglu, SVP Google DeepMind and Chief AI Architect, Google
Independent measurement is more modest than the launch numbers. Artificial Analysis gives Argon an Intelligence Index of 53, level with GPT-6 Astra (max) but at $0.74 cost per completed task against Astra's $1.21. It sits below Claude Opus 5.5 at 57. Google reports 76.2% on SWE-Bench Pro; third-party runs measure 62%. Hal Bench v2 puts its hallucination rate at 15%, well under Astra's 51%. The Decoder concludes it closes the gap without taking a clear lead. Bloomberg reported that some Google staff are unimpressed with its real-world coding. Google disputes that.
The Azrty take
Argon is cheap per finished task and gated like a weapon: GCC teams should buy routing and evaluation now, not commitment to a model.
For a CTO in Dubai or Riyadh, the number worth quoting is $0.74. Artificial Analysis measured that as cost per completed task for Gemini 4 Argon (high). GPT-6 Astra (max) runs $1.21 at the same Intelligence Index of 53. Claude Opus 5.5 is $2.70. Google's pricing sits at $2 per 1M input and $10 per 1M output during the introductory period, then $4/$20. Cached input is $0.10 per 1M. Add the 1M-token output ceiling, up from 64K and roughly eight times the 128K that Astra and Opus 5.5 offer. Whole-repo migrations become cheaper per finished job than anything currently public. The catch is access. Fairwind members first. Enterprise and Vertex AI "in the coming weeks." No date. Plan roadmap and procurement around access windows, not launch days.
Treat vendor benchmarks as unaudited. Google reports 76.2% on SWE-Bench Pro, 61.6% on Terminal-Bench 2 and 68% on CWE-Bench for vulnerability remediation. Those run on a proprietary scaffold. Independent figures come in at 62% and 64% on the first two. The Decoder shows Argon trailing Claude Opus 5.5 on Humanity's Last Exam (49.2% against Astra's 57.3%, and Astra itself behind Opus on GAP-40) and on ARC-AGI-2 (22.1% against 46.1%). The 14-point gap between Google's SWE-Bench Pro score and independent measurement is larger than most gaps between competing models. Its 15% hallucination rate on Hal Bench v2 is a real improvement on Astra's 51%. It is not, however, a licence to let agents publish legal or financial output unverified. Bloomberg's report of internal scepticism about real-world coding fits the pattern: strong scores, uneven feel.
The most important precedent here is not a benchmark. Google is shipping a frontier model that is capability-gated to vetted cyber defenders. Its own model card states the unrestricted build shows a sharp uplift in dangerous cyber capability. That is why ASL-3 procedures were applied, and guardrails are removed only for trusted members of the Fairwind Program. The list includes CrowdStrike, Palo Alto Networks, Wiz, Mandiant, the US Air Force, Deutsche Telekom and 650+ others. The DeepMind blog also commits Google to STIX as its preferred threat-intelligence standard and 1,000+ Microsoft 365 hardening configurations. For GCC telecoms, banks and critical infrastructure the message cuts two ways. Top-end capability will increasingly be rationed through trust programmes that favour US-nexus vetting. Model behaviour is now negotiated per customer rather than uniform. Document your identity, logging and security posture now. Enter through your MSSP's membership where you have one.
How we would approach it: build against a route, not a model. Put an OpenAI-compatible gateway such as our FastLLM Proxy in front of every provider. Pin exact model versions per workload. Enforce a dollar budget per completed task, not per token. Keep a fallback to Opus 5.5 or Astra until Argon is GA on Vertex AI and you have re-run your own evaluation. Cap output tokens well below the 1M ceiling on first runs. Long agentic output compounds error as well as cost. A single runaway run now costs up to 15 times the old 64K limit.
# Policy sketch for an OpenAI-compatible gateway (FastLLM Proxy or equivalent)
routes:
repo-migration:
primary: $ARGON_ENDPOINT # pin the exact model version once GA
fallback: $OPUS55_ENDPOINT
max_output_tokens: 200000 # do not start at 1M
budget_usd_per_task: 1.00 # measured per completed task
gate: "cost_per_task <= 1.21" # beat Astra (max) or keep the fallback
legal-summary:
primary: $OPUS55_ENDPOINT
max_output_tokens: 32000
require_verification: true # 15% hallucination is not zero
budget_usd_per_task: 3.00
Most teams will get this wrong in one of two directions. They wait for general availability and lose three months of harness work. Or they rewrite prompts and agents around Argon's formatting before the API is public at all. The compounding advantage sits with organisations that treat the frontier as a rotating set of interchangeable backends. They invest now in evaluation, verification and gateway policy. In the GCC, where AI spend often runs through annual budget cycles and a single vendor relationship, that flexibility is also your negotiating position. Once cost per completed task is measured on your own code, you know exactly what a $0.74 migration run is worth against a $2.70 one. You know which model deserves the route.
What to do now
- Stand up an OpenAI-compatible gateway (FastLLM Proxy or equivalent) this month. Pin exact model versions per route. Set a USD budget per completed task. Configure Claude Opus 5.5 and GPT-6 Astra as fallbacks until Gemini 4 Argon is GA on Vertex AI.
- Build a one-workload evaluation harness on your own repositories. Use SWE-Bench Pro style or the Artificial Analysis cost-per-task method. Run it so you can compare cost per completed task against whatever is in production today, before you switch anything.
- If you run critical infrastructure or financial services workloads, request Fairwind Program access via Google or your MSSP (CrowdStrike, Palo Alto Networks, Mandiant and CyberCX are members). Register on the Vertex AI waitlist. Assume no public API date this quarter.
- Set spend controls for the pricing step ($2/$10 intro, $4/$20 after). Make prompts cache-aware; cached input is $0.10 per 1M. Cap max_output_tokens at a fraction of the 1M ceiling until you have measured error accumulation on long runs.
