Gemini 4 Argon Is Google Back at Frontier
Argon launched Sep 30 with 77.9% DeepSWE and 68% CWE-bench. Why Google limits it to Fairwind defenders first, plus pricing and builder prep while you wait.
On this page
Google claims it is back at the frontier, but you cannot try the proof yet. Gemini 4 Argon launched September 30 with coding and cyber scores plus a 1 million token output limit, access stays with Fairwind defenders during US pre-release review. Here is what the benchmarks show, why the rollout is phased, and how to plan while you wait.
Takeaways
Scores cited: 77.9 percent DeepSWE v1.1 above Astra, Fable 5.1, and Opus 5.5, plus 68 percent CWE-bench v1 tied first. Output size: up to 1 million tokens in one step, up from 64,000, aimed at migrations and long analyses. Access: Fairwind cyber partners only, no public date, while Google scales guards for misuse and prompt injection. Price signal: $2 in and $10 out with 95 percent off cached input, matching Sol class sticker while you compare per task cost.
What do the benchmarks actually show?
Google positions Argon as a well rounded long horizon worker: real world software engineering, enterprise knowledge work across legal and finance, multimodal parsing of video and charts, plus autonomous find, validate, and patch for vulns. Internal use includes data center memory optimization that freed hundreds of terabytes without new hardware, plus quantum research help. Wiz reportedly used Argon to find a critical hospital system flaw that other frontier models missed, though Google shared no specifics to verify that claim independently.
code: 77.9 DeepSWE v1.1, above Astra, Fable 5.1, Opus 5.5 per Google
cyber: 68 CWE-bench v1 tied first, built on 3.8 Flash Cyber base
work: Vals Index lead on finance and legal style economic tasks
limit: 1M output tokens in one step for migrations and auditsWhere Argon trails its own table
Reuters notes Argon stayed behind on two of four coding benchmarks Google included, and Google dropped the planned Gemini 3.5 Pro after DeepMind changes. Treat the release as a return to frontier breadth, not a clean sweep, and wait for independent runs before moving pinned IDs.
How should builders handle a phased model?
Keep current pins, add Argon as shadow
Leave Luna for volume, Sol for professional coding, and Opus 5.5 for long agents in place. Draft Argon prompts and evals now so you can A/B on day one of wider access.
Design for 1M output early
Chunk audits and migrations so they also run on 64K models today, with a flag to emit full reports when long output arrives. Long output helps only if review and diff tooling keep pace.
Copy the defender posture
Trusted testers get Argon without cyber guards for defense work, while the public build adds monitoring. Mirror that split: full tool power in isolated sandboxes, guarded subset in prod.
Track chain of thought policy
Google says Argon monitors reasoning and can halt out of bounds chains. Log your own traces too, since halt behavior becomes a debug surface the moment evals fail without visible cause.
Is phased release now the norm?
Yes for this capability band. OpenAI briefed similar pre-release access talk around Astra, Anthropic weighs safety evals before any pre IPO model, and Google ties Argon expansion to feedback plus guard iteration. The pattern favors defenders first, then enterprises, then consumers. Budget for that lag in roadmaps that assume day one API access.
Can Google hold the lead it claims?
Does DeepSWE 77.9 settle the model race?
No. It is a vendor reported score on one harness, and Google concedes gaps elsewhere in its own table. The durable signal is breadth at lower per task cost across code, knowledge work, and defense, not a single peak number.
How does this change the September chooser?
It does not change pins yet because you cannot call Argon in prod. Keep the September routing of Luna for bulk, Sol for professional builds, and Opus 5.5 for sprawling runs, then re run the same 10 case suite per workload once Argon opens.
As of October 2, 2026: Argon reads as Google rejoining frontier breadth with a safety first rollout, not a launch you can ship on. Prep evals, chunk for long output, and isolate tool power. Next, read which model to use to hold your routing until independent Argon runs land.