Skip to main content
Web3Fire
Marketmedium ImpakKeyakinan: medium

Gemini 4 Is Here, and Google’s Flagship Tops All Other AI Models on Cybersecurity

|Decrypt✓|Read original →

Ringkasan

Gemini 4 Argon tops 12 of 18 benchmarks in Google's own table, writes a million tokens per reply and resists hijacking best. Cyber defenders get it first, with the guardrails off.

Gemini 4 is finally here, one week after the release of Claude Opus 5.5 and one day after GPT 6.1 Sol, proving American labs are very much committed to slowing down AI development. Please excuse our sarcasm.

Google unveiled Gemini 4 Argon on Wednesday, calling it its frontier model, meaning its most capable, for coding, office work and cyber defense.

On DeepSWE v1.1, a test of whether an AI can finish long, messy, real-world software engineering jobs, scored as a percentage, Argon hit 77.9%. Claude Opus 5.5 got 74.2%, GPT-6 Astra 74.1% and Claude Fable 5.1 67.4%.

For scale, Gemini 3.6 Flash managed 49% on the same test in July. Argon can also write up to 1 million tokens in one reply, up from 64,000. A token is a chunk of text, roughly three-quarters of a word, so that is about 750,000 words versus about 48,000.

Take these numbers with a grain of salt, though. Google computed its own DeepSWE score, while rivals' numbers came from a public leaderboard and company reports. Its table also concedes ground. Argon leads on 12 of 18 benchmarks, ties one and trails on five, a mix of coding, science and computer-control tests.

But the model’s flashy feature is cyber capabilities. Hide a secret instruction inside an email, wait for an AI assistant to read it, and see if the AI obeys the stranger instead of you. That is indirect prompt injection, and it is the nightmare for anyone who wants to hand an AI their inbox or shopping cart.

On Gray Swan's Indirect Prompt Injection benchmark, which hides malicious instructions in content agents read and scores how often the attacks work within 15 tries, Argon landed at 0.7%. Lower is better. Claude Opus 5.5 and Claude Fable 5.1 both scored 1.0%.

GPT-6 Astra came in at 8.5%. Grok 4.6 and Kimi K3 got tricked just over half the time, at 51.8% and 52.7%.

Argon goes to vetted security teams through the Fairwind Program , Google's limited-access cyber defense initiative, which launched September 2 with more than 650 partners including governments and critical infrastructure operators. And it ships "without cyber guardrails," the built-in refusals that normally stop a model from helping with hacking.

Sep 24 Sep 25 Sep 27 Sep 29 Oct 1 $85.3k $84.4k $83.5k $82.7k 24h High High $85,518 24h Low Low $82,951 Vol Vol $1.6B Market projections Odds by Myriad Today $82,000 to $84,000 $82k–$84k 61 % chance This week Below $84,000 Below $84k 59 % chance → Buy Bitcoin with USDT Powered by Jupiter $ 50 $ 100 $ 500 Buy Price data by CoinGecko CoinGecko More Bitcoin news and projections → The logic behind such a move is that defenders need a model that can think like an attacker to patch holes before criminals find them. The catch is that the same skill cuts both ways, so Google says a phased rollout is the only safe path. It is also taking part in the U.S. government's voluntary process for pre-release model access.

Google isn't the first to put a cyber model behind a velvet rope. An early version of Anthropic's Claude Mythos helped find 271 vulnerabilities in Firefox, meaning 271 security holes Mozilla then patched. OpenAI has taken a similar route with its Trusted Access for Cyber program .

Argon's cyber scores jump over Gemini 3.8 Flash Cyber, the restricted model Google launched with Fairwind. On the Wiz Penetration Test Benchmark, an internal Google test that asks an AI to write working exploits against real web-application flaws without seeing the code, and scores the share solved on the first try, Argon hit 70.9% against 58.2%.

Google also says Argon helped security firm Wiz find a critical flaw in healthcare software used by hospitals worldwide, one earlier frontier models had missed.

The launch follows a rough summer for Google. In July it shipped smaller Flash models but skipped the promised Gemini 3.5 Pro, and Alphabet shares fell about 4.4%. Argon also landed the same day President Trump unveiled a voluntary, penalty-free AI accord that Google's leadership signed.

Google says wider release comes as soon as possible, starting with paid API customers and Google AI Ultra subscribers.

Introductory pricing is $2 per million input tokens and $10 per million output tokens. Google hasn't said when that period ends, only that standard rates are $4 and $20.

Berita Berkaitan