RC RANDOM CHAOS

Google ships Gemini 3.8 Flash, plus a defender-only 'Cyber' variant for bug fixing

· via Hacker News

Original source

Gemini 3.8 Flash and 3.8 Flash Cyber

Hacker News →

Google has released Gemini 3.8 Flash, the third Flash model in six weeks, positioning it as a low-cost workhorse that narrows the gap with pricier frontier models on coding and multi-step reasoning. Pricing holds at 3.7 Flash’s introductory rates of $0.75 per million input tokens and $3.75 per million output tokens. Google credits the gains to a model that simply ‘works harder’—running extra reasoning steps and calling tools iteratively—which also means it can burn more tokens at higher effort settings. For cost-sensitive workloads, Google is keeping 3.7 Flash around and lets developers dial effort down. The company cites its own numbers on benchmarks like DeepSWE v1.1, Harvey’s legal agent test, and a 54.9% score on HLE-Verified, though these are vendor-reported results rather than independent evaluations.

The more notable release is Gemini 3.8 Flash Cyber, a security-tuned variant gated behind a new ‘Fairwind Program’ for vetted defenders—government bodies, critical-infrastructure operators, and software maintainers. Google says it deliberately optimized for vulnerability discovery and automated patching over offensive exploitation, reporting a 70%-plus success rate on an internal benchmark spanning 20 languages and roughly matching a leading model on the CWE-Bench patching test (47.2% vs. 47.8% pass@1) at lower cost. Internal case studies claim Chrome’s security team got 2.6x more correct patches than larger commercial models, and that Cloud vulnerability researchers surfaced a critical flaw in under two hours.

The gated distribution is the interesting policy signal here: rather than ship strong autonomous vuln-finding to everyone, Google is restricting the more capable, permissively-mitigated cyber model to trusted defenders, framing it as tilting the balance toward defense under its Frontier Safety Framework. The standard 3.8 Flash ships with CBRN and cyber-offense safeguards and claims improved prompt-injection resistance per Gray Swan testing. As always with self-published launch benchmarks, the real-world defensive value—and whether access controls meaningfully limit dual-use—will need outside verification.

Read the full article

Continue reading at Hacker News →

This is an AI-generated summary. Read the original for the full story.