RC RANDOM CHAOS

xAI Ships Grok 4.7, Leaning on Coding Muscle and a Rebuilt Safety Stack

· via Hacker News

Original source

Grok 4.7

Hacker News →

xAI has released Grok 4.7, positioning it as its strongest model yet for coding and knowledge work. Built on a larger base model than Grok 4.6 and trained with a longer reinforcement-learning run weighted toward multi-hour tasks, it is tuned to persist on hard problems, verify its own output, and hold longer context. xAI claims frontier-level price-performance on the CursorBench 4.0 coding benchmark and gains on professional-work benchmarks like GDPval and AA Briefcase, while keeping the same price and speed as its predecessor at $2 per million input tokens and $6 per million output.

The more notable pitch is safety. Grok 4.7 ships with what xAI calls an entirely new safeguard stack, which it says leads its own testing on refusals and jailbreak resistance. The company reports it tops LatchBio’s biosafety benchmark at 62.4% and allows only 3.3% of risky dual-use prompts through on its internal HackerBench v0.3 while rarely blocking legitimate security work — the familiar dual-use balancing act of staying useful for defenders without arming attackers. Notably, xAI is handing select cybersecurity partners invite-only access to the model’s red-team capabilities for defense research.

The security framing deserves a skeptical read: nearly all the cited numbers come from xAI’s own benchmarks, and a 3.3% pass-through rate on malicious cyber prompts is still a nonzero channel for abuse at scale. The model is available now through Cursor, Grok Build, the Grok API, third-party coding harnesses, and cloud routers, with a fast variant offered at double the speed and double the price.

Read the full article

Continue reading at Hacker News →

This is an AI-generated summary. Read the original for the full story.