xAI ships Grok 4.6, targeting long-running coding agents and one-pass app builds
xAI has released Grok 4.6, an incremental update to Grok 4.5 aimed squarely at agentic workloads: multi-step tasks like codebase work, research, and turning a rough product idea into a working application. The company says the model can hold context across many steps and, on longer runs, has begun to self-test and verify its own output before moving on. On the Artificial Analysis Intelligence Index—a composite of nine benchmarks—xAI claims parity with OpenAI’s GPT-5.6 Sol.
The gains come from a longer supplemental pretraining run over curated reasoning and engineering data with a revised optimizer, followed by SFT trajectories regenerated by Grok 4.5 and filtered with model-based checks, then reinforcement learning across domains including kernel optimization, web dev, and CAD. xAI highlights stronger first passes on visual and interactive projects, positioning the model for a build-something-substantial-then-iterate workflow rather than incremental prompting.
Grok 4.6 is available now in Cursor and Grok Build, plus the API and partners including OpenRouter, Vercel, and Cloudflare. Pricing is $2 per million input tokens and $6 per million output, with a fast variant at double the rate; xAI is offering 2x included usage in Grok Build and Cursor for the first week. The company frames its safety work around legitimate uses such as vulnerability patching and AI research, claiming its widest pre-deployment safeguard testing to date—though, as with most vendor launches, the benchmark and safety claims are self-reported.
Read the full article
Continue reading at Hacker News →This is an AI-generated summary. Read the original for the full story.