RC RANDOM CHAOS

Cerebras and OpenAI Preview Ultrafast: GPT-5.6 Sol at 750 Tokens/Sec

· via Hacker News

Original source

Accelerating GPT-5.6 Sol Ultrafast

Hacker News →

Cerebras and OpenAI are previewing Ultrafast Mode, a new tier in the OpenAI API that runs the GPT-5.6 Sol model on Cerebras hardware at up to 750 output tokens per second. The companies claim the speedup comes with no drop in quality, targeting latency-sensitive work where a frontier model’s response time is the bottleneck. It’s launching as a limited preview to a handful of customers, with access widening as capacity grows.

The pitch is that Ultrafast collapses the usual speed-versus-intelligence tradeoff. Cerebras’s own benchmarks report that Sol Ultrafast worked through all 2,500 questions of Humanity’s Last Exam in about 11 hours versus roughly 78 hours for Claude Fable 5 — a ~7x edge at comparable accuracy — and a 5.6x end-to-end speedup on the GDP-Val knowledge-work benchmark. These are vendor-run comparisons, so the figures warrant the usual skepticism, but the intended applications are concrete: putting agents on the critical path for production incident response, SLA-bound outage triage, and time-pressured security operations like detecting and containing active attacks.

The performance rests on Cerebras’s Wafer-Scale Engine, which packs 44 GB of SRAM onto a single wafer-sized chip. Large-model inference is fundamentally a data-movement problem: on GPUs, weights are shuttled between on-chip and off-chip memory for every token, and memory bandwidth caps throughput. By keeping weights resident on-chip and pipelining tokens across wafers, Cerebras sidesteps that bottleneck — an architecture it argues will scale its speed advantage to larger future models.

Read the full article

Continue reading at Hacker News →

This is an AI-generated summary. Read the original for the full story.