Inception's Mercury 2.5 bets diffusion LLMs can win on latency and cost
Inception Labs has released Mercury 2.5, which it calls the largest diffusion language model ever trained and the most capable diffusion LLM available. Unlike the autoregressive architecture behind most large models, diffusion LLMs generate tokens in parallel, and Inception is leaning on that to compete on speed and price rather than raw frontier intelligence. The company claims a 40% intelligence gain over Mercury 2 — putting it in the same class as cost-optimized tiers like GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite, and Claude Haiku 4.5 — while hitting 1,107 tokens per second on standard NVIDIA GPUs, a 260K-token context window, and pricing of $0.20/$0.75 per million input/output tokens (discounted 80% at launch). Notably, Inception says it tuned this release using production failure cases and customer feedback rather than benchmarks alone.
The pitch is aimed squarely at latency-sensitive, high-fan-out workloads where many small model calls stack up. In search and RAG pipelines, that means query rewriting, reranking, and summarization staying inside a single user interaction. Inception’s customer examples do the heavy lifting: voice-agent firm OpenCall reports median response latency near 170ms and P99 dropping from minutes to about one second, while Augment Code says moving session compaction to Mercury cut latency 82% (from ~150s to 27s) and cost 90%. Alongside 2.5, the company previewed Mercury Voice, tuned for sub-170ms time-to-first-token, and Mercury Router, a diffusion model that dispatches prompts to the best-fit open or closed model.
This is a vendor announcement, so the benchmark comparisons and testimonials should be read with that framing. Still, the broader signal is that diffusion-based LLMs are maturing from research novelty into production infrastructure for coding subagents, voice, and search — categories where token latency compounds and autoregressive models struggle to keep pace. Inception says its next and largest model is already in training, targeting a release in the coming months.
Read the full article
Continue reading at Hacker News →This is an AI-generated summary. Read the original for the full story.