DeepSeek ships V4-Pro and adds 50%-off off-peak API pricing
DeepSeek has moved its V4-Pro model to general availability, positioning it primarily as an agent-focused release with what the company describes as significant production gains. Both V4-Pro and the lighter V4-Flash now expose a tunable reasoning-effort setting—low for trivial calls, high for routine agent workflows, and max for hard problems—giving developers a lever to trade cost against depth per request. The release also ships native support for OpenAI’s Responses API with a one-click setup aimed at Codex, signaling a bid for drop-in compatibility with existing OpenAI-oriented tooling. The model is reachable through the app and web under a new ‘Expert Mode,’ and via API under unchanged model names.
The more consequential change for most users is pricing. Alongside the V4 lineup, DeepSeek is introducing tiered peak and off-peak API rates, with off-peak pricing set 50% below peak to reward workloads that can be scheduled during quieter windows. The new structure takes effect at 16:00 UTC on August 16, 2026.
The move follows a broader industry pattern of time-of-use pricing for inference, letting providers smooth demand while offering batch and non-latency-sensitive workloads a substantial discount. For teams running large agent pipelines, the combination of adjustable reasoning effort and off-peak rates creates real room to cut spend—provided their jobs tolerate deferral and they track the UTC cutover, which lands mid-afternoon UTC rather than at a local midnight.
Read the full article
Continue reading at Hacker News →This is an AI-generated summary. Read the original for the full story.