RC RANDOM CHAOS

DeepSeek ships V4-Flash public beta, leaning hard into coding agents

· via Hacker News

Original source

DeepSeek-V4-Flash Update

Hacker News →

DeepSeek has moved its V4-Flash model into public beta as an official release, accessible through the same API by pointing the model parameter at deepseek-v4-flash. The refreshed build, tagged 0731, is not a new architecture — it reuses the size and design of the earlier V4-Flash-Preview and was simply re-post-trained, so the changes show up in behavior and benchmark scores rather than in how developers call it.

The published numbers make the priorities clear: this is a play for agentic and software-engineering workloads. DeepSeek reports 82.7 on Terminal Bench 2.1, 76.7 on Cybergym, and mid-to-high 50s on repo-scale and full-stack coding tests like NL2Repo (54.2), DeepSWE (54.4), and the internal DSBench-Hard (59.6), while broader autonomous-agent benchmarks such as Agent Last Exam (25.2) remain low. Notably, the coding-agent results were produced using DeepSeek’s own not-yet-released ‘Harness’ framework at maximum effort, temperature 1.0 — a self-selected test rig that’s worth weighing before treating the scores as apples-to-apples.

The update touches only the Flash API; V4-Pro and the app/web models are unchanged, with a Pro release promised soon. It lands alongside DeepSeek’s broader V4 rollout, which added both OpenAI ChatCompletions and Anthropic-compatible interfaces and is retiring the legacy deepseek-chat and deepseek-reasoner endpoints, signaling a push to slot cleanly into existing agent tooling.

Read the full article

Continue reading at Hacker News →

This is an AI-generated summary. Read the original for the full story.