RC RANDOM CHAOS

DeepSeek V4 Flash posts 89% on ARC-AGI-1 at two cents a task

· via Hacker News

Original source

DeepSeek V4 Flash 0731

Hacker News →

DeepSeek’s V4 Flash 0731, released July 31, 2026, has landed verified scores on the ARC-AGI reasoning benchmarks. Running at maximum reasoning effort, the model reaches 89.0% on the ARC-AGI-1 semi-private set for roughly $0.02 per task, and 61.4% on the harder ARC-AGI-2 set for about $0.04 per task. The model ships with three selectable reasoning variants, letting users trade compute cost against accuracy.

The headline story is price-performance rather than raw capability. Near-90% on ARC-AGI-1 at two cents a task puts frontier-level abstract reasoning within reach of high-volume, cost-sensitive workloads, and the sub-50-cent-equivalent economics undercut the pricing typically associated with top-tier reasoning models. The steep gap between the two benchmarks — 89% versus 61% — also underscores that ARC-AGI-2 remains a meaningfully tougher test of generalization, where even strong models have clear headroom.

Because the results are listed as verified on the ARC Prize leaderboard, they carry more weight than self-reported vendor numbers, which have a history of being difficult to reproduce. For teams evaluating reasoning models, the takeaway is that DeepSeek is competing aggressively on the cost-per-task axis while staying competitive on the accuracy axis that ARC-AGI is designed to stress.

Read the full article

Continue reading at Hacker News →

This is an AI-generated summary. Read the original for the full story.