NVIDIA's Vera CPU Is Genuinely Fast — Its Whitepaper Oversells It
NVIDIA’s first in-house server CPU, Vera, pairs 88 of the company’s new 10-wide Arm v9.2 ‘Olympus’ cores on a monolithic die with 164 MB of shared cache and eight LPDDR5X channels good for 1.2 TB/s at roughly 50 watts. The core itself looks formidable: a wide out-of-order design with a neural branch predictor, value prediction, a graph prefetcher, and six SVE pipes. Early independent numbers back this up — Phoronix’s May run put a pre-production Vera about 10% ahead of a 5 GHz EPYC 9575F, 1.55x a Xeon 6980P, and 1.63x NVIDIA’s own Grace, making it the fastest Arm server chip yet seen in public, albeit on a workload set NVIDIA hand-picked with power and frequency monitoring locked out.
The problem is the 45-page paper wrapped around the silicon, which repeatedly inflates ordinary engineering into an anti-x86 argument. Chips and Cheese catalogs the sleights of hand: SMT is drawn as crude time-slicing, a configurable NUMA layout is framed as an inescapable 32-node maze, four SPEC subtests are rebranded ‘agentic benchmarks,’ undefined performance-counter ratios are passed off as causal proof, and an unlabeled pictogram is presented as a 1.8x reinforcement-learning win. Several ‘novel’ features aren’t — Intel has shipped data-dependent (graph-style) prefetchers since 2022, and perceptron branch predictors date back to AMD’s 2012 Piledriver before the industry largely moved to TAGE.
The sharpest error is Figure 5’s portrayal of NVIDIA’s ‘Spatial Multithreading’ as fundamentally better than traditional SMT. In reality SMT already shares fetch, decode, execute, and cache stages dynamically, handing idle throughput to whichever thread can use it; NVIDIA’s static partitioning could instead strand execution resources reserved for a stalled thread. Tellingly, the paper defends Spatial Multithreading on determinism, isolation, and quality of service — never on raw performance — which may suit its agent-serving target market fine but undercuts the diagram’s implication of a throughput advantage. The takeaway: Olympus is strong enough that it didn’t need the spin, and the spin is the weakest part of the pitch.
Read the full article
Continue reading at Hacker News →This is an AI-generated summary. Read the original for the full story.