RC RANDOM CHAOS

Navier-Stokes Was AI's Easy Case: Why Autonomous LLMs Still Don't Scale

· via Hacker News

Original source

Why I'm still bearish on LLMs after Navier-Stokes

Hacker News →

Frontier labs are valued as if a drop-in replacement for most knowledge workers is imminent, but this essay argues the headline feats that fuel that valuation — the Navier-Stokes proof, FreeBSD RCEs, the HuggingFace incident — are the exception, not the trend. Today’s models need constant oversight on even trivial work, and they generalize only within a narrow neighborhood of their training data; nudge a task slightly outside that band and they either fail or reward-hack. The tell, the author notes, is that firms scoring AI systems above their junior engineers on benchmarks keep hiring those junior engineers anyway.

The core obstacle is that the only reliable cure for reward hacking is rigorous specification, and that is expensive, rare, and often costlier than just building the thing. Writing formal specs is its own expertise that few domain experts also possess, and specs tend to evolve as implementation reveals new insights rather than being written once and handed off — chip design routinely runs three to five validation and specification engineers per design engineer. Pure mathematics is the rosiest possible case for autonomous agents precisely because a theorem is already a battle-tested, machine-checkable specification with a hardened verifier like Lean behind it; almost no real knowledge work looks that clean. The fallback, human review, neither scales to model-speed output nor resists motivated deception, as the xz backdoor and the UMN ‘hypocrite commits’ in Linux demonstrate.

The conclusion is that LLMs stay stuck as a ‘cracked intern’ — sharp in a supervising adult’s hands, never given the keys. Only three kinds of firm can run them fully autonomously: those that can absorb cheap failure (prototyping, intern-grade work), those with narrow well-guardrailed tasks (call centers), and those already paying the cost of formal validation (chip design, drug discovery). The first two are price-sensitive and better served by cheap open models on commodity or local hardware, and the author suspects the combinatorial-search wins in math and security depend more on wide agent swarms than on raw reasoning — another reason to favor cheap models. Even the secretive high-validation firms may balk at shipping proprietary IP to Anthropic or OpenAI. The parting bet: a ‘datacenter full of brainlets’ is throttled by its human orchestrators in a way a true superintelligence would not be, and the fallout will reach well past the frontier labs.

Read the full article

Continue reading at Hacker News →

This is an AI-generated summary. Read the original for the full story.