RC RANDOM CHAOS

Cheap, fast AI models hit good-enough — and rewrite the economics of consumer apps

· via Hacker News

Original source

Small Models Have Arrived

Hacker News →

The author argues that small, inexpensive language models have quietly crossed a usefulness threshold. Running a model like the fictional gpt-5.6-luna, he clocks ~100 tokens/sec and struggles to spend more than tens of cents even on research that spans thousands of emails. That price collapse matters because inference cost has been the hidden tax blocking AI-native consumer products: the classic playbook of cheap website, viral growth, then ad monetization breaks when every request carries real compute cost. His benchmark — auto-generating a personalized daily news site — cost roughly $1 per run on prior Sonnet-class models, untenable at consumer price points, but drops to about $0.10 with the newer small models.

The more interesting shift, he contends, is in business work. Borrowing a distinction from his Segment co-founder Peter, he splits labor into rare “IQ 180” breakthrough work and high-volume “token spewer” work — being responsive, nudging people, moving many balls forward. Even a highly effective founder running multiple companies estimates ~95% of his time falls into the second bucket, which mirrors how most hiring skews toward fast, responsive, good-enough contributors rather than lone geniuses.

The takeaway: demand for frontier models will keep compounding for genuine discovery work (hard science, engineering, model training), but a much larger market is opening for cheap-fast-good-enough models that handle routine throughput. Realizing it will require new infrastructure — agent harnesses, prompt-injection defenses, and proper roles and permissions — which the author expects the industry to build out.

Read the full article

Continue reading at Hacker News →

This is an AI-generated summary. Read the original for the full story.