RC RANDOM CHAOS

The Habsburg Frog: How a Deformed-Jaw SVG Became a Quirky LLM Benchmark

· via Hacker News

Original source

My personal AI benchmark: “Generate an SVG of a frog with a Habsburg jaw”

Hacker News →

A developer has turned an absurd prompt — draw an SVG frog with a Habsburg jaw, the exaggerated protruding mandible associated with the inbred Habsburg dynasty — into an informal test of AI models. The site collects each model’s raw SVG output alongside annotations of how it interpreted the request, letting readers compare how different systems handle a task that demands vector-drawing skill, anatomical reasoning, and a grasp of an obscure cultural reference all at once.

The results are revealing in the details. Models like Claude Opus 5 don’t just render green shapes; they encode intent through structure and comments, producing gradients for skin and jaw, bulging eyes, warts, and teeth explicitly labeled as a ‘massive protruding mandible’ with lower teeth ‘jutting over the upper lip.’ Some outputs go further into editorializing, adding ‘droopy regal eyelids’ to imply the royal bearing behind the deformity — a sign the model understood the joke, not just the shapes.

The project fits a growing genre of homemade LLM evaluations, like Simon Willison’s ‘pelican riding a bicycle,’ that sidestep saturated academic leaderboards. Because the prompt is novel, oddly specific, and hard to fake, it functions as a lightweight probe of spatial reasoning and instruction-following that’s resistant to training-data contamination — and it doubles as an entertaining way to watch models reason about something no benchmark designer would have thought to include.

Read the full article

Continue reading at Hacker News →

This is an AI-generated summary. Read the original for the full story.