LLMs Crack 98 of 100 Research-Level Math Problems at Leipzig Benchmark
Forty-nine mathematicians, mostly working during a three-day workshop at the Max Planck Institute in Leipzig this spring, assembled a fresh benchmark of 100 research-level math questions with verified answers. The goal was to test how far frontier language models have come on problems that demand genuine mathematical reasoning rather than pattern matching on training data.
The results were striking. A single-shot run across five state-of-the-art LLMs left 41 questions unsolved. Expanding to 20 runs per model on three of them dropped the unsolved count to 16. A final stage using two heavy-thinking models with just three attempts each cleared all but two problems. The authors read this as evidence that LLM mathematical reasoning is closing in on research-grade competence, though the rapid saturation also raises the familiar question of how long any new benchmark can stay meaningful.
Read the full article
Continue reading at Hacker News →This is an AI-generated summary. Read the original for the full story.