DeepMind Paper: LLMs Can Prove Theorems but Can't Invent the Premises
A position paper by Tom Zahavy of Google DeepMind, presented at ICML 2026, argues that today’s language models are missing the specific cognitive move that produces genuine scientific breakthroughs. Borrowing Einstein’s picture of discovery as a two-step cycle — an intuitive ‘jump’ from raw observation to foundational axioms, followed by logical deduction from those axioms — Zahavy maps current AI onto that framework. Models have effectively mastered induction (statistical pattern-matching) and are quickly getting good at deduction (formal, machine-checkable proof), but they have no real mechanism for abduction: generating the novel explanatory hypotheses that the rest of the reasoning hangs on.
The paper uses Einstein’s development of general relativity as its worked example to attack the popular ‘creativity as data compression’ thesis. If creativity were just efficient compression of observed data, it should break down precisely where the observations are sparse — yet that is exactly the regime in which the most important theoretical leaps happen. Zahavy’s claim is structural rather than about scale: an LLM could plausibly carry out the deductive work of proving results once the premises are handed to it, but it is not built to formulate those premises in the first place. That premise-generation step is the ‘jump’ the title refers to.
The significance is less a hard technical result than a framing challenge to AI-for-science optimism, and it comes from inside a leading lab, which is why it drew attention. Zahavy has been explicit that it is a personal position piece, not DeepMind’s official stance, and that he is not claiming models can never contribute to discovery — the narrower point is that current architectures lack the abductive step and that benchmarks emphasizing proof and pattern-matching can mask this gap.
Read the full article
Continue reading at Hacker News →This is an AI-generated summary. Read the original for the full story.