Language models can’t spark scientific revolutions, but world models might
Can language models spark a scientific revolution? In a position paper titled “LLMs can’t jump,” Google Deepmind’s Tom Zahavy argues they can’t.
They’re missing the cognitive mechanism needed to create something truly new. Zahavy builds his case on a framework Albert Einstein sketched in a letter to his friend Maurice Solovine. Discovery, Einstein wrote, is a cycle: sensory experience leads to an intuitive “leap” toward axioms, and from there, logical deduction produces testable conclusions. Axioms are the unproven foundational assumptions of a theory. To pinpoint where the gap lies, Zahavy draws on a classic distinction from philosopher Charles Sanders Peirce, who categorized all reasoning by how it connects rules, cases, and results. Deduction derives guaranteed conclusions from fixed rules, like running a program that produces a provably correct output. Induction spots patterns in data: observe a thousand white swans, and you generalize that all swans are white. Abduction is the creative leap. It invents a cause to explain a surprising phenomenon. This third form is where Zahavy sees the critical bottleneck, and he draws a line between two levels of it. Ordinary abduction picks the most plausible explanation from a set of known candidates, the way a doctor matches symptoms to a disease. Language models can do this, he concedes. The harder version is what he calls “manipulative abduction”: inventing a cause for which no linguistic template exists yet. That, he argues, is the real bottleneck of scientific invention, and machines can’t do it. Induction and deduction, the paper argues, are well within reach.