LLMs Can't Jump: Why AI's Creative Leap Remains Out of Reach
TL;DR
Google DeepMind researcher Tom Zahavy's position paper, "LLMs can't jump", presented at ICML 2026, argues that while large language models excel at induction and deduction, they fundamentally fail at abductive reasoning—the creative leap required for paradigm-shifting scientific insights. This limitation, rooted in their lack of embodied experience and inability to perform counterfactual interventions, has sparked debate among AI researchers about the future of scaling and the potential need for alternative architectures like world models.
The Paper That Challenges AI's Ceiling
In a thought experiment that has rippled through the AI research community, Zahavy poses a provocative question: given all of physics knowledge predating general relativity, could an AI—as Einstein did in 1907—imagine a freely falling elevator and deduce from it the equivalence principle that gravity and acceleration are locally indistinguishable? His answer is a resounding no.
Drawing on Charles Sanders Peirce's tripartite classification of reasoning, Zahavy demonstrates that while LLMs can process vast amounts of data and make logical deductions, they lack the creative "leap" essential for generating entirely new axiomatic frameworks. This isn't just a philosophical quibble; it's a structural limitation that could define the boundaries of what current AI architectures can achieve.
Why LLMs Are Stuck
The core issue lies in what Zahavy calls the lack of embodied connection to the physical world. An LLM, as one critic put it, is essentially a "Chinese Room" operating in high-dimensional space—processing statistical probabilities between tokens without any perceptual interconnection to reality. When generative video models show an apple falling, they're relying on statistical continuity in pixel changes within training datasets, not on an internal understanding of gravity.
This absence of counterfactual intervention is critical. Humans can mentally "cut the elevator cable" within an internal world model, observe the consequences, and deduce physical laws. LLMs cannot perform this kind of active experimentation. They're limited to what's already been described in language, which is why they struggle with tasks that require simulating physical reality—from safely loading a dishwasher to folding laundry.
Scaling Alone Won't Save Us
This limitation has profound implications for the industry's current obsession with scaling. Ex-OpenAI researcher Adam Hunt shares Zahavy's skepticism, betting that $100 billion will flow into training data because scaling alone won't cut it. Hunt's confidence in this prediction sits at only about 40 percent, acknowledging that technical advances could prove him wrong—for example, if specialized AI models can be combined into something resembling more general intelligence.
Hunt's own perspective has shifted from optimistic to increasingly pessimistic. He argues that the latest models aren't becoming more versatile but more specialized. Programming and complex math capabilities keep improving, while areas like language quality and simple logic are stagnating or even getting worse. This observation of uneven performance aligns with Zahavy's structural explanation for why models are inconsistent.
The Two Paths of AI Development
This creates a stark choice for the industry. On one path, we see gradual, broad improvement across all tasks—the optimistic view that scaling will eventually produce general intelligence. On the other, we witness extreme specialization where a few capabilities spike while core skills stagnate or shrink. Current evidence increasingly points toward the latter.
Hunt's own view of LLMs has shifted from optimistic to increasingly pessimistic. He argues that the latest models aren't becoming more versatile but more specialized. Programming and complex math capabilities keep improving, while areas like language quality and simple logic are stagnating or even getting worse, which lines up with Ho's observation of uneven performance.
Beyond LLMs: The World Model Alternative
As a possible fix, Zahavy points to action-controllable world models that allow for counterfactual experiments. These systems would enable AI to simulate physical scenarios and test hypotheses in a way that LLMs cannot. LLMs could still play a key role in such advanced systems, even if they hit their limits when working alone.
Yann LeCun, Meta's chief AI scientist, has long advocated for a different approach. He argues that humans don't mentally simulate the world in great detail—certainly not pixel by pixel. When we imagine a glass falling and smashing, we don't predict every shard's position. Instead, we run highly compressed models that capture only the aspects that matter. LeCun suggests this is the wrong approach, or at least not the best, compared to reconstructing future observations as faithfully to training data as possible.
What This Means for the Industry
The implications extend beyond research labs. If AI can't be creative in the most meaningful sense, we might be entering an era where Bachelor of Arts degrees become the new computer science degrees. The skills that AI struggles with—creative thinking, abductive reasoning, understanding physical reality—are precisely those cultivated by the humanities.
Yet there's a more immediate concern. As AI becomes more specialized, the gap between its impressive capabilities and its fundamental limitations becomes harder to bridge. The industry's current trajectory of scaling models with more data and compute may be hitting a wall that no amount of resources can overcome.
The Road Ahead
Zahavy's paper doesn't suggest that AI is doomed to mediocrity. Rather, it reframes the challenge: the next breakthrough may not come from making LLMs bigger, but from building systems that can interact with the world—whether through embodied robots, world models, or hybrid architectures that combine the strengths of multiple approaches.
Hunt himself puts his confidence in the $100 billion training data prediction at only about 40 percent, acknowledging that technical advances could prove him wrong—for example, if specialized AI models can be combined into something that resembles more general intelligence. The race is on to find that combination.
For now, the takeaway is clear: while AI can solve century-old conjectures, it can't imagine Einstein's elevator. And that distinction might matter more than any benchmark score.
Related News

Untitled

Mistral's Shieldstral: 3B Open-Weights Model Redefines Multimodal Moderation

Bradbury's 'Soft Rains' Resurfaces: A 1950 Dystopian Warning for the AI Age

Untitled

Untitled

