Leading mathematicians Timothy Gowers and Peter Sarnak have stated that large language models (LLMs) are highly effective at performing calculations and applying established mathematical methods. However, they emphasize that these models struggle to produce original insights or develop new theoretical frameworks. Gowers explained that while LLMs can efficiently combine known techniques and explore numerous paths, they lack the intuitive ability to identify the most promising routes in complex problem spaces. This limitation, he noted, hinders their capacity to generate truly innovative mathematical ideas.

Sarnak echoed these concerns, highlighting that although LLMs can derive results from existing mathematical theories, they fail to create the foundational abstractions necessary for major proofs when starting from basic questions. DeepMind researcher Tom Zahavy, in his paper 'LLMs Can't Jump,' identified the core issue as 'manipulative abduction'—the ability to invent new foundational assumptions without prior linguistic context. Zahavy’s findings align with the broader consensus that current LLMs are constrained in their capacity to produce original, groundbreaking ideas.

The assessments reflect a growing debate within the AI research community about whether LLMs are becoming more versatile or merely improving at benchmark tasks and familiar problem domains. These insights underscore the need for alternative approaches, such as world models, which could potentially overcome the current limitations of LLMs in creative problem-solving.

Source: thedecoder