Puzzles and games have long been central to AI development, serving as a benchmark for measuring progress. As humans test their smarts with crosswords and logic puzzles, developers use similar challenges to evaluate AI models. The term 'machine learning' was popularized in a 1959 article by IBM computer scientist Arthur Samuel, who described an algorithm that learned to play checkers. Chess and Go are also famous AI test beds. In late 2024, a team from Columbia University showed that even the best models could solve only 18% of the infamous New York Times Connections puzzles. By early 2025, some models could solve them near perfectly every time. Despite these advances, today’s models still struggle with subtle changes in classic riddles and visual puzzles. These challenges highlight the differences between human and machine cognition, offering insight into the technology’s strengths and weaknesses. Some puzzles may be as tricky for humans as they are for AI, while others are so simple that they raise questions about AI’s true capabilities. If a person can solve them, they may prove they can out-puzzle an AI—at least for now.

Spatial reasoning remains a domain where humans have a significant advantage. Mental rotation problems, often found in IQ tests, ask individuals to determine whether different images represent the same object from different angles. While language models can analyze visual inputs, they still fail at these puzzles. Despite claims about world models helping AI understand physical environments, large language models (LLMs) still struggle to manipulate 3D objects like spatial thinkers such as architects and engineers. These puzzles test the ability to visualize and rotate objects, a skill that remains challenging for AI.

Memory and adaptability are also areas where models face limitations. Frontier LLMs have extraordinary memory due to their exposure to vast amounts of training data, but this can be a liability. In a 2024 study, researchers from Google and the University of Illinois Urbana-Champaign found that models often fail to notice key differences in puzzles similar to those they encountered during training. This led to incorrect answers based on memorized patterns rather than logical reasoning. The same principle may apply to the SimpleBench test, where models struggle with questions that resemble ones they saw during training. Humans can spot the trick, but even top-tier models often trip over these problems.

Source: mittr