The development of artificial intelligence (AI) has long been intertwined with puzzles and games, serving as benchmarks for measuring progress. Just as humans challenge their intellect with crosswords and logic games, developers use similar tests to evaluate AI advancements. The term ‘machine learning’ gained traction thanks to IBM scientist Arthur Samuel, who created an algorithm to play checkers. Renowned games like chess and Go have also been pivotal in assessing AI capabilities. Recent advancements show that AI’s ability to tackle puzzles is improving rapidly. For instance, a Columbia University team revealed that even top models managed to solve only a fraction of the New York Times Connections puzzles last year, although some have since begun to excel at them consistently.
Puzzles not only highlight AI’s strengths but also expose its weaknesses, providing insights into where human cognition still outperforms technology. Despite significant strides, current AI models struggle with nuanced variations in classic riddles and have particular difficulty with visual puzzles. To illustrate this, readers can engage with puzzles that have historically challenged AI. Some may prove to be as perplexing for humans as they are for AI, while others might seem deceptively simple, prompting questions about the true nature of artificial intelligence. For example, spatial reasoning is an area where humans excel, particularly in tasks that require mental rotation of objects. While language models have made strides in processing visual information, they still fall short in performing spatial manipulations, a skill that architects and engineers navigate with ease.
Additionally, AI’s approach to problem-solving can lead to memory-related pitfalls. Models trained on vast datasets may overlook critical differences in puzzles similar to those encountered during training. Research from Google and the University of Illinois Urbana-Champaign demonstrated that AI can struggle with puzzles like Knights and Knaves, where characters either tell the truth or lie. In contrast, humans can easily discern the trick in these scenarios. As AI continues to evolve, it remains evident that some puzzles, especially those requiring abstract reasoning, still pose significant challenges. Tasks like the ARC-AGI benchmark, which assess the ability to infer rules from examples, reveal that while AI has improved, it often employs convoluted reasoning methods that differ from human intuition. As we engage with these puzzles, it becomes clear that while AI is advancing rapidly, there remain areas where human intelligence still holds the upper hand.
Source: AI models flub these intelligence tests. Can you fare any better? via MIT Technology Review
