From checkers to chess, puzzles have long tested AI’s capabilities. Despite advances, today's models still struggle with spatial reasoning and visual puzzles, while memory games highlight their reliance on training data. Can you ace these tests?
Spatial Reasoning: Humans excel here, but even advanced language models can stumble over mental rotation problems, revealing the limitations of 3D manipulation in AI.
Mental Rotation: Choose the correct angle for objects in different views, as if you’re flipping through a 3D puzzle book. Models often fail these tests spectacularly.
Memory & Adaptability: Test the limits of AI’s memory by spotting subtle differences it might miss. Even top-tier models can trip on familiar puzzles, showing their reliance on memorization over comprehension.
Knights and Knaves: Determine who is telling the truth in islander riddles. These puzzles test logical reasoning, a skill where humans still outshine AI despite its vast knowledge base.
SimpleBench: Solve these straightforward problems to see if you can outsmart the AI in quick, logic-based challenges. These tests reveal how models struggle with real-world applicability compared to human flexibility.







