In the rapidly evolving field of artificial intelligence, traditional benchmarks have often been criticised for their limited scope and relevance. Many current methods rely overly on rote memorisation and cover topics that might not reflect the diverse challenges faced by users. In response, a growing number of AI enthusiasts and researchers are turning to video games as a novel medium to assess AI’s problem-solving skills.

One such innovator is Paul Calcraft, a freelance AI developer, who has designed an application where AI models engage in a Pictionary-like game against each other. In this setup, one model creates a doodle, and the other attempts to guess its representation. As Calcraft explained, the drive behind this creation was the potential of the game to genuinely assess AI’s broader capabilities rather than simple memorisation.

The inspiration for Calcraft’s initiative came from British programmer Simon Willison who experimented with models rendering complex vector drawings, seeking to push AI models beyond conventional memory-based tasks into real strategic thinking. Calcraft remarked, “The idea is to have a benchmark that’s un-gameable.”

Similarly, 16-year-old Adonis Singh has developed a tool named Mcbench, designed to evaluate a model's ability to create structures within the world of Minecraft. This endeavour draws parallels with Microsoft’s endeavours through Project Malmo, which also employs the Minecraft environment to test AI systems. Singh’s approach hinges on the belief that Minecraft presents a unique, less constrained benchmarking opportunity that necessitates creativity and resourcefulness from AI models.

The integration of games into AI benchmarking is not a new concept. Historical figures like mathematician Claude Shannon have long asserted the challenge presented by games like chess in evaluating intelligent software. In modern applications, companies like Alphabet and OpenAI have deployed AI in games such as Pong, Breakout, Dota 2, and Texas hold 'em. These games serve as diverse testbeds for probing different aspects of AI capability, specifically logic and decision-making.

Enthusiasts today connect large language models (LLMs) with these games to explore how these models handle logic-based tasks. LLMs vary widely in their reliability and response, making them intriguing subjects for such game-based assessments. Matthew Guzdial, an AI researcher at the University of Alberta, supports using games as a more intuitive and visual method compared to text-based benchmarks to understand a model’s performance.

Comparatively, models involved in Calcraft's Pictionary-style games are examined for their grasp on concepts like shapes, colours, and spatial prepositions. While Calcraft acknowledges the game's limitations as a full-proof logic test, he believes it encourages strategic thinking and clue interpretation—critical yet often challenging aspects for AI models.

Meanwhile, Singh continues to advocate for Minecraft’s potential, noting the alignment between the game’s tests and a model’s reasoning capabilities. However, Mike Cook, a research fellow at Queen Mary University, remains sceptical of Minecraft’s perceived distinctiveness in testing AI. From his perspective, Minecraft’s realism may be overstated, equating it more to traditional video games like Fortnite or Stardew Valley, albeit with different aesthetics and tasks.

Nevertheless, despite varying opinions on their comparative merits, games such as Pictionary and Minecraft allow researchers a unique insight into AI behaviour and problem-solving. These new approaches are helping to redefine our understanding of AI capabilities in a world where traditional benchmarks have become increasingly inadequate.

Source: Noah Wire Services