A team from Carnegie Mellon, NYU, Stanford, and MIT developed Ataraxos, an AI that decisively beat Pim Niemeijer, the most decorated Stratego player. The AI's victory in a 20-game series with 15 wins, one loss, and four draws ended one of the last human strongholds in board games.
Ataraxos, introduced in a paper published in Nature, is the first AI to achieve superhuman performance in Stratego. Niemeijer, who has won four world championships and 15 Dutch national titles, is considered the best Stratego player of all time by George Franka, the only player who has competed in every world championship since 1997.
Stratego is particularly challenging for AI due to its high level of imperfect information. With over 10^33 possible setups, the game's complexity makes it difficult for AI to determine the value of each move based on prior and current information. The researchers note that methods derived from poker AI could only handle games with minimal hidden information, unlike Stratego.
The Ataraxos team used a custom GPU simulator and achieved results with significantly lower computational costs compared to previous attempts. While DeepMind's DeepNash system required millions of dollars in compute resources, Ataraxos needed less than $8,000, making it much more cost-effective.
The AI's training method, which includes regularization to prevent over-specialization, allows it to vary its strategies and avoid predictable patterns. This is crucial in Stratego, where randomness and adaptability are essential for success. The researchers compare this process to an energy reserve, ensuring the AI maintains a balance between exploration and exploitation.
Despite a built-in disadvantage, Ataraxos achieved an 85% win rate against Niemeijer, with draws counted as half a win. The AI's unpredictable style, including bluffing and strategic stalling, made it difficult for human players to adapt. The results suggest that large amounts of hidden information are no longer a barrier for reinforcement learning in strategic decision-making.
Source: thedecoder