Researchers’ AI Beats Top Stratego Players With Far Less Training Than DeepNash
The Nature study combines self-play training with planning during each turn. Extending that success beyond games will require decisions people can audit.
Loading page…
The Nature study combines self-play training with planning during each turn. Extending that success beyond games will require decisions people can audit.
Listen to this story
Ataraxos pairs a self-play strategy blueprint with planning at decision time, using a generative model to focus on plausible identities of hidden Stratego pieces. In a study published in Nature on September 30, 2026, researchers report that it surpassed elite players with under one-hundredth of DeepNash’s training examples and under one-thirtieth of its self-play games. The results also extended to three other games, but applying the approach to negotiations or cybersecurity remains a proposal; the team says interpretability and human oversight are still needed.
Ataraxos recorded 39 wins and two losses against top human players at the Stratego world championship.
Against the player MIT News calls the world’s strongest, the system’s reported match record was 15-1-4.
The team also reports superhuman performance in Barrage Stratego, Hanabi, and dou dizhu, games with different rules and cooperation structures.
An AI system has beaten elite Stratego players without knowing the identities of their pieces—and with substantially less training than its leading AI rival. Researchers report that Ataraxos combines a strategy learned through practice with fresh planning before each move. Their study, published in Nature on September 30, 2026, offers a more efficient way to tackle games where crucial information stays hidden, with possible applications in negotiations and cybersecurity.
The team spans MIT, Carnegie Mellon University, New York University and Stanford University. MIT News describes the findings in “Scalable decision-making for games of imperfect information.” The central advance has two parts: making training more efficient, then improving the system’s choices during play rather than relying entirely on what it learned beforehand.
Stratego gives each player 40 pieces and a goal: capture the opponent’s flag. Piece identities remain secret until pieces collide. That makes it a game of imperfect information—a setting where a player must act without seeing everything that determines the outcome. Its possible piece configurations exceed 10 to the 66th power.
Hidden information also changes the value of a move. Lead author Samuel Sokota explains that bluffing becomes less useful as an opponent learns to expect it. A strategy therefore has to account for what the other player may infer, not just which move looks strongest on the visible board.
Ataraxos first learns through self-play reinforcement learning: it repeatedly plays against itself to develop a strong starting strategy, called a blueprint. The researchers designed more efficient training algorithms that let it learn faster without getting stuck trying to predict every possible move. That blueprint guides its initial setup and provides a starting point for later turns.
Before acting, the system refines that starting strategy through decision-time planning. A generative model uses probabilities to estimate the likely identities of the opponent’s hidden pieces. Ataraxos then evaluates future choices before selecting its next move. Instead of guessing blindly, it narrows its attention to plausible board states for the particular game and opponent.
The researchers identify planning with the generative model as the missing ingredient that enabled superhuman play. Ataraxos recorded 39 wins and two losses against top human players at the Stratego world championship. It also beat the player MIT News describes as the world’s strongest, with a reported match record of 15-1-4.
The study also tested adaptations in three other games, reporting superhuman performance in each. These broaden the finding beyond standard Stratego because they involve different rules and forms of cooperation:
The researchers suggest the system could be adapted for business negotiations, cybersecurity or military maneuvers. Those are proposed applications, not the game results themselves. Senior author Gabriele Farina points to a shared difficulty: real-world decision-makers often cannot afford to enumerate every possibility when other parties hold information they cannot see.
The team’s next goal is to add interpretability measures so Ataraxos can explain its decisions in terms people understand. Farina says people must retain the final say over whether to follow a recommendation, and that adoption requires a way to audit the model’s choices. He cautions that the researchers still have a long way to go.
Loading discussion...
Join the conversation
Explain when a strong track record would outweigh the missing explanation.
Be the first to share a perspective or an experience.
Reader comments
Newest comments first. Replies stay oldest first.