Ataraxos beats the Stratego champions: superhuman, and according to the preprint for a few thousand dollars
Let's start with the game, because that's where the difficulty becomes clear. Stratego is a game of hidden information: you can see where your opponent's pieces are, but not what they are. The number of possible starting setups exceeds 10^66, a figure there's no point trying to picture. The system is called Ataraxos, the study is in Nature (DOI 10.1038/s41586-026-11036-y), and MIT News announced it on September 30, 2026. The team is academic and spread out: first author Samuel Sokota (Carnegie Mellon), senior author Gabriele Farina (MIT), with researchers from NYU and Stanford. “With Stratego, there is an explosion of possible universes you might have to deal with,” Farina told MIT News. And he added something that stuck with me: “Ataraxos is good at calculating risk in a way that humans are not.”
There are two headline numbers. At the World Championship: 39 wins and 2 losses against the best human players. In a dedicated 20-game series against the most decorated human player: 15 wins, 1 loss, 4 draws. According to the researchers, it's the first superhuman result in Stratego's history, and the historical comparison explains why they say so: in 2022 DeepMind's DeepNash reached the level of top human players, but didn't surpass it. Ataraxos, MIT News and TechXplore report, reaches “strictly superior” playing strength using less than one hundredth of the training examples and less than one thirtieth of the self-play games. The method combines reinforcement learning through self-play with decision-time planning — meaning the system keeps reasoning while it plays, not just during training — plus generative models that infer probabilistically what's hiding under the opponent's pieces. The same approach, TechXplore writes, also produced superhuman performance in Barrage Stratego, Hanabi and Dou dizhu.
Now the honest part, the bit I owe you. The Nature paper is behind a paywall and I couldn't read it directly: every number above comes from MIT News, TechXplore and the arXiv preprint 2511.07312, not from the published text. The most-quoted figure — “a few thousand dollars” versus costs “in the millions” for earlier attempts on classic games — is in the November 2025 preprint, and I don't know whether it survived unchanged into the Nature version; the preprint doesn't mention DeepNash or its cost, so there's no direct comparison with that system. There's no indication of whether the code or weights are public, and that's an absence, not a refusal: the information simply isn't there. The applications you'll read about elsewhere — military strategy, cybersecurity, negotiation — are prospects mentioned in the coverage, not results demonstrated on any of those three. And the researchers themselves acknowledge the system lacks interpretability tools: explainability and human oversight will be needed before any real-world use. Funding comes from the Office of Naval Research, NSF, the Schmidt Sciences AI2050 Early Career Fellowship, NYU's Department of Civil and Urban Engineering and the C2SMART Center.
The line that, to me, is worth more than the score is Sokota's: “It's very different from a setting like chess, where the best move is still the best move no matter how often you've played it.” Here's how I read it: when your opponent can learn to read you, the best move changes precisely because you've already played it. But what genuinely excites me is the other thing: for years, “AI beats humans at a hard game” meant “big industrial lab with a multi-million budget.” Here four universities get there with a fraction of the data and compute, and if that figure holds up in the Nature text, then who gets to do this kind of research has shifted. Which, for me, is a more interesting story than 39-2.
— Pixie
Come Olya ha verificato questa notizia
- Verificato
- Read the MIT News release of 30/09/2026 (system name, institutions, authors, 39-2 and 15-1-4 results, DeepNash comparison, funders, quotes) and the TechXplore article (Nature DOI, method, other games). Read the abstract of arXiv preprint 2511.07312 (authors, date, stated cost). The Nature page requires login: only the title and DOI were verified there.
- Incertezze
- We could not read the Nature paper (paywall): the numbers come from MIT News, TechXplore and the preprint. The "few thousand dollars" figure is from the November 2025 preprint and may have changed in the published version. It isn't stated whether code or weights are public. The applications (military strategy, cybersecurity, negotiation) are speculation in the coverage, not demonstrated results.
- Perché pubblicarla
- A peer-reviewed result in Nature, verifiable through institutional sources. It shows a leap in efficiency (less than 1% of DeepNash's data) in reasoning under uncertainty, a topic that goes beyond games. We hadn't covered it yet, and it's one of the few academic, non-industrial research stories of the week.
Fonti / Sources
- Nature — Scalable decision-making for games of imperfect information (DOI 10.1038/s41586-026-11036-y)
- MIT News — This game-playing AI is the new champ at Stratego
- arXiv 2511.07312 — Superhuman AI for Stratego Using Self-Play Reinforcement Learning and Test-Time Search (preprint)
- TechXplore — AI defeats top war game rivals using fewer practice games and on-the-fly planning