I think that what makes these games beatable repeatedly is that they're static. Not saying an algorithm properly trained won't play better than the average player a game like MtG, or my own https://aethersummon.com (specially now while it has under 90 possible scrolls only) but if you have a regular release cadence (say weekly or bi-weekly) of relevant new "cards", then I think the playing field is much more even for humans.
Those new additions can invalidate the whole training data by a single new "card" that changes completely the dynamics and would be easy for a player to understand and incorporate but not for an algorithm (perhaps with enough compute to re-train it regularly it could) - that along with the decision trees being orders of magnitude deeper, wider and with more conditionalities than go, chess or stratego - even through the same turn with the same cards available and same table state - would probably pose much harder problems for a compute bound algo.
There are very few missing pieces for a game like MTG. The main reasons we don't have a Stockfish for MTG is that it's a PITA to implement the rules and that nobody cares (or at least not enough to make it happen.)
There is nothing that, in principle, makes MTG different from poker or bridge, and we have superhuman engines for both.
> Those new additions can invalidate the whole training data by a single new "card" that changes completely the dynamics
This doesn't follow. You're basically proposing that new combo decks be added all the time, and it's far simpler for an agent to scan the new cards for potential interactions with the thousands of other cards in circulation than for a human to remember all of them.
Your analogy is akin to saying that all you have to do is keep landing new code all the time, and since the agents weren't trained on the code they won't be able to identify and respond to security vulnerabilities in it as fast as humans, which hasn't turned out to be correct
Oh no! Stratego had been on my mind as something we just hadn't tried hard enough to make a winning bot for, including the DeepMind effort from 2022. I was planning to make the first one.
I thought this was slightly less crank-coded than trying to prove the Riemann Hypothesis, but maybe these days you just ask Claude to do that and it tells you there's a counterexample at 1 + πi that no one ever noticed before.
> Now, a team of researchers from Carnegie Mellon, MIT, New York University, and Stanford University has done it. Their AI, called Ataraxos, beat Pim Niemeijer, arguably the best Stratego player of all time, 15 games to one, with four draws. And it took just 16 GPUs and a few thousand dollars to train it.
Just 16 GPUs, and a few thousand dollars?
What about “researchers from Carnegie Mellon, MIT, New York University, and Stanford University” this wasn’t just anyone.
This approach also works for Hanabi, which is a very interesting game. You can't see your own cards, but the other players can. I bought the game because someone on a reinforcement learning podcast [2] mentioned it, and actually played it multiple times.
This puts the earlier "Mastering the Game of Stratego with Model-Free Multiagent Reinforcement Learning", 2022 [1] in some perspective. Apparently the "mastering" in 2022 wasn't quite there yet. Four years later, the new approach seems to actually be better than humans.
I wonder how capable current AIs are with the "silent defense" variant of Stratego [1]? The article states the high level of uncertainty presents a challenge. With silent defense the uncertainty is even higher.
[1] https://www.hasbro.com/common/instruct/Stratego.PDF
"When an attack is made, the attacker is the only player who has to declare the number of his or her piece. The defender does not reveal the number of his or her piece, but resolves the attack by removing
whatever piece has a lower number from the gameboard. Players keep their own captured pieces. Exception: when a Scout attacks, the defender must reveal the number of his or her piece.
As a kid, a friend of mine had "Electronic Stratego"[1], the biggest gameplay change was that you could carry out fights without revealing the strength of either piece to the other side. I found this made for a much more interesting game and we had quite a bit of fun playing it.
Those new additions can invalidate the whole training data by a single new "card" that changes completely the dynamics and would be easy for a player to understand and incorporate but not for an algorithm (perhaps with enough compute to re-train it regularly it could) - that along with the decision trees being orders of magnitude deeper, wider and with more conditionalities than go, chess or stratego - even through the same turn with the same cards available and same table state - would probably pose much harder problems for a compute bound algo.
There is nothing that, in principle, makes MTG different from poker or bridge, and we have superhuman engines for both.
This doesn't follow. You're basically proposing that new combo decks be added all the time, and it's far simpler for an agent to scan the new cards for potential interactions with the thousands of other cards in circulation than for a human to remember all of them.
Your analogy is akin to saying that all you have to do is keep landing new code all the time, and since the agents weren't trained on the code they won't be able to identify and respond to security vulnerabilities in it as fast as humans, which hasn't turned out to be correct
I thought this was slightly less crank-coded than trying to prove the Riemann Hypothesis, but maybe these days you just ask Claude to do that and it tells you there's a counterexample at 1 + πi that no one ever noticed before.
Just 16 GPUs, and a few thousand dollars?
What about “researchers from Carnegie Mellon, MIT, New York University, and Stanford University” this wasn’t just anyone.
[1] https://en.wikipedia.org/wiki/Hanabi_(card_game)
[2] https://www.talkrl.com/episodes/jakob-foerster
[1] https://arxiv.org/abs/2206.15378
[1] https://www.hasbro.com/common/instruct/Stratego.PDF "When an attack is made, the attacker is the only player who has to declare the number of his or her piece. The defender does not reveal the number of his or her piece, but resolves the attack by removing whatever piece has a lower number from the gameboard. Players keep their own captured pieces. Exception: when a Scout attacks, the defender must reveal the number of his or her piece.
1. https://boardgamegeek.com/boardgame/3513/electronic-stratego (We generally banned the use of the 'probing' feature)