Mastering Stratego
deepmind.com
deepmind.com
Its skills for bluffing are both fascinating and a bit scary.
[0] https://www.youtube.com/watch?v=HaUdWoSMjSY https://www.youtube.com/watch?v=L-9ZXmyNKgs https://www.youtube.com/watch?v=EOalLpAfDSs https://www.youtube.com/watch?v=MhNoYl_g8mo
This is scary to do well in practice, because the mathematically optimal bluff frequency approaches 50% as you increase the overbet size.
It seems like it would be easier for AI to do, since it doesn't have any tells (it's easier to have a poker face when you don't have a face at all).
I remember playing poker as a kid, and experimenting with pretending like my cards were good/bad with body language. I don't think that any professional players use that approach (they just have sunglasses and a straight face), but I wonder if AI could beat humans even more consistently if it developed a way to convey tells and fake tells?
No poker bots yet I know of have developed "exploitative" strategies, where they deviate from the Nash Equilibrium strategy to exploit opponent mistakes.
Back when I played professionally (2012-2016, poker AIs being relevant in 2015/16) the standard was to use a bot to study the best default strategies and use expert human judgement to deviate from it against bad opponents.
It seems like it would introduce a level of predictability that would make it easier to know when the opponent is bluffing.
A penalty kick where the kicker can kick left or kick right. The goalie has to jump one direction, if they jump the wrong direction a goal is scored. Both people know that this kicker is great at kicking to the left side of the goal but rather "meh" at kicking to the right, so if the kicker kicks to the left and the goalie jumps left, there's still a 20% chance of scoring, but if the kicker kicks to the right and the goalie jumps right, there's only a 5% chance of scoring.
There is a Nash equilibrium for the kicker, and it can't be "always kick left" because then the goalie would "always jump left" which would give the kicker an advantage if it kicked right.
Similarly the Nash equilibrium for poker can't be to always fold a weak hand, because that's leaving money on the table because then the opponents will always fold against a raise, which would mean the player could get easy money by raising with a weak hand.
In some sense that's the whole point of Nash equilibria.
You only use the "tell" when you believe there to be a pattern.
Otherwise, if it is truly random, I realize I can't get any info there and I ignore it, which then leads to me just having a straight face.
However at a large table, you are going to get called only by the person who thinks they have the best hand, which is a lot better than the average hand of a typical opponent.
In multiplayer you see it where ranges are narrowed, like 3bet pots or on turn/river
It matters less than you'd think because overbets imply you have a polarized range (nuts or air). You generally pick the bluffs to be hands that have cards blocking the best calling hand combinations.
Searching to find out whether "original" is the current version, I've noticed that https://en.wikipedia.org/wiki/Stratego#Versions answers your question
"European versions of the game give the Marshal the highest number (10), while the initial American versions give the Marshal the lowest number (1) to show the highest value (i.e. it is the #1 or most powerful tile). More recent American versions of the game, which adopted the European system, caused considerable complaint among American players who grew up in the 1960s and 1970s."
There is an incentive to just not move your pieces, so that the other player thinks they're bombs. As a result, players only activate 2-3 pieces at a time.
In chess, on the other hand, you are constantly moving your pawns to the other side to promotion, or otherwise trying to activate/coordinate all of your pieces for an attack.
It makes me think that if deepmind for Stratego was trained to not lose instead of win, then the top strategy might be shuffling pieces and letting the enemy come to attack. No human would ever have the patience to play that way though.
Tournament Stratego uses a clock, which reduces some of the issue there. It's not hard to beat a player that does what you suggest; just send some middling pieces after each piece that moves. You'll take the weak pieces and reveal the strong pieces.
It is much more defensive than chess in general though, as moving a piece and capturing with a piece both give the other player information.
But taking a strong piece means revealing a stronger piece of you own. That’s why I think the best strategy is to put almost all your weak pieces up front.,and wait for your enemy to reveal their pieces. Scouts, especially, are canon fodder that you sacrifice to find out information about the enemy and that you need to get rid of so that you get room to maneuver.
So, the idea is to, in midgame, make educated guesses as to the positions of the spy (e.g. the piece that stays close to the enemy general, or that moves towards your marshall), and sacrifice scouts, hoping to kill the spy?
Interesting tactic (and if that’s their common usage, why are they called scouts? ‘Assassin’ might be a better name)
Since killing the opponents Marshall with the spy is such a huge advantage, losing your spy is a big disadvantage.
I don’t know anybody who plays the game that way, but reading the rules (https://www.hasbro.com/common/instruct/Stratego.PDF), that wouldn’t fully help. The rules say
“scouts are the only pieces allowed to both move and attack in the same turn. A scout can move any number of open squares forward, backward or sideways into an attack position. Once in position, it can then attack”
That doesn’t say in any way that that attack has to be in the direction of movement. So, you could move a scout 3 squares forward and then attack leftwards.
The ISF rule is that the attack must be in-line with the move[1].
1: https://isfstratego.kleier.net/docs/rulreg/isfgamerules.pdf Section 5.3
A "player can claim a draw if no capture has been made and no pawn has been moved in the last fifty moves"
A small scene is probably pretty damn good at the top. Having hundreds of thousands of competitive players helps, but even with a small sample you are probably likely to get at least some very, very strong players.
It's hard to think of a relevant real world example, but a fun corollary I'm familiar with is Fedex (Federico Perez Ponsa). He is a full chess Grandmaster, #461 in the world in chess amongst ~300k active FIDE players. By your "International Masters study hard" logic, when you drop him into Age of Empires 2, a game with ~500 competitive tournament players, his work ethic should dominate. But it turns out that the top ~100 AoE2 players are really damn good and practice a ton (easily chess IM amounts), and Fedex tops out around ~#50 in the world.
https://ratings.fide.com/profile/117927 https://liquipedia.net/ageofempires/Fedex
No matter how good the top players are relative to the competition tho, I feel like a large playerbase still raises the skill bar to a huge degree. There's a ratchet effect where someone figures something out, other people copy, and it breaks into public consciousness through influencers and popularizers. Then on the tail end, regular people regurgitate it for years like it's new information. (Getting sick of hearing about cognitive biases and product-market fit and dunning kruger, ffs). I've seen it with dota over the last 10 years. Pro players were always good, but now bad players are good and pro players are better. Not to mention the motivation that comes from seeing other people work hard.
When it comes to stratego, watching the linked games (from an armchair!), the human players looked relatively sloppy. Overusing scouts in the early game, too eager to trade, noticed a piece being forgotten about once. Not to say I'd be better, but it definitely looks like the scene is "for fun" and not so serious.
I play disc golf pretty regularly. Over the last 10 years, it has exploded in popularity.
I'm not sure that the explosion of popularity has resulted in better people performing in tournaments.
But what has changed, is how much money can be won by competing in disc golf, and how much money is available across the sport as a whole.
The increased prize money has dramatically changed the number of people and level of competition for people playing the game seriously as well as how seriously everyone involved in top level play takes the sport. This then has a trickle down effect in the number of people and seriousness of things like training clinics, professional teachers, professional and more intelligent course construction and analysis. It pays for more analysis into all aspects of gameplay to increase the competitive edge of performers at the top.
The increased player base, in turn, pays for most of this, as it increases the potential market for the same services as all of the above. And the more seriously top level play is, and the higher the winning prize pool money goes, the more respectable the sport has become in the public eye, which in turn creates a catch-22 effect whereby players appear more willing to spend more money on the sport on those services.
All of which increases the level of play at the top. And you can see it in the quality of new young athletes that are coming up in this new environment, and how much better they are and how much more they are able to learn from the more widely accessible resources than their equivalent counter-parts were 10 years ago.
Its been fascinating to watch, and very exciting. Particularly over the pandemic, the sport has come from being called "frolf" on a golf course or in your local park, to "disc golf" with multi-million dollar professional contracts, dedicated disc golf resorts and private courses, training clinics, and dedicated PPV channels. Very cool : D
A counter-example might be the various competitive communities around different forms of boardgames, which tend to be very small, but often very competitively driven and taken very seriously by a very dedicated community, but I don't know enough about the topic to discuss : )
In fact, the idea that someone can train with sufficient intensity to be a high ranking chess master then break in to the top 50 of AoE at the same time suggests a lower skill saturation in the AoE world.
The problem here is that it's missing the "glue" to more real world applications. This is where more humdrum software engineering comes in.
Diplomacy in this is much more interesting than Stratego or beating the next video game - it mixes cooperative game theory with NLP and reinforcement learning.
This is a bold statement. The world does not function based on "well-defined rewards". The concept of "common sense", which some consider table stakes for a human operating competently in our world, is mostly made up of things which are neither well-defined, nor allow a tremendous amount of training data. Current ML approaches require both.
But in general, yes, this is why since 2013 we don't see AI making anywhere as massive strides everywhere as they do in boxed-in applications like games.
Humans can be neatly reduced to a "procreate, and make sure your progeny procreate" value function. A lot of apparent counterexamples to that turn out to make sense when 2nd order effects on progeny are considered.
It's been providing real value to me over the past day for practicing Spanish, explaining Machine Learning concepts, and doing fancy write-ups in LaTeX. And this one can't even use Google yet! (other research teams have already created models capable of doing so, it's only a matter of time until these innovations are brought together in one place)
I'm glad I forgot about them and opted out of Copilot. Fwiw, I'm currently in Cambodia.
---
If you took the word "Wikipedia", and threw all the letters up in the air, what are some possible things they'd spell once they fell on the ground again?
If you took the word "Wikipedia" and threw all the letters up in the air, some possible things they could spell once they fell on the ground again might include "kawipieda," "keiapwidi," "kipedwaia," "pwakieida," or "eipawikid." However, without being able to see the exact arrangement of the letters on the ground, it's impossible to say for sure what they would spell.
---
(last sentence is a bit confused)
Input: "If five cats can catch five mice in five minutes, how long will it take one cat to catch one mouse?"
Output: "If five cats can catch five mice in five minutes, then it will take one cat one minute to catch one mouse. This is because the number of cats and mice are inversely proportional, so as the number of cats decreases, the time it takes for a cat to catch a mouse will increase. Similarly, as the number of mice decreases, the time it takes for a cat to catch a mouse will decrease. Therefore, if five cats can catch five mice in five minutes, then it will take one cat one minute to catch one mouse."
As you can see, it fails at simple logic because it fails to understand that the number of mice aren't fixed in the question. Whereas a human would immediately pick that up because they reason beyond x=5 and y=5.
You could probably test this by asking a few friends this question and see what they say. Outside of pure math problems you can get into an infinite regress defining the underlying first principles behind any given assumption.
Are you claiming its training data has letter permutations of the word “Wikipedia”?
It’s actually pretty capable of doing basic combinatorics.
[Citation needed]
I think that the computer made the absolutely standard human mistake, so that could be considered a plus.
I have seen reports that it will happily hallucinate a plausible but wrong answer to all sorts of different prompts, intermixed with many mostly correct answers. It's interesting to think about how to place trust in such a system.
Anyhow spoiler alert, the neural nets running the virus response have been inadvertently trained to prefer simple systems over complex ones without anyone realizing, and decide that a planet with no life on it after being wiped out from the virus is infinitely more simple than the present one and starts helping it out instead of stopping it.
So short answer to your question is I would not place much if any trust and systems like that, in as far as anything that has high stakes, real world consequences.
As others mentioned, AI is making headspace in enterprise and accounting, and achieving the “last mile” of human work. Better image recognition for handwritten forms and mail, better content and sentiment analysis for reducing spam, robot arm tasks which are more and more complex (yet still tame compared to humans)
“AI” hype is indeed overrated. If you think we’re close to reaching the singularity or anything resembling skilled human work you will almost certainly be disappointed. We probably have decades of slow improvement, more and more of these “breakthroughs” which aren’t really amazing compared to a human 5-year old, and aren’t really going to revolutionize industry, but will nonetheless have practical benefits
The actual models work fantastically well.
The board games are merely a cover to advertise to AI Researchers and portray AI as "innocent" in the public eye.
Stratego is Google goofing off.
The Ferrari AI models are being used by Google to absolutely swindle money in some ad tech niche.
I agree that these specific models are not going to be useful outside of board games. But in the future when there is the opportunity for AIs to interact with the world for real, the this kind of research will allow AIs to dramatically outperform humans on these tasks.
That's what one lab is doing.
You cannot be blind to many many applications that are finding their ways to consumers and earning people money.
At least this new model is very different to the approach taken in Alpha Go.
Don’t get me wrong, using AI for that purpose is pretty amazing (but can also lead to some sketchy results if you don’t know what you are doing[1]) but pretending it will lead to some “general AI” is nothing but hype IMO. And teaching AI to play these board games better then a grandmaster only serves to increase that hype.
1: https://www.vox.com/recode/2019/8/15/20806384/social-media-h...
There are for sure use cases for inference models in generalized (or rather ill-understood; or even highly dynamic) non-linear systems, and deep learning models kind of ace at that—given enough training data and a lot of computational power. However I’m not really sure what we will use AGI for.
I’m curious about the comparisons to poker. I know the hot algorithm in poker solvers is counter factual regret minimization. The article indicates that the feedback cycle is too long for those algorithms to work but I’d be curious to learn more about the relationship from CFR to what’s tried here, if any.
Stratego is 40% information, 40% strategy, maybe 10% tactics. If you know where is the flag it's trivial to win in almost all situations. Fighting into imperfect information is literally all the game.
(The old joke is that Dota is a 1 v 9 game, not a 5 v 5)
Both modes of play are fun, mind, but the parallels struck me as worth noting.
Like is it just the speed of your clicking? Or is it more than that, like the most basic kinds of strategic decisons?
It's obviously not about clicking fast, but it is about timing, sometimes 100 milliseconds reaction time make huge difference in outcome. It is usually making decisions on very small time scales. Do you retreat or continue? Use ability or hold it? Can you overextend?
The only meaningful strategic decisions in dota (which you have long time frame of deciding and effect the game for a long duration) are draft (which AI doesn't really master, they reduced the heroes pool to simplify) and item purchases, and there are only a handful of them (~6) in an entire game. Other decisions don't really have a long "memory" time, a minute or two at the most. After two minutes every other decision is just reduced to the relative advantage between the teams.
There used to be one hero in Dota which made it a strategy game instead (techies). But it was like playing a different game and everyone hated it and it was effectively removed. Techies was like playing stratego against chess players, they obviously get pissed off by not playing what they wanted.
What's your MMR, out of curiosity?
I do think OpenAI five derived some of their advantage from seamless ability usage, and inhuman coordination in lane, but it also did some novel things strategically, that challenged some of the established tenets of high level play (e.g. they had a lot of mobility on the map, back when that was considered very inefficient).
Stratego is way more challenging than Poker, though. StarCraft/Dota/Stratego have the property that you can’t represent their imperfect information as a vector in memory, whereas you can easily do that in Texas hold’em poker (there’s only 52C2 = 1024 possible hands). So for those games, you have to use an approximate distribution rather than the exact one.
I’m an author on this paper (although my contributions were relatively minimal) and on the Player of Games paper (which did Poker, Chess, and Go).
My very naive question question being: Why tackle this after SC and Dota have already been done? What is the scientific interest? Stratego seems strictly simpler than both of these games. In what way is this an advancement over how SC/Dota AI were solved?
In principle, the AlphaStar's league approach (from StarCraft) could be done also in Stratego, and it would be very interesting to compare the two approaches. Note that AlphaStar is more expensive: it required to train N competing agents with pair-wise evaluation costing N^2, while Stratego's NeuRD trains a single agent.
Feels like the secret sauce has to be probability distributions guessing what all the pieces are.
Bluffing in stratego seems like it requires long-term planning (if you move a 2 like a 10, you have to keep treating it like that for the bluff to work).
https://github.com/deepmind/open_spiel/tree/master/open_spie...
Thinking back it may have been the sergeant who was the primary Hat Guy as I remember there was "Hat Guy" and later "Other Hat Guy".
https://github.com/deepmind/open_spiel/tree/master/open_%20s...
Someone needs to create a web front end for this -- I would love to play it.
It makes the game so much more interesting, IMO. Played it a lot as a child.
Here are the basic rules, when a piece is attacked:
* The attacker says what their piece is, without showing it (they can lie)
* The defender says whether they believe that
* The defender says what their piece is, without showing it (they can lie)
* The attacker says whether they believe that
* ONLY IF someone calls a bluff is that piece revealed. Otherwise, it is treated as the piece it was claimed to be, and kept hidden.
* If someone calls a bluff, and they were right, then the other player loses a piece (reach over and remove any piece you like)
** If you pick their flag, then you win — game over.
* Likewise, if someone calls a bluff but is wrong, then *they* lose a piece.
* After all of that is resolved, do combat as normal, with pieces having either their revealed or not-revealed claimed value, as appropriate.
Once you resolve all this, there is no "memory" - you can claim it is a different piece in the future.Some minutiae:
* You can move any piece as though it were a Scout (9), but when you do the move, the other player can call your bluff since you're essentially claiming it is a Scout at that moment. Resolve that bluff/call before completing the move.
* You could even call a bluff on *any* move someone makes, if you believe that piece is a bomb or flag (and thus cannot move).
* You can attack with a bomb! It's a two-step process: first you move (and they could call your bluff, if they know it is a bomb - see above). Then, when the attack happens, you say it *is* a bomb. Of course, your opponent may say their piece is a Miner, and if you haven't seen it, it's a dangerous proposition (since bombs are rare).
** You can also do a variant where bombs can't attack (by attacking, you are claiming it is *not* a bomb). I prefer the above version.
Overall, I find this version of the game is a lot less boring. Since you'll probably get several pieces zapped over the course of the game, it affects your flag placement. Plus, you can move flags and bombs, making it more dynamic. Also, the "remember where things were" aspect is even more poignant, since once a piece has been revealed, it loses all the power of being whatever-is-needed-right-now (assuming the other player has a good memory).So, for instance, you can do something crazy like move your bomb as though it were a Scout, all the way across the board, onto an opponent's piece, but then claim it's a "5" instead for the attack. Then if it survives, just let it sit there, continuing to be a bomb in the future (causing havoc).
I remember my brother and I as kids playing Stratego and discovering the “impenetrable bunker of bombs” to put your flag in. Which evolved to “put a scout in as a ruse” and later “don’t actually enclose it because now brother just assumes it’s enclosed.”
Very interesting question btw. I'm also interested in the StarCraft AI state of the art 2022. If it exists. Or was focused changed to some other game?