No limit: AI poker bot is first to beat professionals at multiplayer game
nature.com
nature.com
Judging by his name, I'd assume Swedish is his first language, so that particular aspect isn't that surprising to me
It's not uncommon to speak four languages (often C2 in couple of them) in the North Europe, esp. the Baltic region.
Like mentioned by sibling (sakarisson), that particular part is not impressive, the rest - sure
Clearly we have different experiences (swedish person living in spain currently) but I haven't met that many people who speak four languages and are from a european country (but have yet to been in eastern europe).
That finns speak swedish is a special case though, as AFAIK, they learn swedish in school and being finn-swedish is a thing too.
Our solo flute speaks a whopping 6 languages well, and I suspect our harp player knows even more.
Nordic countries are a special case.
Norden er et spesielt tilfelle.
Norden är ett speciellt fall.
Norden er et specielt tilfælde.
Northern Europe, maybe. French people for instance tend to suck at foreign languages. We rarely go beyond 3 languages (French, English, then German or Spanish. The last two are often forgotten after school.)
I suspect Spain and Italy are similar.
It's hard to find more data beyond my anecdata -- an EdWeek article I found reported that less than 50% of schools report world language enrollment data.
Also, the Europeans who learn three or four languages in school also have the luxury to learn those languages for free* through public schooling, so I'm not sure I understand your point.
I am sure that your implication that every American kid can get a quality free foreign language skill in school is false: just like almost every single other educational outcome in the US, it's generally great in the good (wealthy, suburban) schools and terrible in the bad (poor, rural or urban) schools.
Some of that is the accident of geography: it simply wasn't necessary. Today, we are more connected to our Spanish-speaking neighbors, and the value of learning that language is becoming increasingly obvious. I don't know whether the schools are doing a better job of stressing that than they did when I was in school.
I have indeed chosen to learn other languages, several of them. I wish I'd done it in school, at a time when my brain was more open to it. Unfortunately, that was also a time when I didn't know very much and put my priority on other things that ended up making less of a difference in my life.
I thought it was somewhat delayed, not paid, yet.
The surprising thing for me is Germany having 2. Seems unlikely.
Germany is big. I've heard that the proficiency in foreign languages tends to decrease as your country gets bigger. Because the bigger the country, the less likely you are to interact with foreign languages. Bigger countries also tend to have foreign works translated (or dubbed) into their own language more often.
So, no, I'm not surprised.
But they need to be able to get citizenship afaik... So basically everyone can speak two languages on paper, though their knowledge of the native one is extremely rudimentary
You're also required to learn 2 foreign languages in school if you want to go to university
I recall something like a 2.2 average.
I rarely met someone who could speak four languages fluently.
It's not normal.
Average number of languages spoken: France: 1.8 , Germany: 2.0 , Spain: 1.7 , Portugal: 1.6 , Italy: 1.8 , Greece: 1.8 , Poland: 1.8 , Sweden: 2.5 , Finland: 2.6 , UK: 1.6
French people can usually speak basic English, and a third language is common if that person has ties with another country but that's it. At school, we are normally taught two foreign languages. The first one is usually English, few people actually practice their second one.
The situation is completely different in Scandinavian countries. And it is indeed quite normal to speak 4 languages in Finland (usually Finnish, Swedish, English and a 4th one, often German). Because their native language is only spoken by a few, foreign languages are a necessity for international relationships. And as a Finnish friend told me, learning new languages is a popular way to pass time during long winter nights.
That's the best part in all of this. I'm not convinced by the claim the authors repeatedly make, that this technique will translate well to real-world problems. But I'm hoping that there is going to be more of this kind of result, singalling a shift away from Big Data and huge compute and towards well-designed and efficient algorithms.
In fact, I kind of expect it. The harder it gets to do the kind of machine learning that only large groups like DeepMind and OpenAI can do, the more smaller teams will push the other way and find ways to keep making progress cheaply and efficiently.
But what's the fun in that?
10,000 hands in an interesting number. If you search the poker forums, you'll see this is the number you'll see people throw out there for how many hands you need to see before you can analyze your play. You then make adjustments and see another 10,000 hands before you can assess those changes.
In 2019, it's impractical to adapt as a competitive player in live poker. A grinder can see 10,000 hands within a day. The live poker room took 12 days. Another characteristic of online poker is that players can also use data to their advantage.
So, I wouldn't consider 10K hands as long term, even if this was a period of 12 days. Once players get a chance to adapt, then they'll increase their rate of wins against a bot. Once you have a history of hand histories being shared, then it's all over. And again, give these players their own software tools.
Remember that one of the most exciting events in online poker was the run of isildur1. That run was put to rest when he went bust against players who had studied thousands of his hand histories.
This doesn't take away from the development of the bot. If we learn something from it, then all good.
For any number of hands, my money is on the bot.
Any good player will have their play analyzed and responded to, so there's a feedback loop there - any good player will have their play analyzed, exploited and will have to re-adjust their strategy to respond to exploitative play. The question is: How does the AI strategy adapt over time to players who know the hand history of the AI strategy. That's an extremely important part of being a top level player. To give you an example - if you watch Daniel Negreanu's vlog about his time at the WSOP he actively talks about changing his strategy in response to his analysis of different players' profiles. This is especially important in Sit & Go where at high stakes you'll have regular grinders who build up reputations - less so in tournaments where you're less likely to meet any given player.
Brown and Sandholm's algorithm aims to play a Nash Equilibrium which by deifnition _cannot_ be exploited by a single opponent player as long as all players are playing the equilibrium strategy. As they note in the paper this gives you a strong optimality guarantee in the 2-player setting. It was unclear whether this would transfer to real-world winnings in the multi-player case, and while it looks like it does for now (for current strategy-profiles of human players) humans might be able to adapt to the strategy played by the bot. Given the fact that the bot wins against current human strategy-profiles in the n-player setting, it's likely (but not a sure thing) that human players will have to team-up against the bot to exploit it. That seems rather unlikely to me.
If you read the paper/facebook post[0] (no idea why this worse article is the link here) - you'll see they address this.
>Although poker is a game of skill, there is an extremely large luck component as well. It is common for top professionals to lose money even over the course of 10,000 hands of poker simply because of bad luck. To reduce the role of luck, we used a version of the AIVAT variance reduction algorithm, which applies a baseline estimate of the value of each situation to reduce variance while still keeping the samples unbiased. For example, if the bot is dealt a really strong hand, AIVAT will subtract a baseline value from its winnings to counter the good luck. This adjustment allowed us to achieve statistically significant results with roughly 10x fewer hands than would normally be needed.
0. https://ai.facebook.com/blog/pluribus-first-ai-to-beat-pros-...
Perhaps more famously, Jungleman compiled hand histories from many different people while he was playing Tom Dwan in the 'durrrr' challenge (which I guess technically isn't over....)
My personal belief is that the "no-donk" strategy is an adaptation by fallible human minds to reduce the branching on the decision tree to something tractable.
Another good example is varying continuation betting sizes. A true GTO strategy would mix in a number of different sizings (and I'm sure the bots adapted to do this), but you only sacrifice a very tiny amount of EV by basically betting the same size every time. Doing the latter limits humans risk for making errors which is far more valuable than squeezing out .05bb/100 more by varying the sizes.
That said, it did train by playing against itself (before the experiment against the humans began).
t. former poker pro
Is the bot going for game-theory-optimal play, or trying to exploit weaknesses in other players?
It's going for game-theory-optimal play. It doesn't adapt to its opponents' observed weaknesses. But I think it's cool to show that you don't need to adapt to opponent weaknesses to win at poker at the highest levels. You just need to not have any weaknesses yourself.
that may be true for limit poker, but in a no-limit tournament the best this bot could do is not lose. as the pressure increases with the blinds and the players are forced to bluff and call bluffs how does this bot avoid folding itself to death from a run of bad cards?
I could see this bot doing well at cashing but I don't see how it could consistently place 1st the way the top human players do.
In the fb article linked above.
Since tournaments don't often spend much time with stacks much deeper than 100bb, I would guess that tournaments would be more easily solved. Though tournaments are much more frequently run with 9-10 players rather than 6 at a table.
https://www.cardplayer.com/poker-news/18226-explain-poker-li...
For example, game theory may tell you that in a particular situation, you can't be exploited if you bluff 10% of the time. If the opponent bluffs less than that, you can come out ahead by more often folding when he bets. If the opponent bluffs more than 10%, you can call or reraise when he bets. But if he bluffs the optimal amount, it doesn't matter either way, you can't take advantage of him.
So this bot would bluff at 10% to avoid getting exploited, but wouldn't try to detect whether the opponent is exploitable. (The latter is risky since a crafty opponent can switch up strategies, manipulating you into playing an exploitable strategy.)
If you want your perceived range to be balanced and make x play 50% of the time and y play the other 50%, you look at the watch and if the second hand is in the first 30 seconds, you make x play, 30-60 seconds, y play.
That's just an example but your point is 100% accurate.
Or to turn this around: given enough bots, some bots will place 1st a lot more than others. It’s just unclear which one.
There's a poker strategy we might call 'deterministically' optimal play, which consists of precisely assessing each hand's expected value with little to no bluffing. This is already common in online cash games with both bots and players running multiple games at once. And you're right - it's excellent at running net-positive and not losing, but unlikely to win significant tournaments.
Pluribus, though, is playing something close to game-theoretically optimal poker. In playing against itself, it's attempting to develop a takes-all-comers strategy with no exploitable weaknesses. That includes bluffing and calling bluffs - the goal is simply to find a mixed-strategy equilibrium where those moves are made some percentage of the time, in proportion to their expected payoffs. This can involve doing all of the same basic operations as pro players, like valuing button raises differently than donks or attempting to bluff based on how many players remain in the hand. The distinctive limitation is simply that Pluribus plays 'locally' optimal poker with no conception of opponent's identities or behavior in prior hands.
I could see this being an effective strategy in a WSOP, that ability to perfectly forget the previous hand is probably more valuable than anything the way WSOP champions play. I could see it coming down to whether or not the ability to exploit a reliable tell during a pivotal hand matters more than 10% of the time.
Still well, I suspect, since straightforward theoretically-correct poker will take money off the amateurs efficiently. But it seems possible that playing to wipe out weaker or less consistent players could provide enough margin to bully the more stable AI player.
And it's not like in the movies where if you don't have the money to call a bet, you lose. You simply are considered all in for the main pot and then sidepots that you aren't eligible to win will be created for any bets you can't cover.
>A superhuman poker-playing bot called Pluribus has beaten top human professionals at six-player no-limit Texas hold’em poker...
In fact, the methods we use are designed from the ground up to minimize exploitability. That's a really important property to have for an AI system that is actually deployed in the real world.
All the top high stakes players already have solvers that they've spent a lot of money developing and studying privately. They would definitely be upset with you, but by releasing the code you are democratizing the information to all the midstakes pros who want to study but don't have the resources to pay developers and solve the game privately.
b. How much more efficient is the improved search algorithm? the $150 number sounds like a couple of order of magnitudes..
b. It really depends on the game and the situation. It can be several orders of magnitude in six-player poker. In other games, it can be even more.
https://upswingpoker.com/isaac-haxton-pokerstars-partypoker/
I'm not sure the really provides strong opposition to the GP's claim.
By not releasing it, you're ensuring a higher concentration of money in the hands of a few, IMO.
Anyone with access to this source code could run a bot themselves, or employ someone to do so.
Plus, if you've accomplished this, no doubt someone can replicate it.
This is psychology.
You only need to know the probability of the opponent folding so that you can deviate from the theoretical optimal strategy to win even more money if they are a biased player
Of course, a key part of bluffing is getting the probabilities right. You can't always bluff and you can't never bluff, because that would make you too predictable. But our self-play and search algorithms are designed to get those probabilities right.
This makes no sense. If I am betting for thin value with a weak hand, then I make less money when my opponent folds. Does the bot not know whether it is bluffing or value betting?
Value betting and bluffing aren’t defined by the outcome of a hand — action yet to be completed. Poker is a game of hidden information so betting with “thin value” implies that your component of bluffing is larger. You want your opponent to fold more often than not when you have thin value because more often than not you’re actually beat.
QQ can get KK to fold based upon board texture, street, and prior action. But you don’t know the other person is holding KK when you’re betting for “thin value” on the river.
No, that is simply not true. If I am betting for value, then I want my opponent to call no matter how weak I am or how thin it is.
> But you don’t know the other person is holding KK when you’re betting for “thin value” on the river.
Then it's a value bet. As you said, it's not defined by the outcome.
The bot doesn’t “know” whether it’s value betting or bluffing—it’s not a relevant question. The relevant question is whether to bet, and what amount, in order to maximize value of the particular hand it has, with reference to the board and opponent actions taken.
(This, in addition to what the other comments have said about there being spots where a bet can get better hands to fold with some probability AND get worse hands to call with some probability - see the chapter "The grey area between value betting and bluffing" in Applications of No Limit Hold Em)
I was the one who introduced the term "value betting" to the conversation, applied specifically to weak hands.
Players face mental fatigue and have so over-learned their existing strategies that it takes time to adapt new strategies and even more time for those new strategies to become second-nature.
It reminds me of sports in a way. Teams start running a new wrinkle of offense in the NFL like the wildcat and it takes a few seasons for teams to instinctively know how to play defense correctly against that option.
I feel like a lot of trained ML models have a lot laughable weaknesses, but perhaps they've been trained on every game they're well prepared for any tomfoolery.
I was also confused by the sample videos where everyone had $10K at the start of each of the demo hands. It was unclear to me if that just the simulation of the hands or actual game play. If everyone starts every hand with $10K, then the feat seems less strong as going all-in has less risk.
But yea the sample size is definitely too small imo; when tested the heads up version of the bot some years ago they had it play a bigger sample (50 or 100k iirc?).
The reason is simple: with table stakes, your maximum win for a hand is constrained by your own stack size.
1) What were the reasons for choosing 6-handed play (assuming logistical and costs)? It would be interesting to see how the bot’s strategy would differ in a full ring game. 2) Are there any plans to commercialize the bot as a tool for training human players?
2) I'm quite happy working on fundamental AI research and plan to continue in that direction.
>Regardless of which hand Pluribus is actually holding, it will first calculate how it would act with every possible hand .
Is this information used to form an idea of what other players might be holding based on how the other player acts and how closely that action matches Pluribus's 'what if' action?
iirc, the frequency of bets in that spot is roughly equivalent to the frequency of times you're definitely in front of your opponent in that particular spot, but not always with the hands that are beating your opponent.
The concept is called Game Theory Optimal (GTO) and it's pretty popular in higher stakes games.
Second, I haven't read everything, but I believe you are playing a cash-game and not tournament-style. Is that correct? If that is the case, any chance you will be doing a tourney-style version?
[For those who don't play, in cash, a dollar is a dollar. In Tourney play, the top 2 or 3 players get paid out, so all dollars are not equal, as your strategy changes when you have only a few chips left (avoid risky bets that would knock you out) or when you are chip leader (take risky bets when they are cheap to push around your opponents).]
Also, curious how much poker you folks play in the lab for "research".
Or top 2 or 3 thousand... depends on the tournament but it's usually the top 15% ish.
There's a cash game almost every night at the FBNY office! I don't usually play though -- I'm not nearly as good as the bot.
Is 10,000 hands really considered a good enough sample? Most people consider 100k hands w/ a 4bb winrate to be an acceptable other math aside. However, as your opponent and yourself play with equal skill, variance increases to the point where regs refuse to sit each other.
It is an interesting point that these are pros but their specialities are either tournament or heads up. The current 6 max pros are LLinusLove, Otb_RedBaron, TrueTeller.
We used AIVAT to reduce variance, which reduces the number of samples we need by roughly a factor of 10: https://poker.cs.ualberta.ca/publications/aaai18-burch-aivat...
FWIW, they did some variance reduction techniques that dramatically reduce the number of hands needed to be confident in your results, so the number of hands may be bigger than you think. e.g. the results of 10k HU hands have much higher variance than the results of 10k HU hands where everyone just collects their EV once they're all in.
Furthermore, Chris Ferguson, scumbag aside, is absolutely still a very good player by today's standards, and one way higher than the mean participant in a research experiment.
10,000 hands is an effective enough sample at a certain win rate and analysis of variance of play; the n-value alone is not enough to tell you if it was enough hands.
The desire to master this sort of game has inspired the development of entire branches of mathematics. Computers are better at maths than humans. They're less prone to hazardous cognitive biases (gambler's fallacy etc.) and can put on an excellent poker face.
As a layperson who's rather ignorant about both no-limit Texas hold 'em and applicable AI techniques, my intuition would tell me that super-human hold 'em should have been achieved before super-human Go. Apparently your software requires way less CPU power than AlphaGo/AlphaZero, which seems to support my hypothesis. What am I missing?
Bonus questions in case you have the time and inclination to oblige:
What does this mean for people who like to play on-line Poker for real money?
Could you recommend some literature (white papers/books/lecture series/whatever) to someone interested in writing an AI (running on potato-grade hardware) for a niche "draft and pass" card game (e.g. Sushi Go!) as a recreational programming exercise?
Fortunately, these techniques now work really really well for poker. It's now quite inexpensive to make a superhuman poker bot.
That's interesting, could you share an example? Most of my search results are anecdotal Reddit threads about how many people cheat in online chess.
From what I've read, they work by comparing the player's moves against chess engines, and if the player is picking engine's choice too often in positions where there are multiple roughly equal moves, they get flagged.
Not really. With perfect information you know the correct strict equity plays assuming normal opponents. This doesn't give you the ultimate answer, because a player's reads and inference about another player is definitely an input - especially at the highest level - but it is more than enough to give you a winning/losing player at the small/midstakes.
source: worked for an online poker company that had these tools... and far more available to us
I think player-dependent strategy is more important at lower levels because the players are much further away from what you call "normal opponents", so there's far more opportunity to exploit their mistakes.
That said, I suppose it would be possible for the bots to become so prevalent that all this sort of opportunity is effectively used up, so the return vs time and risk for a human player is no longer worthwhile. (That already happened long ago for most players, as the initial online poker boom faded and most casual players left.)
On the other hand, all the major platforms have terms prohibiting using bots, so their numbers might be sufficiently limited to prevent that scenario.
Reaction times ought to be one of the easiest things to fake. All it would take is a bunch of monitoring of large numbers of games to create a nice model of real player reaction times, which in all likelihood are normally distributed anyway.
Oh, right. I was thinking along the lines of 100m dash, where people often do have negative reaction times (which we penalize as false starts).
In poker we don't have much of an incentive to react instantly to any play.
You might be surprised by the lengths people go to in order to bypass bot-detection just for ordinary games. All of the things you mentioned are pretty standard. Considering there is serious money on the line here, I am positive that plenty of poker bots will be virtually indistinguishable from professional players, if they aren't already.
A successful bot shouldn't get caught for "playing like a bot" because the moment it's actions are that predictable it would presumably no longer be effective.
But it will get caught for operating like a bot. So, don't run it 24hrs a day. Sites also randomize things to keep bots at bay, even card imagery.
If your performance and success drops whenever they randomize something that gives the bot false inputs, then you might get caught.
Inputting all of the poker events manually would be really tedious I'd imagine.
Of course, if you're winning millions, they can interview you about your poker history and how you got so good.
It sounds like easy money, but probably not.
People have tried it and online poker sites know they've tried it, so they'll randomize images and other data. If you take a dive when the randomizations are triggered and outperform otherwise good luck trying to collect your winnings.
Edit: In fact, if we're talking worst case, circumventing their anti-bot restrictions would presumably be illegal under the CFAA. So if you're in the US you could even be charged criminally, although I expect in reality that would be less likely.
It would probably be possible to figure out the types of detection being performed by the poker sites and use adversarial training methods to train a machine learning solution to mimic human input patterns. Or, more pragmatically, have the bot analyse the state of the game and give orders for a human to perform at their own natural pace.
The problem would be if i was a pro i would rather run 1000 bots than play myself. Which means the only players left are AI and fish. Once the fishies learn of this fact, they will abandon in drove.
It's all gonna go back to live poker soon.
OP discussed it but while this is true, it is not necessarily true or straightforward when it comes to games with hidden information like poker. This is more of a game theoretical problem (Economics) than it is a purely mathematical one, which had less support in the AI/ML community, hence the delay.
The lower CPU/GPU/resource use supports that fact as does your intuition. Breaking poker required a lot of manual work and model design over brute force algorithms and reinforcement learning.
- Are the action and information abstraction procedures hand-engineered or learned in some manner?
- How does it decide how many bets to consider in a particular situation?
- Is there anything interesting going on with how the strategy is compressed in memory?
- How do you decide in the first betting round if a bet is far enough off-tree that online search is needed?
- When searching beyond leaf nodes, how did you choose how far to bias the strategies toward calling, raising, and folding?
- After it calculates how it would act with every possible hand, how does it use that to balance its strategy while taking into account the hand it is actually holding?
- In general, how much do these kind of engineering details and hyperparameters matter to your results and to the efficiency of training? How much time did you spend on this? Roughly how many lines of code are important for making this work?
- Why does this training method work so well on CPUs vs GPUs? Do you think there are any lessons here that might improve training efficiency for 2-player perfect-information systems such as AlphaZero?
- Are the action and information abstraction procedures hand-engineered or learned in some manner?
- How does it decide how many bets to consider in a particular situation?
The information abstraction is determined by k-means clustering on certain features. There wasn't much thought put into the action abstraction because it turns out the exact sizes you use don't matter that much as long as the bot has enough options to choose from. We basically just did 0.25x pot, 0.5x pot, 1x pot, etc. The number of sizes varied depending on the situation.
- Is there anything interesting going on with how the strategy is compressed in memory?
Nope.
- How do you decide in the first betting round if a bet is far enough off-tree that online search is needed?
We set a threshold at $100.
- When searching beyond leaf nodes, how did you choose how far to bias the strategies toward calling, raising, and folding?
In each case, we multiplied by the biased action's probability by a factor of 5 and renormalized. In theory it doesn't really matter what the factor is.
- After it calculates how it would act with every possible hand, how does it use that to balance its strategy while taking into account the hand it is actually holding?
This comes out naturally from our use of Linear Counterfactual Regret Minimization in the search space. It's covered in more detail in the supplementary material
- In general, how much do these kind of engineering details and hyperparameters matter to your results and to the efficiency of training? How much time did you spend on this? Roughly how many lines of code are important for making this work?
I think it's all pretty robust to the choice of parameters, but we didn't do extensive testing to see. While these bots are quite easy to train, the variance is so high in poker that getting meaningful experimental results is relatively quite computationally expensive.
- Why does this training method work so well on CPUs vs GPUs? Do you think there are any lessons here that might improve training efficiency for 2-player perfect-information systems such as AlphaZero?
I think the key is that the search algorithm is picking up so much of the slack that we don't really need to train an amazing precomputed strategy. If we weren't using search, it would probably be infeasible to generate a strong 6-player poker AI. Search was also critical for previous AI benchmark victories like chess and Go.
Did you guys set any rules as to whether or not members of the team that worked on this are allowed to use it?
[1] https://www.cigaraficionado.com/index.php/article/robotic-po...
It feels like a more down to earth version of the sci fi super human running impossible differential equations to predict exactly what you will do given knowledge that he knows what you know what he knows... etc. ad Infinitum. But since it doesn’t actually consider the person it’s predicting, it may simply be a really really good approximation of the game theoretic dominant strategy.
At what complexity of game and hidden information should we feel like the bot can’t win by running a lookup table?
There is a cliche about how poker is about playing your opponents and not the cards. Is this AI is only focusing on its cards and ignoring its opponents?
Also I would be curious to see how it performs against people that aren't "elite human pros". Would this AI win at a higher rate in a game against average recreational players compared to the rate a pro would win?
Lastly it is also possible that the pros simply didn't have enough time to adapt to the AI which would be extra important considering the AI plays unlike humans and therefore is harder to predict.
We played 10,000 hands over 12 days in the 5 humans + 1 AI experiment. That's quite a long time, and there's no indication that they even began to uncover any weaknesses in that time period. So I'm fairly confident the AI is robust to exploitation, and I think that's a very important quality to have in any AI system.
Out of curiosity, how does a bot deal with oddities things like this?
While this particular bot may not have those programmed in, a more powerful variant eventually will.
None of this is meant to diminish what you all accomplished, I'm just highlighting areas of poker in which this AI would be less successful than humans even if it is more successful overall.
You can't out-think or adapt to a rock-paper-scissors opponent who selects at random. All you can do is also select at random and accept that the two of you have even odds.
In the case of poker, it appears that adaptability is not as good as pure mathematical optimization. Humans can adapt their strategy, but it’s basically just worse regardless because this thing has cracked the code.
I’m surprised that you managed to beat pros without adaptability. It’s pretty impressive and says a lot about how we define strategy. If human adaptability is just not as good as machine optimality across all games, we could imagine discovering that an adaptable poker AI can’t outperform this one. It raises a whole lot of interesting questions because lots of criticism towards something like Starcraft AI is that it is strategically stupid and doesn’t adapt. Now the Starcraft Ai is admittedly kind of stupid now, but we may hit a wall on its creativity simply because creativity is, despite human intuition, a dumb idea.
Because of how Poker is not sub-game solveable (it is not possible to self-locate within the tree), this bot's play has to get into its opponent's mindspace in a sense. To not be exploitable, it essentially has to infer the other player(s) hidden state and paths from observed actions. This isn't something I've seen in Dota, Starcraft, Chess, Go bots.
It's true that it doesn't learn online to find exploitable patterns of other players, but doing this without also making yourself exploitable in turn is a very difficult other problem. Low exploitable near optimal play according to game theoretic notions is considered strategy.
While you're correct that online learning is powerful and something machines are not currently good at (in complex spaces), you can avoid being exploited without learning if your experience is rich enough and you know how infer what your opponent is trying to do and anticipate them. I'd argue this lineage of poker bots are the closest to playing that way of the major game playing bots.
Adaptability is beaten by perfect strategic play in games with clear victory conditions.
My familiarity with optimal control theory is nil but Kydland (1977) applied it to monetary policy to show that the right rules dominate discretion. What the right rules are for monetary policy is still an open question though, because while the victory conditions in economic policy are clearly defined the surrounding environment is very far from static so you deal with out of training set data regularly. Once AI can deal with these kind of out of context problems it seems plausible GAI is a matter of time.
http://www.finnkydland.com/papers/Rules%20Rather%20than%20Di...
> Rules Rather than Discretion: The Inconsistency of Optimal Plans
> Even if there is an agreed-upon, fixed social objective function and policymakers know the timing and magnitude of the effects of their actions, discretionary policy, namely, the selection of that decision which is best, given the current situation and a correct evaluation of the end- of-period position, does not result in the social objective function being maximized. The reason for this apparent paradox is that economic planning is not a game against nature but, rather, a game against rational economic agents. We conclude that there is no way control theory can be made applicable to economic planning when expectations are rational.
Is this true in any meaningful sense?
For heavily studied games there's usually a theoretically optimal play independent of the opponent's interior state, this is obviously true for all the "Solved" games, which includes the simpler Heads Up Limit Hold 'Em poker (solved by Alberta's Cepheus project) but it seem pretty clearly true for as-yet unsolved games like Go and Chess too.
I'm very impressed by this achievement because I had expected good multi-player poker AI (as opposed to simple colluding bots found online making money today) to be some years away. But I would not expect "adaptability" to ever be a sensible way forward for winning a single strategy game.
I would think if professional players are utilising this information, a bot could benefit from it. I don't see how they would ever lose out from this information, even if it only uses situations where the opponent has a history of 100% of the time responding a certain way.
I am impressed by the bot but I have to laugh a bit because years ago I joked with a friend about making an "amnesiac bot" that had no recollection of previous hands, it seemed so useless we obviously didn't make it, we've evidently been proven wrong. (pointless tangent there)
The theoretically optimal play just skips that meta and meta-meta play and performs optimally anyway. Because poker involves chance the optimal play will be stochastic and so you can stare at the noise and think you see a pattern, that just means you'll play worse against it, because you're trying to beat a ghost.
For example, suppose in a certain situation optimally I should raise $50 10% of the time. It so happens, by chance, that I do so twice in a row, and you, the note-taker, record that I "always" raise $50 here. Bzzt, 90% of the time your note will be wrong next time.
1. Coming up with an unexploitable strategy, then scaling it up by playing as many hands as you can, earning the slim expected value each time.
2. Picking a good table / card room / 'scene', and then trying to extract as much value from it as possible.
You most often see 1 online, and 2 live, for obvious reasons.
A skilled human would be a lot more successful, I believe, than a bot in case 2. For 2, important skills are:
1. Be entertaining. You have to play in a way that is entertaining to those playing with you, such that they want to continue playing with you (and losing money to you). Good opponents (i.e. that are bad at poker but want to play high stakes) are hard to find, it is vital that you retain them.
2. Cultivate a table image, then exploit it. Especially important for tournament play, where you have the concept of "key hands" that you really need to win to potentially win the tournament. With the right table image, you may be able to win hands you otherwise wouldn't have won.
3. Exploit the specifics of the players you are playing against. Yes, that also makes you exploitable, but the idea is to stay one step ahead of your opponents.
Furthermore, you can kind of account for such players by including more random or aggressive profiles in the inference/search stage.
Now say I have thousands of hands viewed against you, and you raise pre-flop 50% of the time. That is pretty significant information about the types of hands you play. If I have only 10 hands I've observed, that same stat means nothing.
The theoretical optimal play depends on who you're playing, as more value could be extracted in certain situations vs certain players.
For example, if I've seen you face a pre-flop 3-bet 1000 times and you've folded 99% of the time. That would be a good opportunity to recognise that 3-bet bluffing this player more often would have value, and be a more optimal play than some default. Contrast playing someone who called pre-flop 3-bets 75% of the time it wouldn't be optimal to 3 bet bluff here. Different opponents, different optimal plays.
I don't think you play very much, which is fine, but makes this discussion a bit pointless.
That said, for this bot, I wouldn't say it's playing completely independent of the other players's interior state. Pluribus must infer its opponents strategy profile and according to the paper, maintains a distribution over possible hole cards and updates its belief according to observed actions. This is part of playing in a minimally exploitable way in such a large space for an imperfect information game.
This is what interests me. It doesn’t do this. In fact because it played against itself only, it is should be assumed that the only strategy profile it considers is its own.
In an n-player game, a table can be in a (perhaps unstable) equilibrium which the "optimal" strategy will lose at. This has been demonstrated for something as simple as iterated prisoners' dilemma (tit-for-tat is "best" for most populations, but there are populations that a tit-for-tat player will lose to). I don't play poker but I've definitely experienced that in (riichi) mahjong - if you play for high-value hands the way you would in a pro league, on a table where the other three players are going for the fastest hands possible, you will likely lose.
Interesting is up to you, but effective is definitely wrong.
ICM-perfect bots crush small tournaments, which do not take into account opponent behavior - merely modeling the gamestate. The faster the blinds and the smaller the stacks, the better, but even normal structures get killed by these so-called "expected value" only bots.
Game Theory Optimal (GTO) attacks are incredibly effective at all levels of the game. The AI need not incorporate opponent feedback to be a winner. It can make it better, but it is not at all required.
I have a few basic questions. I would like to implement my own generic game bot (discrete states). Are there any universal approaches? Is MCMC sampling good enough to start? My initial idea was to do importance sampling on some utility/score function.
Also, I am looking into poker game solvers - what would be a good place to start? What's the simplest algorithm?
Thanks
It depends on context. 4.8bb/100 is quite good for high-level online play, but wouldn't be enough to make a living at live poker. The biggest game that runs on a regular basis in most areas is $5/10. At ~33 hands per hour, that's 1.6bb or $16 an hour.
And I'd assume there was no rake in your game? That would take a big chunk out of the rate.
$16/hr X (VM|microservice thread) could become astronomical profit.
To answer your question, no, I don't think human players would play at their best when not playing for actual money.
What are your thoughts on a poker tournament for bots? Do you think it could turn into a successful product? I've always wanted to build an online poker/chess game that was designed from the ground up for bots (everything being accessible through an API), but have always worried that someone with more computational resources or the best bot would win consistently. Is it an idea you've thought about?
Another person asked "What took you so long?", and i had the same question. :) I really thought this milestone would be achieved fairly soon after i left the field in 2007. However, breakthroughs require a researcher with the right amount of reflectiveness, insight, and determination.
Well done.
The training aspect has some improvements but is at its core similar to Libratus. The search algorithm is the biggest difference.
There aren't that many great resources out there for helping new people get caught up to speed on this area. That's something we hope to fix in the future. Maybe this would be a good place to start? http://modelai.gettysburg.edu/2013/cfr/cfr.pdf
With one AI and multiple professional human players sitting at a physical table, the humans outperform the probabilistic model because they take advantage of each other's mistakes/styles. Some players crash out faster but the winner gets ahead of the safe probabilistic style of play.
So this bot is better at the current professional player meta than the current players. In a 1v1 against a probabilistic model, it would probably also lose?
Am I understanding this properly? Or is playing the probabilistic model directly enough of a tell that it's also losing strategy? Meaning you need some variation of strategies, strategy detection, or knowledge of the meta to win?
The bot played like 10 000 hands. There is no way that is enough to prove it's better or worse than the opponents.
More so in no-limit where some key all-ins can turn the game up side down. The variance is higher than limit or fixed, right?
I did a heads up Texas holdem fixed bot with "counter factual regret minimization" like 8 years ago from a paper I read. It had to play like 100 000 hands vs a crappy reference bot to prove it was better.
Strategy detection in so short games is probably worthless.
The edge is probably in seeing who are tired or drunk in paper poker.
> Although poker is a game of skill, there is an extremely large luck component as well. It is common for top professionals to lose money even over the course of 10,000 hands of poker simply because of bad luck. To reduce the role of luck, we used a version of the AIVAT[1] variance reduction algorithm, which applies a baseline estimate of the value of each situation to reduce variance while still keeping the samples unbiased. For example, if the bot is dealt a really strong hand, AIVAT will subtract a baseline value from its winnings to counter the good luck. This adjustment allowed us to achieve statistically significant results with roughly 10x fewer hands than would normally be needed.
I'm thinking in particular of unbalanced tables with an ever-changing mixture of TAG and LAG play. I've changed my mind three times about whether that's humans' best refuge -- or a situation that's a bot's dream.
You've done the work. Insights welcome.
I see on chess channels that grand masters have to rethink their whole game preparation methodology to cope with the "Alpha Zero" oddities that have now been introduced into this ancient game. They literally have to "throw out the book" of standard openings and middle games and start afresh.
I would say that it's thoroughly rebounded to play the game not the player in poker and this isn't because of super bots like the one used in this paper.
Ever since game theory invaded poker players that play in highly visible events such as tv tournaments try as hard as possible to make their game unexploitable.
However, even according to the former world champion (Viswanathan Anand) the run he's been on is something quite shocking: “His results this year is simply [great].... difficult to find words. [It’s been] completely off the charts. I think the chess world is still in a bit of a shock. The rest of the players are struggling to deal with a phenomenon [like him]. Even in 2012-13, his domination was less than it is this year. Everyone is still processing this information.” [1]
Carlsen is basically on route to breaking 2900 Elo - at 2882 Elo with a clear upwards trend - while there's only two other active players even above 2800 Elo and struggling to keep it above that treshold. (Elo is the rating system used in chess. Above 1500 Elo is an average player, 2000 Elo is a good player, 2500 Elo is a grandmaster. Anything above 2700 Elo is basically godlike.)
Oddly enough, instead of playing more like a machine, it seems like Carlsen has been playing chess that is much more about the human aspect of the game rather than trying to find the top ranked engine move on every turn. (The current traditional top engine - Stockfish - makes an assumption of each move's validity using a point system, which the chess world has been more or less obsessing over for the past decade. Alpha Zero doesn't have such a point system whatsoever.) He's been playing a drastically more aggressive and dynamic variety of chess compared to what has been seen in a long time at the top tournaments.
He's been playing to create dizzying positions on the board, making a few moves that aren't necessarily liked by the traditional top engines, but still finding himself in a winning position several moves after. It definitely looks like some sort of black magic, but it seems like the big thing Alpha Zero has brought to the general philosophy on how to approach chess at the top level is that it's possible to play aggressive chess, take risks and win in 2019. Magnus Carlsen is the first player to successfully reinvent that style of play, more than likely partly inspired by Alpha Zero. So, I'd say the big thing about Alpha Zero isn't necessarily that it could beat the other top engines, but more importantly that the 'artistic' aspect of its play is something that has never been seen from another chess engine. The fact that it proved that sort of style superior to the play ever before played by another chess engine is just the icing on the cake.
Garry Kasparov on Alpha Zero's chess persona: "I admit that I was pleased to see that AlphaZero had a dynamic, open style like my own. The conventional wisdom was that machines would approach perfection with endless dry maneuvering, usually leading to drawn games. But in my observation, AlphaZero prioritizes piece activity over material, preferring positions that to my eye looked risky and aggressive." [2]
[1] https://sportstar.thehindu.com/chess/viswanathan-anand-on-ma... [2] https://science.sciencemag.org/content/362/6419/1087
For reference, I once folded top boat to quads (he showed) to a river all in raise in PLO to a dude who had a 100% win showdown when raise river stat over several thousand hands. Other stats confirmed he was a nit, so it was an easy fold. Iirc, this was PLO 200 or PLO 400 — I never saw anyone that nitty at the PLO 1000 or PLO 2000 tables.
FWIW, I did a lot more than “look for an opening”, although I did a lot of that. I tried to play GTO as much as possible, but I would adjust to people who were exploitable when they called too much, folded too much, or were too aggressive into weakness.
I spent a lot of time away from the table analyzing stats of the regulars to find leaks to exploit. It was worth the time, and it made it much easier to play 8-12 tables of PLO.
No longer online. I quit with UIGEA. I hate Bill Frist.
> I used to railbird FT back then, watching patrick antonius/sahamies/durr/jungleman etc. exciting times.
I never played with those folks. For reference (and maybe I was vague, and maybe I show my age), plo 200 is 1/2 blinds plo, and 2000 plo is 10/20 blinds plo. My heyday was when 10/20 blinds were the max. When they created the 50/100 and higher limits, they killed the 10/20 blind (former max) games. I never played over 25/50, because I didn’t consider myself properly rolled for it. That said, in retrospect, I should have gone modified Kelly criterion and take shots at the higher stakes — some of those dudes were total donks.
I did play against huck seed on FT (nit), and I played against Doyle and Todd on Players Only (I think their room was a skin). They also played tight, but they may have been doing required hours. People donated to them religiously with light calls.
> How many hours did you put in for you to be good?
I would argue that I am still not “good”. There’s a hierarchy in the poker world, and you don’t feel comfortable until you’re at the top — and even that is fleeting.
To answer your question, though, I was profitable at 5/10 and 10/20 after maybe 1000 table hours (usually 4-8 tables) and 500-700 study hours. Note that this was when poker was super soft, and note that I am a specialist in learning (degrees, experience, and whatnot), so I learn things like new games more quickly than most people. My job lent itself to a lot of study away from the table, so I availed myself of that time.
I remember several breakthroughs for my game:
1. I had a dream one night in which I finally understood the bidirectionality of plo8. This was the game that I built my bankroll on (after cashing in a few free rolls). That took me to plo8 100 ($0.50/$1 blinds) in short order. After that, I just grinded to 200 and 400 plo and plo8.
2. I remember getting crushed by a LAG player in plo 400 one night. I went to four plo 50 tables and played 12 hours straight playing with a 55% or so VPIP. I broke even in that session, but it helped me understand LAG players a lot better. In retrospect, that session helped me understand how to exploit LAGs really well, and that paid off a lot at higher stakes. It also helped my SLAG game a lot.
3. The next big leap was realizing that there were three lines to exploit in poker — players who are too weak (fold too much), too passive (call to much), and too aggressive (bet/raise too much). Being able to exploit these tendencies is optimal. Being able to induce these tendencies is insanely profitable. The above is easy to say, but not always so easy to do.
4. The last phase of my development was understanding “gears”. Changing gears is the ability to switch between being passive/aggressive and tight/loose depending on the context. Most people change gears predictably — for example, if they lose a big hand, they tighten up (or some players loosen up). I played my best when I was able to adjust to the texture of the game and play the way that my opponents least expected me to play and/or wanted me to play. It’s a lot of psychology, but when I mastered this, I felt like I owned the table. No one could read me, and I read them like an open book. This is the high that skilled poker players live for, imho.
To close, I twice considered becoming a pro poker player. Once before UIGEA, and once after.
Before UIGEA, I didn’t because I realized that I was only good for about 20 top notch hours per week, and I could play those hours after work. Furthermore, the tables were only juicy for maybe 30 non-consecutive hours, so I didn’t feel like i was missing much. I was also worried about the non-legalization of poker in the US, so I wanted to keep my day job.
After UIGEA, I thought about moving to Thailand or Canada, but I (rightly) thought that games would get much worse without the US market. My earn would have been a solid $100-200k based on some of my former peers, but that’s not terribly exciting money for me. Anyone who can make $100k or more in online poker can make way more than that by being a programmer or by doing some sort of tech business (SaaS, e-commerce, consulting, etc.) or financier.
Ok, that’s a wall of text. Feel free to ask follow ups.
I also played against Mike the Mouth (either party or FT). I think that this was when they limited the 5/10 and 10/20 games to two tables each.
I played Mike in both plo8 and plo, and he was supposed to be a specialist. He was an absolute donator in the games I played in. He took really bad lines, and he was a net loser over a statistically insignificant number of hands. That said, if he was at the table, I wanted to play, and I wasn’t leaving until he got up. He was very exploitable.
To be fair, I don’t know what his life situation was like at that time (it was up and down from what I heard). That said, I wanted him at my table 100% of the time.
I have a question about the conspiracy. For the 5 Human + 1 AI setting, since the human pros know which player is AI (read from your previous response), is it possible for human players to conspire to beat the AI? And in theory, for multi-player game, even the AI plays at the best strategy, is it still possible to be beat by conspiracy of other players?
Thanks.
Will it just become increasingly sophisticated bots playing each other online?
The fact that there's a published recipe for a superhuman bot that can be trained for $150 and run on any desktop computer sounds like an existential threat to their business.
The main mitigating factor I can think of is that you'd need to also adversarially train it so it isn't distinguishable from a skilled human. But that doesn't seem like it would be too difficult.
I'm sure the sites have been crawling with bots as long as they've been around, some better than others. As long as it doesn't drive away too many customers I doubt the sites care. They still take a rake on bot games. However better AI could change that as the "dumb money" slowly dries up.
There are a bunch of such threads over the years where through statistical analysis, users have identified groups of dozens of bots.
While years ago many of the pros could theoretically beat these bots, it may not have been by enough of a factor to overcome the rake. Of course if the bots are practicing any game selection they can take money out of the economy even if they can't beat pros.
Anti-bot measures is an arms race and the sites aren't always ahead of the game.
Thinking about online poker again gives me ideas now that I actually know how to program. I actually thought up and wrote out a good way to subtly steal money from people, but I'm deleting it because I don't want someone else to do it. (And I wouldn't do it myself because I have ethics.)
OpenAI Five beat the world champions in back-to-back games...
Cool achievement, but hollow marketing doesn't make it better.
Was it online? the picture on the article seems to imply IRL.
If IRL, what inputs did it have, simply cards shown or could it read tells? Did those players know they were playing an AI?
"$50,000 was divided among the human participants based on their performance to incentivize them to play their best. Each player was guaranteed a minimum of $0.40 per hand for participating, but this could increase to as much as $1.60 per hand based on performance."
So the humans weren't betting their own money, but they still made more money if they won.
Maybe a bot technically qualifies as an opponent in durrr's challenge [0]? :)
How would bluffing influence the outcome? Both these players who are considered very strong, are known to play all kinds of hands.
[0] - https://en.wikipedia.org/wiki/Tom_Dwan#Full_Tilt_Poker_Durrr...
Are we now saying that a computer can do this all in simulation? if so, it's a great break through in human history.
Poker is about exploitative play against people who base their play off emotions, and unperfect game theory optimal against players who don't base their play off emotions. The more perfect the GTO play is, the higher the winrate against the latter group, but higher stakes games are built around one or more bad players - pros will literally stop playing as soon as the fish busts.
Chris Moneymaker got some damn good hands. Its part of the game. Its why this feat is unremarkable and why poker is a crap game for AI. The outcomes are very loose, especially when the reason these guys are pros is partially because of their ability to read.
You are taking away a tool that made their proker players great and then expect them to be a metric to test the AI. A better test would be to have pro players play a set of 1, 2, 4, 7 basic rule bots and the AI does the same. Then you compare differences in play. With enough data points you can compare situations that are similar but the AI did better or worse. This is a fair comparison of skill.
Also if there are professional players at a multiplayer game the AI is getting help from other players. Just like Civ V I get help from the AI attacking itself. Im sure this AI got help from the players attacking eachother (especially if they were doing so and making the pot bigger for the AI to grab up, think of a player reraising another player after the bot does a check all in).
That's not luck. See also: https://news.ycombinator.com/item?id=20416099
Also, Chris Moneymaker is a good poker player. He's no Phil Ivey or Tom Dwan, but he's still very good and has had decent results after his WSOP win.
More links for reference: https://ai.facebook.com/blog/pluribus-first-ai-to-beat-pros-... https://science.sciencemag.org/content/early/2019/07/10/scie...
The trick is how to create natural mouse click movements or keyboard inputs. This is the part that I'm most shaky on but the pokr.live API works by sending screenshots which it will translate into player actions at the table
disclaimer: pokr.live API is a WIP
It might pay better than a full time job.
That's exactly how the brain operates.
Poker is an incomplete information game with crushingly high variance. The bots strategy is likely not quantifiable.
> Pluribus disagrees with the folk wisdom that donk betting (starting a round by betting when one ended the previous betting round with a call) is a mistake; Pluribus does this far more often than professional humans do.
Is it overly simplistic to think that humans could improve their game by incorporating some strategies like this more/less often than they were previously?
Whys is using CFR better than training based on real data?
Also, as you add more players it becomes harder and harder to evaluate because the bot's involved in fewer hands, we need to have more pros at the table, and we need to coordinate more schedules. Six was logistically pretty tough already.
I have a question about the conspiracy. For the 5 Human + 1 AI setting, since the human pros know which player is AI (read from your previous response), is it possible for human players to conspire to beat the AI?
This is for 6-man games. The article mentions 10,000 hands - this is a very small sample size to draw any real conclusions, as anyone who has dabbled in online poker for more than a few thousand dollars can attest to. Regardless - it's trivial to write a bot that'll beat 90% of the players, as site runners can all attest to (bots are a serious problem that is not new). What does it matter that a bot can beat 'the best' or 'professionals'? It's enough that it can do better than the vast majority, outside of dystopian woes about robots taking over or being 'superior' to human beings.
Glossing over all that - I am curious if this can be used for something other than ruining online poker, which has largely already been ruined by allowing multi-tabling professionals with custom software that gathers statistics on players (data mining), existing bots, US government and irresponsible (criminal) site runners (looking at you ultimate bet)