The Steely, Headless King of Texas Hold ’Em
nytimes.com
nytimes.com
This article is filled with bold claims by people that want to sell the idea IMHO. I'm not buying it because even limit texas Hold'Em has never been solved mathematically by super computers, let alone a single machine. Limit Hold'Em is close to being solved but if you check out the last match of pros against a supercomputer you'll find that it's close, but the computer is not a clear winner: http://poker.cs.ualberta.ca/man-machine/Competitors/
I'm just not buying it. Can it beat the average player? Probably. Pros do it everyday. But can it beat a pro consistently and claim a clear victory over humans? To be determined.
Edit: I found a take from High Stakes Limit Pro Anthony Rivera on 2p2: http://forumserver.twoplustwo.com/showpost.php?p=29762774&po...
His stance is that it's probably a break-even situation where you are basically playing another pro, and that he'd rather be playing craps. I guess they really have made a great machine. :D
Moreover, if we could, the opponent may choose a different strategy. It may be like a game with nontransitive dice: if you figure out which dice I use, you can beat me, but if I figure out that you figured that out, I can pick different dice and beat you (http://en.wikipedia.org/wiki/Nontransitive_dice)
I think it is possible to write software that statistically beats any opponent, but proving that it does is way harder.
The definition of solvable is "Does there exist a strategy that, regardless of opponents play, is not a losing strategy" though (because the game is symmetrical). It is not possible to solve for "Does there exist a strategy that, regardless of opponents play, guarantees maximum profits". You can solve it for not losing money, but you cant solve it for making the max amount of money.
Trivially it's not possible to write software that beats any opponent (because what would it do playing against itself?) . Less trivially, any game that has a finite amount of decisions (and limit hold-em does) has at least one Nash equilibrium, so there exists a strategy that will at least have you break even.
The way to solve the game is to calculate your odds of winning based on previous actions and ensure that you take actions that make any future decision of the opponent have the same outcome (to reach a Nash equilibrium).
That strategy hasnt been calculated yet, but the best limit players are most likely playing very close to it, at least if you compare to the best no-limit players playing no-limit (the variable amounts possible to bid in no-limit multiplies the possible strategies massively).
Given two strategies, you can determine which one is expected to win and by how much. That allows you to find the strategy that's hardest to take advantage of. In the case of non-transitive dice, the hardest strategy to take advantage of is likely to be "pick each die with equal chance".
That being said, solving these games is extremely expensive. Super-exponential in the size of the game, if I remember correctly. Way worse than games like chess. See this paper: http://www.sciencedirect.com/science/article/pii/00220000849...
Let's say for all the possible hand/board combinations and betting patterns you evaluate all of them (with current computing power this is not possible) then you can take the line that maximizes your value assuming your opponent maximizes his. Then you both end up close to even over a long period of play.
But it's even more simple than that: Nash equilibrium strategy says that you have to bluff 1/3 of the time on the river in a certain spot and have a highly valued hand 2/3 of the time.
It doesn't matter whether your opponent calls you or not. If he calls you too much, your value hands get paid off more often. If he doesn't call you enough, he loses too much to your bluffs. If the bet is the size of the pot, then he has to call one half of the time as part of his Nash equilibrium strategy to not get exploited. If he deviates, he will lose to the bot in the long term.
And since this is limit it's actually very simple. Bet sizes in regards to the pot are tiny (1/5 to 1/10 of the pot) and there is no having to choose the bet size or allowing for players to overbet the pot (make a bet larger than the pot)
No limit is a much bigger challenge than limit for computers, because a small edge in limit is too hard to exploit since you'll get tired of exploiting a computer for pennies per hand
Is this even possible in a game of imperfect information?
Starting out you have 2 cards and you know your opponent has one of 2450 possible combinations of cards.
You know exactly how your 2 cards stack up in chances of winning against those 2450 combinations, and you know how much money is on the table initially. Each time an opponent has an action he can take one of two or three actions (call, raise or fold) - your job is to calculate a strategy that makes either of those decisions have the same expected value for your opponent. That is doable, but quite complex, because it requires you to consider not only what your opponent has, but also what possible hands you could have based on your previous actions.
The complexity is obviously enormous (hence why it hasnt been computed yet) but it is not impossible to compute by any definition.
http://www.econ.ohio-state.edu/jpeck/gametheory/gameL8.pdf
http://robotics.stanford.edu/~koller/Papers/Koller+Pfeffer:I...
It's a problem that is incredibly challenging.
The earlier part of the article (when it was full strength and in testing) has no mention of it being beaten.
I've thought long and hard on this, it's one of my life goals to create a piece of software that could mimic good players in texas hold-em, and I think I might have found a way to do just that using big data, hand histories, neural-networking, and a ton of input by actual players. Or at least a good start.
How would I accomplish this? Well a very high-level overview below. Basically it starts out extremely stupid and grows as a player:
1. Dump hand histories from pros into a large database. The number of hand histories would run into the billions. Users could dump online logs in bulk or create one-offs using input software.
2. Create an input system for a single user to choose a random hand-history and then classify it using tags. For instance tags might help categorize the player's style, the opponent's style, the "street" of play they are tagging, common name for the situation (Facing donk-bet on flop after raising pre-flop), etc. I would leave this fairly flexible and allow users to create new tags. Think similar to Galaxy Zoo but a little less rigid.
3. Using these tags/classifications the system would create a poll and present it to users with a question. What would you do in this particular situation? Where that situation is point in time of a hand history.
4. Eventually the computer player would then have a huge number of situations to use as examples with input from humans on how to proceed. This obviously will be very fuzzy and that's where, IMHO, the strength of the bot actually lies. The system would not lead to a rigid "Do X when Y occurs". It would decide from a large range of choices that have been entered by humans in the polls described above, leaning towards the most common answers first, but trying outliers also.
5. Use a neural network to create pathways based on previous successes.
How would this work once it's all together? An example:
The computer player is dealt AJ off-suit while being last to act on the dealer button. It would ask the database for a set of situations where AJ was dealt to players in the same position. It would then choose one of them and look at the results of the polls. How do most people with the highest success rate play this hand in this particular spot? Choose a random path to take based on that data. Observe and record outcome. This data becomes the true empirical data that the system will eventually rely on. If the situation has been encountered before in it's own play it will look at that and use it or it might choose randomly like it did above from the polling data. Eventually the weight of the empirical data it has collected might outweigh the data from the polls and it will "know" the right move based on its previous pathway choices. If this move is recognized by the opponent and exploited the results would dictate that the system falls back on another random choice. As the data set grows and the system plays the game, it could hypothetically be tuned to play consistently well.
The main hurdle is actually user-input. There would need to be incentive for users to enter the data they think is correct. The system is also open to manipulation through input data so that would need to be thwarted. And then on top of that the system would have to play an enormous amount of hands to create known successful pathways. But I think the sheer randomness and human-like qualities of the system would create a truly awesome experience.
(Joking) Now, does anyone have the $1M+ USD I need to fund this? I promise I'll pay you back. ;)
But I think it's worth looking at to only use hand histories and if possible only validate when something unexpected results. Maybe then I could use the human input to validate the computer's decision. I'm still hashing that out.
Personally, I'd rather see resources go into a robot that can cook for me but development follows the money I suppose.
Although Super System was written for all different types of poker, he is referring to no-limit holdem which still involves playing the player, with heavy math elements as well, but still much more psychology involved. Hence the name, you can bet any amount at any time changing the dynamic entirely. There is still no computer that is efficient at winning no-limit hold-em.
... and there won't be for a long time yet, as all hobbyist and academic efforts have been focused towards a consistently winning limit bot, because of the reason you outlined.
I'd settle for a robot that makes me a salad. That's one of the projects I might work on, once my current project, a robot that grows me a salad (automicrofarm.com), is successful.
A simple example, common stats software for poker will record what % of hands an opponent raises on the button if it's folded to him. If this figure is 80% you know his hand range is extremely wide. "Playing the man" simply is recognising that in this situation your opponents hand range is very wide so you can often raise the bet and take the pot without further resistance.
'Playing the man' is essentially building a mathematical model of one's opponent internally inside one's head.
Notice they tackle limit texas hold'em to keep the mathematics within reach. That's the time of crunching we're taking about. No-limit is impossible at this time. Not only will they need incredible computing power, but they also need the math to run it all, which remains unsolved. Could they use Sklansky's theory of Poker along with other optimal playing mathematics to win? Sure. But creating an unbeatable computer player? That is the incredible challenge.
Isn't this obvious?
Not at all. This article certainly doesn't imply it and it would be illegal.
Do you have any evidence of that being done?
You are telling me that nobody would program a secret algorithm, standing to win loads of money, to tilt the chances at the right moment? Because it is illegal!!!? Ohhh! Illegal! Surely nobody does illegal things? Not in a casino, not in a company, not in a government ...
No, I'm not.
>You are telling me that nobody would program a secret algorithm, standing to win loads of money, to tilt the chances at the right moment?
No I didn't make that generalized statement. What I did say was that in this particular case that hypothesis is of low probability:
1) The people building the machines don't make money on poker playing, but by selling them to casinos
2) The article explicitly mentions that they actually have a machine that plays too well and that they've had to dumb it down to give players enough of an incentive to play. So there's no need to cheat.
3) Casinos are heavily regulated and their hardware is known to be verified more thoroughly than both ATMs and voting machines. You don't risk a billion dollar business to nickel and dime a few customers.
4) Casinos make their money milking gambling addicts, making sure they don't take their money fast enough that they'll give up. Fixing the games would only reduce their variance not their final outcome and they have enough scale that the variance isn't high at all.
>Because it is illegal!!!? Ohhh! Illegal! Surely nobody does illegal things? Not in a casino, not in a company, not in a government ...
Adding a bunch of exclamations does not an argument make.
So, according to you, I, Mr. unbeatable hold'em player, can go to this machine, bet a million dollars and be sure that, in that perfect moment when I know I am going to crush it, it will not play tricks against me?
When I lose, how do I know? How can I be sure that it has not dealt itself favorably? It is not a matter of whether they are doing it: it is a matter of whether they can do it. If there is no independent dealer, this is not for me. It does not matter what the law says, which are the incentives, how they are generating profit, ... As long as the machine can theoretically deal itself a good hand, I am not playing it.
And, by the way, as long as the machine can know what cards I am holding, I am not playing it either.
Give me an independent dealer, and then we talk.
No, I never said you could be sure of that. What I do argue is that no one has much of an incentive to cheat you in this case.
>It is not a matter of whether they are doing it: it is a matter of whether they can do it. If there is no independent dealer, this is not for me.
Don't change the subject. The discussion was around if they were doing it. We all know it could be done, that's why we discussed this in the first place.
>And, by the way, as long as the machine can know what cards I am holding, I am not playing it either. Give me an independent dealer, and then we talk.
That's fine. You require 100% certainty of not cheating and this machine doesn't offer it. That is in no way an argument to say that they are in fact cheating. It's not even an argument to say that the probability that they are indeed cheating is very small (which is what I argued).
It's perfectly legal to make games that, bound by physics and probability, pay more to the casino than the players. There is no need to cheat, get it?
The only thing casinos need to work at is attracting a bigger slice of the pie to their games.
You, and nobody, will not notice small probabilistic variations. Whenever you discover it (let us say, 30 years from now), you will be told that there was a difficult to find bug in the random generator. Nobody will be prosecuted.
It could be a couple of lines of code in a subsystem somewhere, available only to a handful of engineers, and understood only by two of them - both of them with nice bank accounts in the Cayman Islands.
I suspect it does perfectly count cards that it has seen in its hand or community cards. This seems okay as well-trained human players can do the same. I suspect the machine would get even more play if this were obvious, maybe if it played at a table with a real dealer and could read community cards on the table. The technology to do this certainly exists.
The computer calls ... and deals itself a full house. I am bankrupt :( How do I know? How can I ever trust playing poker with a machine, when it is doing the dealing?
Or, let me put it this way: I will play any machine, no limit, if I can do the dealing (secretly, that is, as the machine is doing).
Another issue entirely is chess: no secret dealing going on. All are playing with the same in-game information.
A default position of mistrust when it comes to gambling is a healthy and safe attitude.
1) This is (fixed) limit Hold'em; you can't suddenly go all-in.
2) You ask how you could trust the machine. It would be trivial to keep track of all hands played and show that in the situations like the one you describe, the computer only makes a full house the expected ~9% of the times.
2. When this machine showed up ~2? years ago, there was a thread on 2+2 about it, and it would sometimes do some weird things like not value bet in obvious situations. They explain it in this article by saying it's, "playing dumb" but that seems like it would be a huge leak against Limit Hold'em HU specialists. I am guessing that they assume that they can make up for it in the weaker players losing consistently against this machine.
3. It is kind of annoying that the guy is proud that he "broke" a 24 year old player.
That said, if there's a mistake in there, a pro can eventually figure it out.
As you allude, the math of poker does break down in no-limit.
>Casino commissions, however, mandate that a gaming machine cannot change its playing style in response to particular opponents.
The algorithm is encoded in a neural net; there is no calculation of raw probabilities involved, at least in the sense programmers are used to.
I will never forget the feeling of "you've got to be kidding me" when the code I had written was able to successfully classify a huge percentage of the validation image set. I really should head back in that direction and get with the machine learning.
Re: poker software. Before the feds crushed online poker, I had already given up on low-limit limit hold-em. It had clearly become a bot and augmented-player race. I would assume given game theory and Bill Chen's book that some enterprising hackers would encode pretty decent no-limit bots as well. If any hackers know anything about successful no-limit bots, I'd enjoy reading about it.
But ultimately, the whole enterprise just seems distasteful: "Look, there was a 24-year-old who had beaten it for a while. Now he’s broke. And I think this machine had something to do with his demise."
There's more poker situations than amount of atoms in the universe
Yet the article states that the machine has multiple personalities (passive, aggressive, etc) and will sometimes throw hands.
In other words, machines are allowed to shark $player, but not allowed to target John Doe specifically.
Computers can easily calculate probabilities and expected values for every hand. So why they are not beating professionals?
Poker has no optimal strategy that wins against all other strategies. To play poker in higher level you must model the strategies of others, including them trying to model your strategy.
1. In the lowest level it's just maximizing expected value based on hand probabilities.
2. In the second level you try to learn the strategy of your opponents and maximize value against those strategies.
3. In third level you try to figure out how much your opponents have figured out your strategy and you change the strategy so that you maximize value when your opponents play against what they have learned from your strategy so far.
4. ... and so on. It' goes meta.
If you want to create ultimate pokerbot, it has to be able to model the minds of it's opponents. it hast to be able to detect leaks in it's own game and close them down. It must go meta all the time. It must learn how to understand how others think and how they think about it and change it's behaviour constantly.
Opponent modeling is purely advantageous, but not necessary
* Q: What’s a Nash Equilibrium or “game theory optimal” strategy? – Failed Math, Port Perry, Ontario A: An equilibrium strategy is one that wins the most money possible against a perfect opponent (this does not mean an opponent who can see your cards, but one who always knows your range whenever you take an action and makes the best choice against that range). In the game “rock, paper, scissors,” the equilibrium strategy is to randomly choose between the three options, choosing each one a third of the time in the long run. Finding equilibriums in poker is much more complicated, but the concept can be useful when you’re playing lots of hands against tough opponents. For example, if your opponent bets half the pot on the river after a particular series of actions, the pot is offering him 2-1 on his bluff. If he were a perfect player, the right thing to do would be to call his bet a third of the time, since if you called more he’d exploit you by never bluffing and if you called less he’d exploit you by always bluffing. In reality, of course, our opponents are never perfect, and so the idea of playing an equilibrium strategy at the table is usually pretty academic. *
You can certainly use Nash equilibrium when you have figured out the strategies your opponents are using. This is what Bryce Paradis is talking about. It can have practical value when playing Heads Up.
But If we are talking game theory and "solving poker", there is no single winning strategy that works against all other strategies and you can't calculate single Nash equilibrium that would be optimal in actual game against specific strategies.
Your opponent has 1352 different hand combinations. Assume he is playing the Nash equilibrium strategy. Make the perfect plays based on this. If he plays worse than the Nash equilibrium strategy, you beat him. If he plays perfectly, you tie.
Assuming your opponent plays perfectly works in chess. Chess programs are stronger than the best humans now.
You you can't do that assumption because you don't know what the strategy is. You can calculate Nash equilibrium only if you know the strategy your opponent is using. In full no limit hold em there is no single strategy winning strategy, so you don't know the strategy your opponents are using.
An optimal strategy’s goal is to loose the least against any arbitrary strategy. It is a strategy that is impossible to exploit in poker because poker has antes.
Poker players must seek maximal strategy. A maximal strategy’s goal is to win as much as possible against a specific strategy.
I'd like to add that how poker players generally define as their main goal, to maximize expected value: http://en.wikipedia.org/wiki/Expected_value
For example, you would raise more often when the antes are higher, regardless of the other player's strategy.
The optimal strategy beats everything but an equally good strategy, and ties against itself, but it doesnt necessarily maximise profits against other bad strategies.
If you are able to identify flaws in your opponents strategy then you can play non-game theory optimal to increase your profits against that perceived strategy. Doing so comes at the cost of you yourself no longer playing the best strategy though.
There exists strategies that gives higher yields vs certain unbalanced strategies than the game theory optimal strategy (or strategies - for all we know there are several optimal strategies in limit hold'em).
For instance - in limit poker if your opponent will never raise, call every street, but not call the river with anything less than a pair, regardless of what you do, then bluffing every river is a more winning strategy than the game theory optimal strategy. The game theory optimal strategy would include times when you do not bet the river, for balance, but knowledge of your opponent's flawed strategy would tell you that betting 100% of the time has a higher yield.
You can solve equilibrium for simplified poker games like just Heads Up with only shove or call an all-in options though.
It's theoretically possible to find Nash equilibrium over all possible strategies but that's not winning strategy. You just lose as little as possible. You lose against most/all strategies.
Take for example Kuhn poker (https://en.wikipedia.org/wiki/Kuhn_poker). It's very simple but first player has several optimal strategies.