Computers Conquer Heads-Up Limit Hold'em
spectrum.ieee.org
spectrum.ieee.org
The article states that this algorithm is weak to bad players but that's more an artifact of resources and training method; one advantage of minimizing regret on games instead of using linear programming is that online learning versions can adapt to exploit poor play with payoff larger than the game's value.
I've also posted here before that RM solves 2 player Zero sum game more efficiently than linear programming and how it's related to boosting, portfolio optimization and as an abstraction of natural selection.
Seems like a useful algorithm for dating sites.
(only half-joking)
The book http://www.amazon.com/One-Jump-Ahead-Jonathan-Schaeffer-eboo... One Jump Ahead details the checkers effort.
I highly recommend this book.
The "solved" aspect of resembles enumerating all possible outcomes.
Full table (9 or 10 person) limit hold 'em has been mostly dead in casinos for well over a decade and it only had a shelf life of a few years online, even during the boom.
There are better edges (for the sharks) and more excitement (for the fish) in no-limit and pot-limit games. I don't see a bot being close to be able to conquer heads-up no-limit, let alone a full table of no-limit. I suspect most heads-up limit happens in private games or the end of a limit poker tournament. It's played in scenarios that will not really be that exploitable should a person grok this bot's abilities.
The article also mentions that it's opponent was another bot that was also playing a very strong strategy. That suggests it's not really able to adapt to individual play, which is essential to being at profitable at all but the lowest levels of poker.
The other thing is that poker has an unexpected (at least to me) social aspect to it. The people who play against each other are largely friendly to each other, and they play against each other every day so playing 3-6 limit lets you spend more time at the tables than 1-2 or 3-5 No Limit, etc.
HU limit is probably only online, and theres a fair amount of bots prevalent in the games already today.
> The article also mentions that it's opponent was another bot that was also playing a very strong strategy. That suggests it's not really able to adapt to individual play, which is essential to being at profitable at all but the lowest levels of poker.
The computer bots attempt to play Game Theory Optimal - meaning it plays an (optimal) strategy that does not lose overall, regardless of stake/opponent. Against another GTO player, it will breakeven - otherwise it will be profitable. That is the holy grail of poker (and any game). If you are not playing GTO, there are leaks/mistakes in your game that can be exploited.
Check out https://www.youtube.com/watch?v=VHcrsMPQtgo , a poker theory course that talks about solving a simplified poker limit game
In the case of a poker tournament where there are at least a few non-optimal players, it's advantageous to exploit and take chips from them before other players do. A player who plays as though all of his opponents are playing optimally probably plays too conservatively in that early stage and no doubt faces a significant (non-optimal) disadvantage. Your comments still apply to heads-up (two person) poker, though.
There are two main categories of poker strategy: exploitive and balanced (also known as game theory optimal or GTO). With an exploitive strategy, your goal is to find the mistakes your opponents make (eg their tells) and figure out how to profit from them. With a balanced strategy, your goal is to 'balance' the choices you make (eg don't have any tells) so that it's impossible to your opponent to gain any information about your hand.
A perfectly balanced strategy will never lose money over a sufficient sample size, but it also limits the amount of money you can win. Exploitive strategies will maximize the amount you can win at the risk of being exploited by other players. From my understanding, your best bet against the weakest players is to play a purely exploitive strategy, and against good human players, you'll want to blend exploitive and balanced strategies.
My understanding of poker bots (the Alberta group from the article and PokerSnowie are the best-known) is that their goal is to develop a true balanced strategy. This may not be the way to maximize profits, but as long as humans aren't capable of playing perfectly balanced, the bots will slowly win in the long run, even against top pros.
Hard but not impossible. As long as 5-6 years ago there was a number of bots[1] beating mid stake (15/30 & 30/60) 6 max limit holdem on Ultimate Bet (before it imploded). The games were fairly soft but pretty aggressive and those bots were taking a lot of money out of the game. They were also very frustrating to play against.
[1] They all had incredibly similar playing styles so were likely operated by a single person/group.
PokerSnowie plays NLHE, but only uses 3 bet sizes (half, full, and double pot) to make the number of decisions manageable. I believe it plays heads up and 6max. Of course, I don't know if we have any evidence to suggest it's "conquered" those games, and we don't fully know how much the fixed bet sizes limit its potential.
It's a nice social outlet, but I would definitely prefer no limit.
The article and my comment are referring to a two player limit holdem game vs an opponent who does not make mistakes. So I presume this is equivalent to a high stakes match with two professionals.
Moreover, bluffing isn't as crucial in Limit b/c you have little leverage to push somebody off their hand. Many people will chase because the pot odds or their own playing style demand it.
Finally, NL or Limit are both poker. In fact, they're both Texas Hold'em. So "quite literally" i respectfully disagree with your assessment.
I'm not a professional poker player, but I did work with one to build a bot about 9 years ago on top of WinHoldem. It was more or less a break even bot on PokerStars if left to play itself, but its real value to that team was as a force multiplier. It made multi-tabling easier, raising interrupts whenever the bot woke up with a hand.
I'm not one for online debates, so feel free to have the last word. To put it another way, I check in the dark.
Heads up limit poker was /quite/ popular before the US cracked down and online gaming and many players made a ton of money off it. Towards the end, it was clear botters were getting in on the action, and bots were already playing an extremely difficult game.
You are correct that no-limit, especially heads-up no-limit remains a tougher nut to crack from an AI perspective, but this is pretty damn impressive
NB: I played online poker professionally for a few years.
So why don't you f king use that instead of making your own program?
Also, how do they know that play is optimal if poker hadn't been solved before? They're comparing it to perfect play, but if we can compute that, then the problem is done anyway.
Your suggestion makes no sense, sorry.
As to the claim of "conquering". While there is no reason to not believe them let's see how they fare vs other near optimal AIs. There is a lot of scope for numerical mistakes when you are dealing with solutions that big and it may well be that they missed something along the way. Some other teams claimed they solved HU holdem some time ago (and without supercomputers). They compete in Alberta yearly championship every year so it will be easy to see how it goes for Cepheus.
They mentioned that:
Burch warned that human poker players should take the Cepheus strategy with a grain of salt. After all, Cepheus honed its strategy by playing the equivalent of a near-perfect opponent that made practically no mistakes. Certain strategies that wouldn’t work against such a powerful opponent could still prove very profitable for human poker players when exploiting the mistakes of other human players.
Only playing hands you are +EV against a full range would be exploitable as it is way too tight in heads up.
(Related thought experiment - consider a game where each player has $101 in front of them before the hand starts, and then post $50 and $100 blinds. 100% of hands should be played from each position.)
The point I wanted to make was that you play hands where your odds of winning is <50% a large amount of the time.
Also, HU LHE is (much) more complicated than full ring. You have less knowledge about your opponents range, not more. You have to play more hands, a wider range of hands, and you frequently need to make marginal decisions.
It does that with the assumption that the opponent is also very strong, so it doesn't attempt to exploit weaknesses. If you try to exploit weaknesses, you have to vary from optimal play, thus opening up your own weaknesses. The software doesn't do that, it just plays a perfect game.
Perfect play also exists in games like chess and checkers. Checkers has been completely solved and computers can play it perfectly. This is not true for chess, though computers play quite well. Go is even harder and the best humans easily beat the best Go computers.
Otherwise, how could they show it's close if they can't compute perfect?
This is a complete -solution- for a limited poker version. Really strong, but stochastic, Hold'em algorithms have existed for a long time.
The algorithm sounds like it basically identifies the path with the least negative EV for any given action, which is essentially how humans play limit too. "If my opponent plays perfectly, what path will let me lose the least or break even"
Interestingly I get pegged for a bot rather frequently anyway (well, perhaps once a month). It is EXTREMELY annoying when you have to answer a captcha while playing 40 tables, you time out everywhere no matter how fast you are.
Wow. I can't believe that sentence made it into an article, let alone one on IEEE Spectrum.
First, normal people don't know how much RAM is in a phone. Second, the numbers are reasonably identical (only differing by 2.4%) so what's the point of specifying it out that way? Why not figure out what an average amount of RAM in a new computer is (4GB?) and say "That's the RAM of about 70k new Desktops".
For many years, everything was compared to "The Library of Congress". I got real tired of reading cliches like how some fiber link "can transfer the entire 2 million books of the library of congress in 10 seconds". Or how some new storage device can "store the entire library of congress". Comparisons like that never made any sense to me. Did they mean just the text characters? With or without compression? Surely images weren't included? Doesn't the library also have stuff like maps? Why weren't they included? Etc. :-)