AI Beats Four Top Poker Players
bbc.co.uk
bbc.co.uk
http://www.pokerlistings.com/libratus-poker-ai-smokes-humans...
The above article spells out some of the details of the competition. The winrate (14.72bb/100) that the AI achieved over the 120k hand sample is almost certainly not due to luck. It is a huge winrate that most pros have to employ strong game selection techniques to achieve (only play against bad players).
Here's a layman's explanation of how a poker AI can be trained: http://www.pokersnowie.com/about/technology-training.html
And some details about the weaknesses resulting from how they've abstracted the game: http://www.pokersnowie.com/about/weaknesses.html
Ivey was pretty good at managing his image, too.
Almost all pros who have had success both online and live agree that live games a ridiculously soft compared to online games. The reason more online pros don't play live is because you have to live in Las Vegas or Macau to play a the highest stakes, or you have to be invited. Also, the big live games are usually mixed games that don't afford a huge edge to online pros who specialise in a few games.
He is a top pro. Not in HU cash. He couldn't hang with these guys at HU cash, at least not with his current skill in the discipline. But hey, poker is not that narrow of a term. It includes all kinds of disciplines, live and online, horse and stud and hold 'em, the list goes on.
Those same "top pros" who Negreanu wouldn't play HU Cash wouldn't sit in the big game with him.
Arguably bankroll management is the most important skill of a top poker professional, and that's maybe what big-game players are best at.
Also, a small caveat to all of your comment is that live games are soft compared to online games of the same limit. There is no online equivalent of the big game, or at least there wasn't when black friday hit and knocked me out of the professional poker scene.
And while bankroll management is certainly an important and required skill to endure as a professional poker player at any level, it is by no means what differentiates high-stakes pros from pros at the lower levels. There are lots of small stakes and mid stakes players who practice sensible bankroll management, but who will never acquire the skill necessary to make it at the highest levels. If you want to get a sense of how much skill goes into playing poker at the highest level, you should watch some of Phil Galfond's strategy videos on YouTube (see for instance [1]). Poker strategy has come a long way since Super System[2] and The Theory of Poker[3]. Even players who were considered very good just five years ago, can no longer compete at the highest levels.
[1] https://www.youtube.com/watch?v=khdoSFCQ9iA [2] https://en.wikipedia.org/wiki/Super/System [3] http://www.twoplustwo.com/books/poker/theory-of-poker/
The Big Game and online HUNL are both technically "poker" but they are truly completely different games.
The post I responded to slighted Negreanu, saying he wasn't a "top pro." He certainly is. Not a HU cash pro, of course, but a top pro? Absolutely. A match for the "best pro players?" Certainly.
The "best pro players" in a given niche wouldn't have sat in the big game, and he wouldn't have sat in their specialty. Overall he's the better poker player.
Up to 2010, Negreanu had some up and down and as poker theory developed, his game showed a number of pretty big flaws--which he himself admitted. Sure, he was a competent player, so calling him just a 'celebrity' might be misleading, but there were many players in the top online scene who did a lot more of volume and analysis that Negreanu, who pushed more edge and who leaked much less.
Poker is solved using a very large game tree, just as with the other games. The structure of the tree is modified to support the notion of hidden state, but beyond that it is essentially the same as the other games. The structure for representing hidden nodes was developed in the 1950s by Von Neumann. Most of the algorithmic innovations related to how to update the game tree.
My guess is that the primary innovation for the Libratus strategy was that of scale.
It's also important to keep in mind that the best AI can still lose, and the worst AI can still win (and everything in between). Poker involves randomness, obviously whereas chess/go/etc does not.
That's wrong. Even when you're holding a good hand, your opponent could hold a better one and reading them is a key element of poker. The opponent's hand is an important variable to decide whether you hold the winning hand or not.
If you look at the experiment in detail, you'll find that it was set up in the AI's favor.
>When a hand was all-in before the river no more cards were dealt and each player received his equity in chips.
While all that is less important when you can avoid all-in situations, the main statement -that the other player's behavior is irrelevant- is still wrong.
Could you elaborate on this ?
As expected, the AI is good at making technically correct decisions and "draining money" from a table by playing hands with sufficient data almost perfectly.
However, in decisive all-in situations with little information available, it supposedly wouldn't do so well, regardless of all the learning, but that's what it often comes down to.
>Nash Equilibrium is a strategy which ensures that the player who is using it will, at the very least, not fare worse than a player using any other strategy.
How do you make this work for situations that can cost you the game in one hand, with little information available? Without observing the opponent's behavior you can't, and for the AI that means it can be forced into making bad calls by playing aggressively, unless the game mode allows for avoiding such decisions, which was the case in this test.
Even limit poker is too large to solve directly. In 2015, limit poker was essentially solved with a new technique in game theory that allowed them to find a simplified model.
http://spectrum.ieee.org/automaton/robotics/artificial-intel...
Heuristics make analysis practical, and well chosen ones make the difference. This isn't to minimize the accomplishment, but rather to say it has strong similarities to other AI games.
You mean Libratus' strategy used a very large game tree. That is not the only strategy. Take a look at research from the University of Alberta [0]. Also, I'm not certain Libratus' strategy can be simplified to "very large game tree" as I haven't seen the paper, yet.
While finding a Nash equilibrium means no other player can beat you, it doesn't mean you're going to make the most money in a big ring game. A different strategy might lose money to an equilibrium player, but exploit a different, weak player so much that it's worth the loss.
This happens because some entrants aren't 100% random, and the worse of them can be exploited by the better of them. What happens is that the results involving any random AI essentially degenerate into noise, while the tournament is really contested between the nonrandom entrants and will be won by the one of them with the best strategy.
Put another way: to win or place highly in a tournament, you don't just want expected win-rate, you want variance. If there is no difference in reward between a 50% win-rate versus a 10% win-rate (both are far out of the money), but there is a big difference between a 50% win-rate and a 90% win-rate (the latter wins the tournament), you will seek the 90% at the cost of potentially ending up at 10%.
The 33%-each Nash Equilibrium is the mixed strategy Nash Equilibrium of the micro game (i.e. a single round of RPS, averaged over all possible randomizations).
This is in no way the Nash strategy of the tournament game, which is "win the tournament given a pool of unknown participants and a set of rules for whom you face when". You have to add additional assumptions (e.g. that everyone else is going to play the uniform random strategy) in order for uniform random to be the Nash strategy for the tournament game.
If the pool includes players who deviate from the single-round nash equilibrium strategy, there is opportunity to exploit them (and in doing so, open yourself to possible exploitation). This is why pure random play can often perform very poorly at the tournament game.
Isn't that literally what a Nash Equilibrium is though? It's my understanding that if there is players playing exploitably in the game then it cannot (by definition) be a Nash Equilibrium, so the Nash strategy may no longer be the optimal or maximally exploitative one.
Not as much "quite poorly" but more specifically, they will provably land at exactly the median position in the ranking (let's assume there's only one pure random bot in the tournament, no reason to have more than one, but the argument also works with multiple).
While it's impossible to win more than 50% of the time from a pure random RPS bot, it's also impossible to lose more than 50% of the time.
So all the other AIs will on average score exactly 50% against the random RPS bot. Whether they end up in the final ranking above or below this median line depends on how well they do against each other.
"AI beats top 10 hedge fund managers"
to
"AI run hedge fund blows up due to black swan event"
regardless it's an incredible feat. It really casts questions into what our edge as humans are which is slowly disappearing and we didn't even need to put a brain in a jar and hook it up to a computer....it's deep learning reinforced algorithms that is appearing to outlearn, outthink the best of humans.
I just can't emphasize what a monumental period in history we are at. Humans are producing specialized algorithms that learn and hold information about the deep web of relationships between myriads of parameters to produce superior performance than humans.
It's almost like we've uncovered ways to automate our intelligence very much like we've been automating human and animal labor in the past couple centuries.
So the question is, how does an average joe hacker like me exploit and leverage this wonderful thing called deep learning? I'm not interested in reading PHD papers with advanced calculus.
I want to have a map of what AI, ML, DL, NN methodologies to use and when and who to hire based on that. This is no time to be a luddite and don't count on basic income from appeasing the masses anytime soon. Much like people took the most hit in the early rise of industrial revolution, our generation and immediate generation will be hit the hardest.
Knight Capital, 2012:
https://www.bloomberg.com/news/articles/2012-08-02/knight-sh...
I'd think one would make a lot of money shorting such outfits.
The other capital that blew up from the comments suggest it was more of a human error that led to deployment failures which is exactly why AI hedge funds is such a paradox. At the end of the day, it's humans that's deploying and building the model to be approximately right with high degree of accuracy but in the day trading environment it is a poor model for success-not only do you have to be right your monetary exposure must be right...."but proper money management and you will be fine" said everyone who took part in this zero sum game.
I don't know, prove me wrong, I'm sure those ivy league engineers on wall street aren't getting paid dimes for what they do.
Knight's blow-up was a process/sysops failure, not due to their strategies going wonky.
[As a side note I believe their core strategies were human-analyst designed, not ML based, but I could be wrong.]
https://en.wikipedia.org/wiki/Knight_Capital_Group#2012_stoc...
Why's this an argument against AI run hedge funds? Funds run by humans also blow up due to black swan events.
Well reading paper is a must to get to deep learning. Those papers may not be that math heavy once you are used to it. Most of the time, it is about network architecture and loss objectives.
Hedge funds are already using A.I. successfully. I talked one manager who's team was trying to layer successful individual neural networks into a "bigger brain".
These aren't silver bullets. Both cost more, in fees, and have a habit of leaking order information across informal channels. They're also, generally, slower--there are, on average, fewer buyers and sellers for a given security in a dark pool than in the open market.
What I do know is that it wasn't worth it for me personally to occupy my GPUs. By contrast, larger trading firms would probably like it.
https://www.udacity.com/course/machine-learning-for-trading-...
I averaged a return of 7-13% a year. With less than 1000 in capital, it wasn't very profitable.
I don't see why it shouldn't scale.
FFS, if you're not willing to read a paper with BASIC CALCULUS (it's HS/college, not advanced like fractional), then I'm not sure machine learning is the right place for you.
Math is good but my time won't be best used if I have to learn calculus all over again just to begin understanding the linguo. Rather have a generic model of what to use and when, hire those that have that capacity to dive deep when needed and implement the expected business outcome.
To me, your comment reads similarly to "Much like most business people don't spend their time reading code to do the job, I question why such academic rigor is to be demanded by someone who is translating business requirements into database operations".
You don't need to be an oil rig worker to understand where the opportunity lies is what I'm saying. It's possible to understand systems from a blackbox point of view and still be able to exploit them. Afterall, this is what we do whenever we hire a new developer, it's impossible to know what they know so I leave that up to smarter and talented people while I allocate the capitals and set a vision for what's required to capitalize on the business opportunity that AI presents with it's superior performance, scalability and economies to scale.
You outsource what you don't know or can't do and if you can't afford it you do end up doing it yourself. This is why capitalism works. It's not necessary for you to dwindle with all the underlying complexities when you model the business opportunity in terms of basic expenditure and investment returns.
Dude, I'm not even asking you for anything that difficult here, how would you like it if I walked in and said "well, I don't need to learn any programing, my time is better spent magically making it work".
You don't need to know programming if you understand your workers and they understand it enough to convey to you the time & resource constraints.
I have a million dollars for this project. What are the features and functions that will maximize my returns? Hire a couple of smart engineers who already holds the technical knowledge, able to communicate when asked all the pitfalls, potential roadblocks and unknowns so that their productivity is maximized. Set expectations on how they will be able to trade their time for money. At the end of the day, it's a naked call option with limited risk and unlimited reward by utilizing the correct knowledge workers.
Why spend time learning the reference when you can look it up or hire somebody who already has a distilled version in their head which you can utilize?
It's, like, inadvisable, to leave yourself vulnerable like that.
In an orchestra, the conductor is not replaceable but the instrument players are. Similarly, someone who allocates his or other's capital cannot be easily replaced very much like the worker's.
Your engineer leaves there are 100 others ready to take the place. Business owner leaves you'd have to find a buyer to keep everyone going.
This is different from a CEO funded by VC's where you are still replaceable but nevertheless a lot harder to replace than the cost intensive worker.
Being on the profit generating part of the business and cost intensive part of the business helped me understand why such discrepancy exists-those who are able to read and perform in chaotic and uncertain environments will always be valued than those who directly affect the margin's as a result of the monetary figure attached to their time which is always far more costlier because cost centres do not generate new revenues. Your top sales guy makes $9 for every $1 spent on him (+$8) where as your top engineer costs $3 for every $0 he generates (-$3).
tl;dr: revenue generators are king while cost generators are easily swappable due to the high supply of it as a result of it being a far safer and financially stable perception.
(it cuts the other way too: if an oversupply of labour in a field as difficult-to-learn as ML does arise, the demand shortage causing wages to fall is almost certainly because ML techniques aren't giving companies and traders as much of a financial edge as they hoped for...)
Added bonus for the ML experts: in a lot of the possible areas they can choose to work in, their individual marginal contribution to the company's profitability is at least as quantifiable as that of a salesperson. In many areas of finance you can quantify the impact of an individual line of code!
On the other hand, it's easier to piggy back off the hard work of highly paid experts in this space now commoditized into 1/10th of a cent per application of an ML algorithm from Microsoft or Google or IBM.
Commoditization is going to be in full force as businesses realize they can just have a junior developer to plugin IBM Watson or Azure than hire an expensive PhD student who may know everything under the sun but cannot compete against an army of highly paid experts completely under the control of a highly collusive labor market.
Silly google, they must've skipped Econ 101 since they hire tons of engineers and are doing TERRIBLY!
Why would I go to you and not a VC firm? I sure as hell don't want to work for someone like you....
(oh and there are plenty of others like Karpathy, like Facebook, Baidu, Microsoft, Uber, Amazon, a metric fuckton of startups, etc. who are being paid ridiculous amounts of money to build nets)
A typical Econ 101 textbook says:
higher demand --> higher price
higher price --> higher supply
higher supply --> lower price
lower price --> higher demand
In Econ 101 we pretend that cycle eventually reaches an equilibrium. In grad school we analyze the dynamics.But even if we believed your demand causes supply causes zero price model, why is that unique to engineering and not business acumen? And why do programmers get paid anything, don't they work for free?
Programmers have largely driven themselves to zero, just take a look at the workers on freelance websites and how much difficulty a native English speaking freelancer has against an army of commoditized labor.
There are still six digit earning engineers and always will be but you can't look at that as a metric for the massive commoditization that has occurred with generic programming in the past 15 years.
I'm a freelancer and fully booked, so maybe those low cost competitors aren't so competitive after all.
Guess who's getting replaced once the VC figures out who's doing the work...
In this case you should immediately fire all your engineers, your CFO, and your COO. Because why would you want to employ a team of engineers who are costing you money but bringing you no value?
You should also fire your cleaning staff, who have never closed a deal in their lives, plus the plumber who most days dosen't even come in.
Granted, engineers are required to create products and maintain it. You need product to sell.
But at the marginal level for every dollar your sales person makes your engineer is not. You might argue but the product is generating revenues but that's not what drives a sale. A sale is a function of value derived from the product and the price paid for it. Engineers aren't driving the sale unless the product itself is a developer tool. But even in that case the top level business controllers will always have the final say. In the end, buying such developer centric tool is about minimizing cost expenditure.
There is critical proprietary knowledge and experience that is formed from a sales person that is driving revenues. Businesses do not want to lose the hen that lays the golden egg as they are hard and expensive to replace.
Engineers on the other hand are easily replaceable and are expensive because they cannot generate revenues. They are not efficient when pulled away to a sales meeting from the work they are doing. Their knowledge and skill costs businesses time and money with no way to get it back directly. It must be sold, money collected and redistributed to all the payrolls. That's the main risk a business owner or a CEO is dealing with where as the engineer collects a cheque at the end of the month with little to no concern or exposure to that risk.
Society rewards risk takers disproportionately at the corporate level. Even if it's tough to measure to the Board and investors losing a C-level executive is always going to be more impactful than a senior engineer who has far more workers to replace him as a result of being cost intensive.
Find me an engineer that can code and sell, now that is a truly rare hybrid, a mewtwo, but it will still be more expensive than a guy who just sells (and does well).
A poor salesman understands this better than anyone, his weight is worth the revenues he generates. For engineers it's the efficiency / dollar or output / dollar that they are competing against which is always headed towards commoditization and any business would replace them with an AI that can code if they could if it cost less.
Then exactly what value are you bringing to the table? Money? Sales skills? If that's the case then why focus on ML/AI technologies? Any company needs money or sales expertise.
I say this is an arbitrage opportunity because AI is essentially commoditizing human intelligence. It didn't take much brains for people to realize cars could do away with horses. You don't need to understand how a car works down to the formula. It's enough to view it as a black box and still know all the outputs and constraints associated with it in order to orchestrate capital that allows you to seize the business opportunity.
Again, the arbitrage is buying something for cheap and immediately selling it at a higher price...if you can't understand why a deep level understanding of the academia is not necessary in order to capture the opportunity then less competition for me which is great.
That should be enough for you to have an idea what expertise you'd want to hire for.
Granted, it sounds you're just looking for ML scientists or ML engineers. Data Scientist might also be able to help you out, so that's what you want to hire.
But let's indulge your point of view for the sake of the argument. So you've got a black box AI that does something. You know what inputs it needs and what outputs it produces. How do you arbitrage it? What creates the price discrepancy that allows you to arbitrage at all, given that your incapacity to understand the black box renders you unable to gauge its value beyond what the market is telling you?
In that case, why not trade any random commodity?
> In that case, why not trade any random commodity?
Its much easier to gain control of markets that are new as opposed to established ones.
"Listen I have this great idea for a web app, it's going to be like the Facebook for hardware stores [or whatever] and people can rate/vote/tag ... I got the whole idea worked out, I just need someone to program it for me" etc. Every programmer has heard these kinds of proposals many times.
One (of many) problems with this kind of proposal is that the idea-man, unable to program (and not wanting to learn) has no idea about the complexity of what they're asking. And even if they think they have the whole idea worked out, not being able to program means they probably missed a lot of shortcuts, possibilities and best practices. The worst ones hand-wave this with the possibility of a fresh outsider-look on things! Except that the whole (fully worked-out) idea probably needs to be reworked entirely before it's even feasible or competitive. The programmer would probably have been better off without the idea, cause that's where it started to go wrong.
Similarly, if you believe math isn't worth your time but instead want to hire ML experts to implement your ML-related ideas (formulated in terms of "expected business outcome"), puts you in a very similar role.
(Tongue-in-cheek) I got a really great idea, the expected business outcome is: make lots of profit. It's fool-proof. I checked with some of my people and they agreed, making profit is a solid idea. I just need some business dude to implement for me. I'm not really sure what sort of business person, my time can be spent better than learning about subtleties between sorts of business people. I just need to hire a good one. But remember, the idea was mine first.
But for instance understanding and specially producing meaningful language, let it be natural language or programming language? I hope not so.
Because otherwise, your average joe hacker might as well shut down their IDE and say good bye.
Actually, a lot of mental work isn't based on intelligence but rather training in pattern recognition, thinking faster, translation of information between different coding systems, and developing instincts about the behavior if complex systems. This isn't very different from manual labor when you think about it.
Now it appears the essence of that is being distilled into parameters for AI algorithms which can produce superior output by teaching it with decades of professional human knowledge which maybe dumbed down with simple machine learning.
Renaissance Technologies already is one of the top hedge funds in the world and it's fully automated.
I wonder if humankind will be the creator of a new form of life. One that is not bound by chemical processes. One that moves at the speed of light and spans the whole planet. This being, or these beings, with access to billions of sensors, will know everything, see everything, hear everything. They won’t be limited to visible light, nor to the frequencies we can hear. There are no limitations except for what sensors can record. They will be able to interpret, predict, plan. They are not bound by time as they themselves don’t physically age. They will use vessels to carry out tasks and build the infrastructure they need to live. In comparison, we will look like single-celled organisms. So dumb. Until we disappear.
You've got a lot of comments like "if you're not willing to read a paper with BASIC CALCULUS ... then I'm not sure machine learning is the right place for you" to this, but I second you on this. I do have a PhD in CS and not afraid of calculus in papers. I just don't think reading PhD papers is the best path to getting practical results in DL/ML. "Not right place for you" is the sign of the "exclusiveness" problem.
I'd recommend checking the following course by Jeremy Howard and Rachel Thomas:
It is designed to make DL accessible and achieves its goal very well. I'm not sure if the average Joe Hacker will be able to write top-pocker-players-beating program after this course, but it definitely gives enough insight and practice to get started in DL and be able to design and code solutions in a large number of problem classess.
But, an artificial intelligence probably has different black swans than humans, as their perception is inherently different. A tweet from Trump might not surprise us (anymore) but to an AI player, it might not be evident that a small amount of text from one Human can cause an uproar on the market. And it might not even have access to the relevant data (Twitter) at all.
It's hard to know how much that affected the strategy, but in the Reddit thread, the human players said the overbet frequency was what they were most surprised by.
You'd happily go all-in pre-flop AA vs KK. On the other hand if you got 4-bet pre-flop by a 22 and you're holding AK, you might ask, "Check it down?" This is assuming the opponent really likes small pocket pairs and will call an all-in, etc.
In a heads up game? Really? "Check it down?" ???
Also, if you're first to act on the next round, it might be worth asking, even if you don't think they'll agree.
Kelly's criterion and the sharpe ratio is not relevant for the format played in these human-AI heads up games. They play with an unlimited bankroll.
My point exactly about the unlimited bankroll. The experimenters may not have realized that an unlimited bankroll would significantly affect the strategy.
https://www.reddit.com/r/IAmA/comments/5qi3i9/we_are_profess...
This sounds more like an "advanced chess" setup, where a human teams up with an AI to play. The title of the article should really be "amateur poker players + AI defeat professional poker players". The real test would be if the AI self-corrected over the length of the tournament, without human intervention.
or conversely, operators shouldn't be allowed to tweak the AI logic, but if they program the AI to tweak itself, like review its moves and adjust its algorithm automatically, that would be reasonable, then the AI and nothing but the AI is responsible for the victory.
This doesn't sound like they had the goal of conducting a fully controlled experiment here, but it's still interesting none the less.
They way things are going with AI. You have a good algorithm, you can get rich very quickly by being a one man business with hardware rented in AWS.
Libratus AI player is modelling it's human counterparts and predicting how they think to outsmart them.
When Google started, they got the page rank algorithm and distributed algorithms good enough to run on shitty unreliable cheap computers. They are well on the way to become the world's largest company overtaking Apple someday.
Their ad algorithms already know that I am applying for a house loan and are blasting me with ads every fucking page I visit on the Internet.
I can totally see Google and Facebook personalizing ads per person and taking advantage of the person's vulnerablaties. Like psychologically modelling them to make them click ads and buy random shit. I can see the start of ultimate God algorithms for marketing.
The ability for AI to create drug like experiences for us that we can't stop craving.
The probability space of poker is such that 4 competitors isn't going to tell you much.
And defining "top" is difficult because "top" may be more celebrity than anything. Everyone has their different objectives. If you are "top" then you sure as didn't get there by building a case history against AI poker bots. Give these guys a chance to adjust, and give them a chance to figure out why the effort to adjust might be worth bothering with.
> A poker-playing AI has beaten four human players in a marathon match lasting 20 days.
20 days seems like plenty of time. No ?
People talk a lot about number of states in poker, but the hand and visible cards can be easily (to a statitician) reduced to a scalar "probability of having the best hand".
At least in terms of crunching the possibility tree this should be a far less computationally intensive challenge.
At this stage, AI wins in competitive games simply will not impress me. From here it may simply be a tour of force showing the breadth of fields AI can dominate in.
After that enthusiastic agreement; poker wins are evidence for the laypeople of something I suspect most people in AI research or game theory already know - a human bluffing isn't a magic advantage in a fair game. The AI can outperform on fundamentals.
It's just a different kind of game
I still think the hardest part of poker is the grind. The computer doesn't get bored or tired and play hands it doesn't have a reason too just because they are stuck in a dead streak. Taking a mostly conservative style, you'd expect the computer to out perform over time.
It would be curious to know what percentage of the time the players bluffed the AI successfully and vice versa.
An additional player requires some sort of modeling of the interaction between players, which may not be feasible. For example, Player A may have multiple co-optimal strategies with respect to his own EV, but that affects the EVs for players B and C differently. Player A can choose arbitrarily between these strategies at any frequency, but players B and C can't predict this choice at all.
On top of that, Nash equilibrium requires that each player acts independently in their best interest. Teams can be modeled as a single player if necessary, but shifting alliances over the course of a series of hands can't be easily.
And on the pure computational complexity side, the state space explodes when you can have more than one opponent in a hand simultaneously. Combinations of players in a hand scales as n!, not n.
How likely is it your opponent is bluffing, based on their past behavior?
Should you try and bluff this hand?
"Oh you have AJ don't you?"
no-limit and tournament play were deliberately placed outside the scope of their poker-bot projects--at least during the period of time i was following it which was approx. 2004 - 2010.
anyone know if the CMU team trained their rig on these variants?
This is a great achievement in AI, don't get me wrong, but the headline should read, "AI beats the best four poker players we could find who were willing to play for a mere $200K".
All the actual best players play for millions and have a reputation to uphold. They would never agree to do this.
They four guys they got are pretty good, and could certainly destroy me, but they aren't the best of the best.
I'd love to see the bot play in the World Series of Poker for a few million.
All these players are high stakes players. I think you're underestimating the fun factor. As for reputation. For a poker player having a reputation as being beatable is a profitable thing to have.
>but they aren't the best of the best.
Who do you think is? Like how many people do you think rank above this group at HUNL?
Same as in programming. There are some great programmers that we all know and that are public figures. But the absolute best? Probably making high 7 figures working in a dark room somewhere.
They aren't the best since I think that probably is safely in Polks hands but they are near the best.
Or any of these people: https://en.wikipedia.org/wiki/List_of_World_Series_of_Poker_...
If someone asked you who the best mobile app developers were in the world, your list should comprise solely of mobile app developers.
Jamie Gold won the WSOP, but he's a pretty bad player.
I imagine it's just reinforcement learning where the inputs are the actions of the individual players (hold/fold/raise, timing etc) and the statistical probabilities in terms of expected cards. Train a neural net to predict probability of the opponent's hands and act accordingly.
Is it just that the professionals all act similarly enough that the bot can learn based on other players?
http://www.pokersnowie.com/about/technology-training.html
And they go into some of the problems that stem from how they've abstracted the game here:
Even if poker sites could somehow perfectly detect automated players(which they can't of course), highly skilled poker is profitable enough that some people would be willing to manually execute the actions themselves as directed by the AI.
But certainly in all HE variants (that I'm aware of), counting and remembering is unnecessary.
[0] It's necessary, but not sufficient.
Another part has to do with the fact that both you and your opponent have surprisingly many legal "moves" at each turn. because not only must you decide to fold, call or bet, but if you bet, you also have to decide how much.
the first authors twitter account https://twitter.com/polynoamial/
I am sure the AI they built is a mighty and spectacular achievement, but poker is about the only game I know where the worst player can easily beat the best player. Prove me wrong and I'll happily accept your insult of "ignorance"
edit I should have been more clear, as I am extremely impressed with the AI's results of hands over time. I am referring to the context of a standard tournament where the loser is eliminated after losing their chips.
They will not be able to consistently beat a better player, especially not professional players.
That's what this AI is doing. It didn't beat them once or twice. It beat them consistently over the course of 120,000 hands in a 20 day event.
Is it theoretically possible that "the worst" player could do that based entirely on luck? Sure, probably in the same realm as monkeys, typewriters, and Shakespeare.
So the deck for hand 1000 might be the same as hand 113853 but they swap who is the button.
That is, if the game is 100% luck, the better player will have an expected win percentage of (100-100/2) == 50%
If the game is 0% luck, the better player will have an expected win percentage of (100-0/2) == 100%
If the game is 95% luck, the better player will have an expected win percentage of (100-95/2) == 52.5%
Conclusion: If there is any non-zero amount of skill in the game, the better player will win in the long run.
Why would you refer to that? That's not at all what happened here.
In the game that was played (long-term HUNL cash) the worst player cannot beat the best player at all. Not easily, not at all.
Not to mention they didn't run the cards out.
You're commenting without reading the article, seemingly.
Find the right balance, and your opponent can't exploit you. With no-limit hold'em this is extremely complicated, and until recently the best humans have always beaten the best bots.
There's no such thing.
An algorithm playing straight hand value based on probabilities is more susceptible to bluffing, not less. And this is no limit, where a single hand can swing all the chips.
Any poker AI that isn't a loser is going to have some pretty sophisticated modeling of the opponent.
It makes much more sense for the strategy space to be the set of probability distributions over game moves (i.e. mixed strategies).
I think that the optimal mixed strategy for each hand is immune to bluffing (over many hands it will have larger expected winnings against a bluffer). If that wasn't the case, there would exist no Bayes-Nash equilibrium for the game, contradicting Nash's theorem.
It'd be far too easy to recognize when they have a good hand and fold and to push them off all their marginal hands.
> I think that the optimal mixed strategy for each hand is immune to bluffing (over many hands it will have larger expected winnings against a bluffer). If that wasn't the case, there would exist no Bayes-Nash equilibrium for the game, contradicting Nash's theorem.
I believe that's true. I know for sure that heads up limit hold'em has been solved. That said, I think this context is similar to the iterated prisoners dilemma contest. There's certain to be an equilibrium, but what's interesting isn't the perfect strategy in a min/max sense, but rather a slightly suboptimal strategy that can detect and exploit suboptimal behavior in other players. It sounds like you know this area well, perhaps you can shed some light if I'm on the right hunch?
That's what you see online poker players do. They model their opponents (in the sense of labeling them as fun player, too tight, too loose, etc), then try to predict their hands based on their moves. Otherwise I guess they would be losing money: poker is zero sum by its nature, and the casino's cut on top of that makes it negative sum!