CMU's Libratus builds substantial lead in Brains vs. AI competition
cmu.edu
cmu.edu
An ecology of Artificial Intelligences, unbounded by our evolutionary history and neural architecture, could evolve to suit each particular task more effectively than our brains can.
Promises and perils abound.
[1] http://www.cs.cmu.edu/~sandholm/ section "Algorithms and complexity of solving games"
what's the peril? I can only see promise ahead, if AI's progress continue unimpeded.
For once. But back when we were thinking about the perils of the internet we didn't correctly anticipate homogenization of media consumption and filter bubbles. People rather thought that it would be the other way around, everyone would have access to high quality information and enjoy diverse (long tail) media content. So we're likely to be wrong about the precise perils here, too.
How about a machine that can beat someone at Smash Bros, a game with varied characters, complex comboing mechanics, and a nontrivial computer vision task?
Or--more difficult by a few orders of magnitude--what about a robot that can beat someone at tennis? Or a team of robots that can best a professional basketball team?
When do you suppose we'll begin to see these sorts of things? Within our lifetime, I hope?
http://kotaku.com/overpowered-bot-puts-top-smash-bros-player...
(I don't even think the base AI can actually see your location in the air very well, and it has no air DI itself so it really could not win for long.)
on a related note, hopefully a benevolent gai emerges before human society implodes.
I think robotic and A.I. are not the same and shouldn't be compare. IMO, the only missing part in beating someone at tennis is having a robot that can compete with any athlete.
like using a few degrees of freedom and little strength to chop onions and some vegetable, crack eggs, whisking them, pour a bit of oil into a pan and lighting a stove, pour the omelette, flip it onto the plate, throw the eggshells in the trash unless it's too full (and take the trash out if it is) and wash the chopping board and pan with a sponge (adding a little dish soap) and not too much water, then rinse them thoroughly but without too much water. Which poses no challenge to most adult humans who take just a bit of time to learn (if you never learned to make an omelette and you're an average adult, by Wednesday you can make a perfect omelette every time.) Even though objectively humans have very weak hands, see things very, very slowly compared with machines, and cannot do any single mechanical action as reliably and predictably as robots.
computers might be able to find and count all the primes between one and a million before I can count the ones up to ten[1], but they can't even scrub my bathtub with a sponge given a whole afternoon to do it - not without a lot of specialized robotics anyway.
[1] https://www.quora.com/How-long-does-the-fastest-algorithm-ta...
The human system is pretty incredible.
Chess can be modeled with each game serving as a single observation. Poker must consider each player's entire lifetime as a single observation. Not simply a hand, but all hands that player has ever played, including what that player knew about all opponents ever faced. This quantitative increase in data creates a qualitative difference.
Poker is less like chess and more like repeated rock-paper-scissors.
see: https://en.wikipedia.org/wiki/Nash_equilibrium#Nash.27s_Exis...
Check out the "Occurrence" section in the article you linked to.
Nash equilibrium may not exist if one of the players follows, say, a Markov switching process. If that process causes the opponent to stop seeking equilibrium or to settle into a false equilibrium, then the switching process may have been a better strategy than seeking equilibrium.
To your point, I wonder how they account for mental and physical fatigue. To a computer it makes no difference to play thousands of hands over such a long period of time or hundreds of hands over the course of a single day. Humans on the other hand don't have the same attention span as a computer.
which is: "If after 120,000 hands either Libratus or the humans are one standard deviation above break-even, they will have won the competition with “statistical significance.”"
(I'm a professor at CMU, but I have nothing to do with this research or competition.)
This challenge is very unfair to players so I wouldn't say it won, players have a massive disadvantage, every professional player has tracking software and a database to analyse every decisions that has been made.
This is of course extremely important because you can model your strategy to exploit the suboptimal decisions made by your opponent, yet players have no access to any of these tools, so the bot adjusts it's play based on their human opponents but humans cannot do the same and are left with a guessing game.
If they want to make a proper challenge then players need to have access to the tools they usually use playing in the Internet.
Do we know if it actually does it? I imagine it's much simpler to build a bot that plays a balanced profitable strategy rather than one that tries to build a model of their opponent and exploit it.
I assume there are probably papers that specify what form of learning is taking place, but the article didn't go into that level of detail and I haven't tried to track it down.
That being said, the AI seems pretty impressive. Not sure how they picked the players I could think of a few HUNL players I'd rather see but they might not be interested in a 200k freeroll.
This happens in algorithmic trading, where traders would make a large number of low-valued, bad bets to mislead the algo. Then bet big and go the other direction.
As to your point about the EV, this is why collusion can work. By colluding over a long enough horizon, the AI can believe that the average expected value to be something that it is not. If only one individual feign a weakness and the rest do not, then the strategy doesn't work.
For example attempting to feign weakness by betting small in a spot where your entire range should bet large is not tricking the AI, it's just passing up on EV for the players, good players are not going to play poorly in hope of tricking the bot for future mythical EV gain.
Also, there is no 'colluding' in heads up poker.
An example, say the humans are getting to a river situation with too many bluffs for a given betsize, an exploit for the AI would be to always call. The opposite is also true, if they are bluffing too little it should always fold. The players notice that the AI has adjusted, and adjust their frequencies - now exploiting the AI. By taking an exploitative approach the AI leaves itself open to be exploited, this is not the goal.
If this were rock paper scissors, the AI is doing the equivalent of always throwing each at 1/3 - even when it's opponent throws rock every time. It could switch to paper, but a thinking opponent will now switch to scissors, this will continue until we are back at equilibrium. The AI aims to play poker in this same fashion, having the correct frequencies of actions for a given range in every spot.
Poker isn't about equilibrium, it's about misdirection and exploitation. When the table gets cold, you liven it up by convincing everyone to do a round of straddle.
"Tricking an opponent into thinking it has metaphorically thrown rock" extrapolated into a poker example would be betting larger/smaller, calling more/less, folding more/less than is optimal in a given scenario in the hope that your opponent makes a (bigger) mistake. You're simply hoping he makes more errors than you, the AI instead choses to just make zero mistakes and let the opponents do the rest. You can see this in action for yourself in Heads up limit holdem by playing Cepheus (http://poker-play.srv.ualberta.ca)
I agree that would not happen if two equilibrium-seeking computers played each other. Since the human strategy is unknown, it is possible that equilibrium may not exist or be optimal. Even if it's two computers, if one of the computers has the possibility of choosing a non-equilibrium strategy, then again the optimal strategy may not be to seek equilibrium.
It does totally eliminate variance, but they also take that into account and correct for it when looking at final outcomes usually. Right now the bot is up by something like 800K over 60K (out of 120K) hands. If that rate continues, it will win by around 1.6M or 400K per human. The blinds are 50/100, so that would equate to roughly 33 millibets (thousandths of a big blind per hand). That isn't too far off from standard win rates in bot vs. bot tournaments [1].
I'd say it's likely that the results of this tournament will be a statistically significant win for the bot.
[1] http://www.computerpokercompetition.org/downloads/competitio...
Typical winrates in human vs human are between 1-5ptbb/100 where 1ptbb = two big blinds. At 1ptbb the variance is pretty big and north of 1million hands are probably necessary to establish an edge, whereas at 5ptbb the variance is much smaller and 100k hands are usually enough to converge to the expected value
Of course it also makes games shorter and introduces a lot more variance, which isn't so good for assessing how well a computer plays.
It is akin to only playing the first 10 moves of a chess game, then resetting.
The chess analogy would be more akin to resetting after the flop.
It's a great solution, and at that stack size, I'm sure it's better than nearly every human competitor. But until they solve all stack sizes down to one big blind, their strategy is practically incomplete.
While it is true that solving a smaller stack size is cheaper, you have to solve many stack sizes from 1 to N to get good coverage across the space of all play.
I've been following this work for over 15 years, and they certainly deserve credit for what they did. But what they have done falls short of the banner headline.
I'd have to say that playing poker for long stretches of time is at least as, and possibly more so, than grinding for 12-16 hours writing code.
Even still, it's a stretch to call it unfair. It is brains vs. AI, after all, and that will always be true of brains no matter the task - fatigue is a factor, and this is where AI will have important advantages in the future.
absolutely not, these players play online, there are no tells online
So in that sense , online poker is very similar to live poker, the strategy doesn't change very much
Against a fish, yeah, I'll ham it up and they eat that stuff. If they're on the edge of a decision, some good acting can push them the direction you want.
I guess I'm agreeing with you -- a pro would never pay attention to my "tells".
Maybe, but it seems like impressive performance nonetheless.
That's true for most any level of serious player, but I see it often in casual games in peoples kitchens with casual players who don't play often or are just starting. Even I can't control myself sometimes when the adrenaline starts pumping. I'd rather say lack of self control is just one of the most amateur things you can do in poker, and what's nonsense is the idea that it's a meaningful aspect of any kind of serious poker game.
These are online players. They're playing the game they are best at.
Second, in the WSOP it is key to exploit weaker opponents. This bot was able to find almost perfect play against expert opponents but exploiting weak play is a different ballgame, especially if you are facing a mix of strong and weak players at a multi handed table.
There's little heads-up play in live games, other than at the end of the tournament, and the stack sizes are completely different. These are not tournament players, they're heads-up specialists. They most assuredly play online the majority of the time.
Second of all, the computer doesn't read the human players, why would it matter? It's all up to actual strategies at that point.