How A.I. Conquered Poker
nytimes.com
nytimes.com
This is actually one of the best poker articles I've ever seen in generalist media. Not too clickbaity, reasonable high level overview of game theory, a (very accurate IMO) quote from old pro Erik Seidel about the state of the game just 15 years ago, a discussion on variance vs. EV and results, and most importantly, an emphasis on math, randomization techniques and emotional control rather than the TV image of staring in someone's eyes and reading their soul. Probably the one biggest misconception people have is that pros have sick reading abilities since TV likes to emphasize staredowns, when the actual single biggest skill long term pros have is the ability to lose hand after hand for hours and still play their best game.
Incidentally, the anecdote at the top of the article is pretty intuitive game theoretically. Basically on the river you need to bluff with some portion of your hands, or else nobody will ever call when you have a good hand. The natural portion of your hands to bluff with is the absolute worst ones - you don't want to bluff with your middling hands because you have some small chance of just winning a showdown when it checks around. On a board of Kc4c5c2d2c, 7d6d is quite likely the absolute worst hand you can hold given the action that's taken place, therefore it's the one you bluff with. (For some pot/stack sizes it's possible that you bluff so rarely that you have to choose between 7d6d/7h6h/7s6s, but that's getting into details.)
1. Emotional control and pushing through the variance are key, I would play all nighters and it was mostly a subtle rise up vs big pots (except the one that did me in)
2. You absolutely do read players, I didn’t need to be a pro to learn how to observe the degenerates. Even decent players signaled things here or there (obv not how they eat an oreo , but I had a knack for knowing when someone was off their style). And it’s not even close wit a normal group of friends.
3. the sharks and their meta game… I had run up to about 50K within 3 weeks from my starting 2-3K. Had the pipe dream. I thought I had a growing bond with the pros. When I say pros I’m not talking grinders, I’m talking consistent big winners. About a handful of them. None of them tv personalities. - There were few times we’d run into each other, but I picked off a few donk bets here or there. However their persistence is what did me in. Went the entire month observing their advanced play style, which the meta was harsh, there were few rivers that mattered more than pre flop and continuation/donk betting. This style of play was like nothing I was seeing on tv. But I was winning with them. The last time I played, about a month into my run, I was on a 12h 25/50 game winding down, turned 30k into 70k (had 20k sitting back) pulled QQ and went up against a pre-flop donker. I face a 3 bet and get that gut feeling but put on lower pair or AK. Flop rainbows, large betting action. Turn, junk, I face an insane Re raise after testing a raise of my own, I know it’s off but it’s either AA/KK or AK at this point. My guts off but I’m too far in to quit. On the river I’m defeated, check, face bet, ultimately end up throwing 50K into the pot, lose to aces.
I never had a gut feeling so intense though. Like I was screaming at myself to get out. But I was caught.
Such a draining hand I never went back. I don’t have the discipline to grind games, and the tech industry pays more.
For cash games, I was only playing on the casino. For the 1/3 (1$/3$ blinds) the earnings rate seems absolutely abysmal and you have to account for rake. There are grinders who play these games, along with 2/5. My assumption is any grinder playing here is down under or not making enough to compete with a professional career in the tech industry.
Next is 2/5. Here I could’ve sustained a decent salary. A pro at these stakes showed me his past year earnings which was 180k.
It’s hard to tell with the 10/25, 25/50 pros. There were definitely dudes sitting on 500k+, maybe way more, chips that they kept in the casino and I assume similar amounts of cash at home. There were guys leaving with big nights but still losing 90k in a night. I’d hear about some being down under 200K at any point. There were probably some money launders. There was the occasional private investment fund guy that came in, was short with slick back hair, very loud mouth, would brag and show off one of his 17M personal banking account. Then there were the very quiet, calculated guys. It seemed that they were operating on some formula I never cared to figure out. They’d avoid big hands for the most part. Some form of grind I didn’t have the discipline for. I also didn’t have the discipline to handle large swings at those stakes, nor the bankroll, or desire to play lower stakes to build a bankroll.
This is absolutely true, but let me add to this: live pros like Phil Ivey, who have played the game in person virtually their entire lives, do have sick reading ability. However, it mostly works against amateurs. Actual pros know how to hide their tells better. If you take an online pro who has been clicking at the computer all day on multiple tables, that pro will not have anywhere close to the reading ability of Phil Ivey.
[1] https://youtu.be/dS_uv88YuPs?t=179 [2] https://youtu.be/IN1bW4jo4i4?t=72
It would be interesting to see a GitHub style graph of poker activity, with won and lost corresponding to added lines and removed lines.
My highest and lowest sessions were in the 5 figure range (cash game player, not tournament), so I was never up in the really nosebleed stakes, but I was high enough up that a lot of the nosebleed players would play in my games. Some of them have an impressive ability to play their A-game even when switching between games an order of magnitude in stakes apart. Some of them, uh, don't. People you've seen on TV are overwhelmingly more likely to be in the second category and were generally welcome in my games, at least for monetary reasons. (There was an old screenshot of Phil Hellmuth sitting at a full 9- or 10-handed 300/600 limit hold'em table online. He's out of chips and every other player is sitting out, waiting for him to reload. I have a pretty similar live story about him.)
I don't understand this logic, can you elaborate a bit more? What do you suggest to do with middling hands then?
- Worth betting, because you have the best hand
- Worth calling, because your hand is good enough to beat a bluff (and maybe some value bets as well).
- Intending to fold, because your hand is bad, but maybe it's good enough to win if the opponent's hand is worse and they don't bluff.
- Worth bluffing, because your hand is so bad it can't win otherwise.
The middling hands would literally be the middle two buckets in this example. Call with the better ones, fold with the worse ones. (To complicate this more, in a real world situation the worst part of the "intending to fold" bucket might become a "planning to bluff-raise" bucket, due to similar logic as to why you bluff with your worst hands.)
You try to get to a showdown. With middling hands there's a chance the opponent has a worse hand. It's when your hand is so bad your opponent has you almost certainly beat that you get to the bluffing territory, as that's the only way for you to win.
At least, that holds if you assume the opponent will always call with hands above a threshold and fold with hands below it. Which is correct play (though kindly ignore raises, for simplicity - but they don't blow up the argument). Only if they behave very strangely, e.g. folding with the very best hands and worst hands but calling with the decent ones, can I see a better connection with semi-bluffing on early streets.
People used to make fun of me for playing limits and sucking out on people. And no limit I would have some crazy strategies to fake out a table and make them anger-bet into my connected straight flush draws, when I was “representing” a two pair etc.
On large tables, often you want many people to stay in, until the river, because then the pot is big enough that multiple people will go all-in. You just have to be very sure you probably have the best hand (such as a flush that you were drawing for). I woul actually “goad” people with raising small amounts, whenever I had suited connectors for example, and if I hit a draw or the nuts on the flop, I would act as if I am “protecting” a top pair or so. On the other hand, when you are on a draw and if you want to see the next hand on the turn, you have to bet the flop much harder, because then it signals to most “basic” people that you’re “protecting” a hand — therefore gaining you an informational advantage when you do hit — as well as the “intimidation” factor that lets you “check” the turn to see the river — because they often think you’re checkraising them and also check, giving you another free card. You pay more on the flop but have two free cards after. Also If you don’t hit your draw, sometimes they didn’t hit either, and you can take it down with an aggressive bet since you were “representing” that you flopped a strong hand. To summarize:
1) play large tables
2) play suited connectors, or suited Ace-something
3) limp in or if you are late-position, double the bet to make people call and grow he pot and be more excited to bet later, you also represent a reasonably high pair
4) after a flop - if you are one card away from hitting a very strong draw (eg ace high flush) then bet hard on the flop — late position is always best for this… otherwise you have to bet less to keep people in
5) regardless of whether the turn card makes your draw, check on the turn … you may have to call any bet if multiple people stay in, and play the odds
6) if you made the nuts on the river, checkraise all in (since you represented that you are protecting a flooped hand).
7) there were also timing factors, people get more annoyed if you take too long. And you can play multiple tables online until you hit these kinds of starting hands
People said I was playing wrong but I intuitively felt like this would give a lot higher payoffs than just straight play with no bluffs.PS: Against a really intelligent table where people don’t churn, you’d have to leave after a couple big wins and do the same thing elsewhere. The whole idea in poker is to take advantage of most people just executing some common strategy or mindset.
No winning players nowadays would advocate for a strategy that doesn't involve bluffing.
Not to be harsh, but just to point it out for any novices who might be seduced by the simplicity of what you propose: the "strategy" you outline is simply nonsense and would not be winning in the long run even vs a fairly weak field. Maybe vs the level of play in the pre-solver world it did okay (though I doubt it would win except vs the very weakest fields), but today it would simply be burning money, even at microstakes.
Some of the heuristics you describe work in some situations, but you make no mention of accounting for other players positions, other players ranges, or the cards on the board; it is simply based on your hand and your perception of the player populations tendencies (which have changed dramatically since the pre-solver days). Ignoring the majority of the publicly available information in a given hand, in this game of incomplete information, is a grave strategic mistake.
I noticed in the article that piosolver can incorporate a mixed strategy with multiple bet sizes. Pretty neat, is that a recent feature? As I rememember it, a few years back people would discuss what the optimal bet size was in various situations with the implicit assumption that there is just one.
What happened in chess is now happening in poker: at first people were all about the best theoretical move and then gradual shift to "not necessary the best just sound enough but less likely to be analyzed by the opponents" started happening.
As to the first part of your question: the solutions are pretty much perfect. The only assumption is about possible bet sizes. As to accuracy we measure it in theoretical exploitability (how much a perfect adversary who knows our strategy exactly could win against us). You can easily go so low that even a theoretical perfect adversary wouldn't get close to beating the rake vs the solution, even at high stakes.
Disclaimer: I'm doing phd in this area, generalizing to harder games than poker.
We considered it very unlucky that the science paper was published as we thought more people will implement solvers then (ours was ready in mid 2014 and we released and a working solver with a GUI in early 2015, 3 months after the Limit Holdem paper). As it turned out though it didn't really matter much.
I had an idea to code a poker solver around 2008. Unfortunately I was yet to learn to code back then. The biggest challenge was to overcome implementation issues. Once I got the crucial part to be fast enough I knew the solver is possible. I think not having much background in the field allowed me to come up with more natural (and better) tree representation and memory efficient (even if not the fastest) algorithm as I was thinking more as poker player than a computer scientist back then.
It’s an interesting history.
> Koon will often randomly select which of the solver’s tactics to employ in a given hand. He’ll glance down at the second hand on his watch, or at a poker chip to note the orientation of the casino logo as if it were a clock face, in order to generate a percentage between 1 and 100.
Another classic is to use the suit of your cards, but that has problems with being correlated with the state of the board. It works fine preflop.
The two most popular games (no-limit Texas holdem and pot-limit Omaha) are still unsolved.
It's shocking that the reporter didn't mention these results or anything else more recent than 2015.
Further, thousands of the hands that Pluribus played against the human pros are available online in an easy to parse format [0]. I've analyzed them. Pluribus has multiple obvious deficiencies in its play that I can describe in detail.
It seems like it's very difficult to set up any kind of proper repeatable and controlled experiment involving something as random as poker. Personally, I would be much more convinced if Pluribus played against real humans online and was highly profitable over a period of several months. This violates the terms of service / rules of many online poker sites, but it seems like the most definitive way to claim terms like "solved" or "superhuman"
https://www.science.org/doi/abs/10.1126/science.aay2400
https://par.nsf.gov/servlets/purl/10077416
and it just says
> Finally, we tested Libratus against top humans. In January 2017, Libratus played against a team of four top HUNL specialist professionals in a 120,000 hand Brains vs. AI challenge match over 20 days. The participants were Jason Les, Dong Kim, Daniel McCauley, and Jimmy Chou. A prize pool of $200,000 was allocated to the four humans in aggregate. Each human was guaranteed $20,000 of that pool. The remaining $120,000 was divided among them based on how much better the human did against Libratus than the worstperforming of the four humans. Libratus decisively defeated the humans by a margin of 147 mbb/hand, with 99.98% statistical significance and a p-value of 0.0002 (if the hands are treated as independent and identically distributed), see Fig. 3 (57). It also beat each of the humans individually.
Surely the correct strategy here is for the human players to collude to give as much money as possible to a single player and then split the money afterwords, no?
Also, the fact that they players can only gain money without losing anything likely changes their play somewhat. By default I'd assume (and have generally observed) that most players on a freeroll (or better than a freeroll really) tend to undervalue their position and gamble more than is usually wise.
I'd definitely be interested in seeing a "real" game where the humans are betting their own money.
Top pro poker players understand the value of money. They weren't treating it as a freeroll and anyone that has seen the hand histories can confirm that.
But if you are rooting against the machines, don't worry: it is almost certainly impossible to calculate a full equilibrium policy for no limit multiplayer, so we will instead be debating over the virtues of various types of imperfection for a long time. And even if an Oracle gave us convenient access to equilibrium strategy, it would still not be the optimum at a table full of imperfect players. Your poker game is safe for a while!
I don't think you're correct saying it doesn't affect poker as people were able to notice and analyze this before solvers. It's true though that no-limit Holdem as played today (two blinds,no ante,deep stacks) is likely not strongly affected by the phenomena. I don't agree Pluribus experiment shows much when it comes up practical play. Not enough variety of skill levels, not enough hands and not enough time for metagame (people adjusting to how others play) to develop. I do agree pure equilibrium play is most likely not terrible in cash game nlhe but definitely not in poker in general.
> Machines have raised the stakes once again. A superhuman poker-playing bot called Pluribus has beaten top human professionals at six-player no-limit Texas hold’em poker, the most popular variant of the game. It is the first time that an artificial-intelligence (AI) program has beaten elite human players at a game with more than two players
Pluribus had approximately zero impact on that.
I wrote probably the first online texas holdem game, played on IRC, around 1994. Back then in rec.gambling blackjack teams were forming and instead of joining a blackjack team I started to play poker, because it was more fun than counting cards. I wrote the game so I could play against other people and was immediately a winning player in low limit games of the time in lake charles and vegas. However even in the 90s, it was about statistics.
I do regret not allowing people to play for money (people were constantly asking).
Even back then the primitive bots were mostly killing human players.
The online poker community is extremely hostile towards poker bot software in general; you would be very hard-pressed to find an existing poker website that would be willing to encourage and host such tournaments.
It's a fun project though, especially trying to design in a way that proves I'm not manipulating the deck behind the scenes.
Email in profile if anyone is up for being an alpha player.
I see it like HFT. You may have multiple bots running multiple playing strategies and an entire tournament could occur in minutes/hours. A large weekly tournament might even be like a sporting event for hackers.
The human vs bot scenario you’re interested in would be possible on the same platform with a longer timeout period.
This article is about how humans have been memorizing the variations - akin to chess openings - in order to play as perfectly as possible.
If you want to quickly improve to a proficient level there's no quicker way than to read Theory of Poker and Hold 'em for Advanced Players. Both of these books focus primarily on limit poker, but the concepts are critical for no limit as well. And you'll realize there's a lot more strategy and nuance in limit poker than you thought.
Training videos
1-on-1 coaching
PioSolver (and other similar software)
For any of these to stick, you need to spend some amount of time studying by yourself; just consuming learning material and playing isn't enough.
Applications of no limit by Janda or mathematics of poker are recommended by a different reply here. I would caution that these are extremely academic and extremely dense texts that would be a very tough read for a newer player. Mathematics is less practical and more of a math book than a poker book and applications is Janda showing how to work out solver-like solutions before solvers existed and also contains a lot of math. I think the other two above are more practical and aren't going to lead you to put the book down a 10th of the way through.
Applications of No-Limit Hold 'em, by Matthew Janda - this one goes into specifics about how to think on each street, how to build a good strategy etc. Some of the details will be different from the latest theory because the book was released in 2013 but it's still a very solid read if you want to level up.
After that, perhaps watch some Doug Polk on Youtube to see how he uses the concept of putting people on ranges of hands and what his thinking is in a specific spot.
I believe Alphastar would generate more interesting strategies if we limited alphastar to a bit below human APM and forced it to emulate USB K+M to click (instead of using an API, which it currently does) and adding a progressively increasing random fuzzing layer against its inputs so that as it clicks faster the precision/accuracy goes down.
By "interesting strategies" I mean strategies that humans could learn to adopt. Currently its main strategy is "perfectly juggle stalkers" which is a neat party trick, but that particular strategy is about as interesting to me as 2011-era SC AI[0]. Obviously how it arrived at that strategy is quite interesting, but the style of play is not relevant to humans, and may in fact even get beaten by hardcoded AI's.
I'm also very curious what Alphastar could come up with if it were truly unsupervised learning. AIUI, the first many rounds of training were supervised based on high level human replays -- so it would have gotten stuck in a local minima near what has already been invented by humans.
This may be relevant if Microsoft reboots Blizzard's IP. I would love to have an alphastar in SC3 to play against off-line, or have as a teammate, archon mode, etc. I think all RTS' are kind of "archon mode with AI teammate" already. The AI currently handles unit pathing, selection of units to attack, etc. With an alphastar powering the internal AI instead, more tactics/micro can be offloaded to AI and allow humans to focus more on strategy. That seems like it would be super cool.
Examples: "Here AI, I made two drop ships of marines. Take these to the main base and find an optimal place to drop them. If you encounter strong resistance or lots of static defense, just leave and come back home"
"Here AI, use these two drop ships of marines to distract while I use the main army to push the left flank. Take them into the main, natural, or 4th base -- goal is to keep them alive for as long as possible. Focus on critical infrastructure/workers where possible but mostly just keep them alive and moving around to distract the opponent."
0: Automaton 2000 AI perfectly controls 50-supply zerglings (2.5k mineral) vs. 60-supply (3k mineral, 2.5k gas) siege tanks: https://www.youtube.com/watch?v=IKVFZ28ybQs
No Alphastar definitely had misclicks, and it had a maximum cap on APM regardless of average that was far lower than the max burst of APM (or even EPM) of top players. When I have the time I can go dig up some games where Alphastar definitely has misclicks, and I believe the Deep Mind team has said before that it will misclick. Its APM limits are already lower than pros both on average and in bursts (and are reflected in its play, Alphastar will often mis-micro units in larger, more frantic battles such as allowing disruptor shots to destroy its own units, but it will never make the same mistake with much smaller numbers of units).
> Currently its main strategy is "perfectly juggle stalkers"
Definitely not. That was its strategy in its early iterations against MaNa and is no longer feasible with the stricter limitations in place. Its Protoss strategy is significantly more advanced than that now (see its impressive series of games against Serral with an amazing comeback here: https://www.youtube.com/watch?v=jELuQ6XEtEc and a powerful defense against multi-pronged aggression here: https://www.youtube.com/watch?v=C6qmPNyKRGw) (and of course by "now" I mean when Deep Mind took it off the ladder). Both of these involve an eclectic mix of units with Alphastar effectively using each type of unit and varying it in response to what Serral puts out and its own resource constraints.
A lot of commentators have difficulty distinguishing Alphastar from humans when the former plays as Protoss (its Terran and Zerg play is weaker and often more mechanical).
> I mean strategies that humans could learn to adopt.
My main takeaways from watching Alphastar were "pros undervalue static defense and often have a less than optimal number of workers (where Alphastar's seeming overproduction of workers lets it shrug off aggressive harassment)," but I don't know if those have picked up in the meta.
If the question is rather "what characteristics of Alphastar's Terran and Zerg play style make me say that its Terran and Zerg play is worse than its Protoss play," the simplest answer is that Alphastar just feels a lot more like a bot. Unlike when playing Protoss, it seems to get into certain "ruts" of unit composition and tactics that are a bad match for the opponent it's facing and can't seem to reactively change based on the game is going, whereas with Protoss it seems more than happy to change its play style over the course of the game based on what the opponent is doing.
Edit: 7 minutes after writing this I re-read the original paper[-1]. https://arxiv.org/pdf/1708.04782.pdf page 6 and 7 make it clear that DeepMind limited themselves to SpatialActions, so they cannot tell units "Attack Carrier" but have to say "Attack point x,y" (and x,y also has to be determined visually, not through carrier_of_interest->pos.x ). It's still not clear in the paper if any randomness is added to Attack(x,y).
Additionally, I have some serious concerns about assuming that the design decisions made in this 2017 paper were actually used in the implementation of the 2019 Alphastar demo vs TLO and MaNa. The paper claims "In all our RL experiments, we act every 8 game frames, equivalent to about 180 APM, which is a reasonable choice for intermediate players." I would agree with this choice! But [5][6] indicates that Alphastar's APM spiked to over 1500 APM in 2019! And even in moments when a human reaches that APM, their EPM would be an order or magnitude lower, whereas Alphastar's EPM matches its APM.
Original post:
Thank you so, so much for adding to the discussion! Would love to chat more about this if you see my reply and feel like it.
Regarding "mis-clicking", my understanding was that AlphaStar used Deepmind's PySC2[0][1], which in turn exposes Blizzard's SC2 API[2][3].
Here is the example for how to tell an SCV to build a supply depot:
Actions()->UnitCommand(
unit_to_build,
ability_type_for_structure,
Point2D(
unit_to_build->pos.x + rx * 15.0f,
unit_to_build->pos.y + ry * 15.0f
)
);
where unit_to_build->pos.x and unit_to_build->pos.x are the current position of the SCV and rx and ry are offsets. It's possible to fuzz this with some randomness, and indeed in the example, rx and ry are actually random (because the toy example just wants to create a supply depot in a truly random nearby spot, it doesn't care where). But the API doesn't attempt to "click" on an SCV and then use a hotkey and then "click" somewhere else. The API will never fail to select the correct SCV. It will also build precisely at the coordinates provided.Point 1: Even if DeepMind added a fuzz to this method to make it so AlphaStar can "misclick" where the depot gets built, it cannot accidentally select the wrong SCV to build that depot. (Possibly wrong, as they could be using SpatialActions, see below)
Point 2: Most bot-makers wouldn't add a random fuzz to the depot placement coordinates to make their AI worse and I'd be super surprised if there was hard evidence somewhere that Alphastar had such a fuzz. (This is my main concern.)
My personal conclusion was that anything which looks like a "misclick" is, in fact, a "mis-decision". A human can decide "I want my marines to attack that carrier" but accidentally click the attack onto a nearby interceptor. I didn't think Alphastar could do that because I assumed it would use the Attack(Target: Unit) method instead of Attack(Target: Point) in that scenario -- and even if they used Attack(Target: Point) it would be used as Attack(Target: carrier->pos.x).
However, I realize now that they could be doing everything with SpatialActions (edit: it does, see paper[-1] pp. 6-7) (select point, select rect's)[4], and that they could have a implemented a randomness layer to make alpha star literally mis-click.
I suppose I would need to test this API and dive into the replay files to first see if its possible to discern the different between Attack(Target: carrier_of_interest) and Attack(Target: carrier_of_interest->pos.x). Then, even if Alphastar is using the latter, it's still not clear that there's an additional element of randomness outside of the AI/ML control.
Has anyone already done an analysis of the replay files on this level, or has DeepMind released hard info on how they're controlling the bot?
-1: https://arxiv.org/pdf/1708.04782.pdf
0: https://www.youtube.com/watch?v=-fKUyT14G-8
1: https://github.com/deepmind/pysc2
2: https://github.com/Blizzard/s2client-proto
3: https://blizzard.github.io/s2client-api/index.html
4: https://blizzard.github.io/s2client-api/structsc2_1_1_spatia...
5: https://www.alexirpan.com/2019/02/22/alphastar.html
6: https://deepmind.com/blog/article/alphastar-mastering-real-t...
7: https://ychai.uk/notes/2019/07/21/RL/DRL/Decipher-AlphaStar-...
If you look at an Alphastar Protoss game from the latter half of 2019, it's not relying on cheap tricks to win (such as the impossible stalker micro). Nothing it's doing leaps out as superhuman. Instead it just grinds down its opponent through a superior sense of timing and macro strategy. The two games I linked against Serral it wins by punishing when Serral overextends his reach or by altering its unit composition to better fit what Serral throws at it, rather than some ungodly micro. Nothing it's doing there couldn't be done by a human. In fact I would say in most of the battles, Serral's micro was better than Alphastar's.
Now it's also worth pointing out that Serral is playing on an unfamiliar computer, rather than his own, so there's a bit of a handicap going on and even Alphastar Protoss will still lose to humans, so it's not superhuman, but it's definitely an elite player and its play style is very difficult to distinguish from that of an elite player.
And then of course it's also imperfect information both in the sense of your opponent hand but also his deck. The cardpool is also very large for some formats.
I actually don't think it's solvable just by throwing MCTS at it with todays hardware but would love to know more about this, if someone else has more insight please reply.
EDIT: Oh and there is also the meta-game / deck building aspect. If you are going to win a tournament you have to have favorable matchups against most players in the room.
Even the meta-game/deck building aspect doesn't seem all that insurmountable as it doesn't seem fundamentally different from say a build order other than that it cannot change dynamically on the fly.
I think it'd be much harder for the AI to do deck building in a vacuum. You could model it as every game starting with you building your deck, but I can't see that converging stably.
He was able to reach Mythic, the highest ranking tier on Magic Arena. Of course this is a different problem to actually playing the game (and probably significantly easier). That being said, this is one guy doing it as a side project with restricted resources.
I imagine that MTG could played quite successfully by an AI if someone where to dedicate the resources. Imo much of the difficulty is in laying the groundwork. Large amounts of data don't exist publicly and laying the framework for a bot to play itself would be quite difficult (and then the computation costs would be extremely expensive).
https://arxiv.org/pdf/2112.03178.pdf
Abstract:
"Games have a long history of serving as a benchmark for progress in artificial intelligence. Recently, approaches using search and learning have shown strong performance across a set of perfect information games, and approaches using game-theoretic reasoning and learning have shown strong performance for specific imperfect information poker variants. We introduce Player of Games, a general-purpose algorithm that unifies previous approaches, combining guided search, self-play learning, and game-theoretic reasoning. Player of Games is the first algorithm to achieve strong empirical performance in large perfect and imperfect information games — an important step towards truly general algorithms for arbitrary environments. We prove that Player of Games is sound, converging to perfect play as available computation time and approximation capacity increases. Player of Games reaches strong performance in chess and Go, beats the strongest openly available agent in heads-up no-limit Texas hold’em poker (Slumbot), and defeats the state-of-the-art agent in Scotland Yard, an imperfect information game that illustrates the value of guided search, learning, and game-theoretic reasoning"
This probably shouldn't surprise you. If you are researching how to attack a new/difficult class of problems, typically you look for the simplest versions of that class first.
The remaining challenge is getting it to play well with human partners. Doing that requires modeling human conventions rather than learning weird bot conventions. That's hard because while you can collect essentially unlimited data through self play, it's hard to collect a lot of data playing with humans using reinforcement learning. AI algorithms are really bad at sample efficiency.
Frustratingly, this seems to be an unsolved problem. Pluribus is only 6-max, requires fixed starting stack sizes and requires a lot of hardware.
But also I think Pokersnowie is a bot made with more traditional ML methods (i.e. training on hand histories rather than solving an entire game tree with CFR), which basically fits the description.
For poker, people can still play face to face, and it being a game of mixed strategies means it's probably a lot less important that poker bots are better than people than in chess. Top chess players use the computers to set up preparation bombs on their opponents, with rebuttals to moves for specific positions created by the computers. That doesn't work in a game of imperfect information and mixed strategies.
The norm for online chess is to not play for money, whereas the norm for online poker is to play for money, often for large sums.
Of course the game with higher stakes is more likely to attract cheaters.
That is definitely not the norm. Most people are playing recreationally. Heck it is illegal to pay poker online with real money in 44/50 states in the US.
I don't see this, I used to really enjoy online poker tournaments for play-money. It's a fun and strategic game in and of itself with nothing more than minor bragging rights on the line, to me at least. I honestly don't see why it should be any different from chess in terms of purely recreational play. (You could say chess is a deeper game, and that's probably true in some ways, but that's not a positive feature for casuals like me with no desire to invest many hours/years into mastering the game.)
Would you play the same way if you were betting your life savings versus $5?
e.g. I'll lose 100% of my matches against a Chess GM but against the best poker player in the world I could go all in preflop every hand and still win the match ~20% of the time.
Only way to counter this variance is to play lots and lots of hands (Law of large numbers) It could take weeks/months in a single matchup to determine the better poker player.
In terms of playing for real money, in general you'll find more challenging competition at higher stakes. But then there's an incentive for bots.
Think about playing rock paper scissors agains an opponent who always chooses rock. If you still play the optimal strategy (choosing randomly) you gain nothing...
This is not true. If you play optimal strategy, you will win against any opponent except one that plays optimal as well, in which case you'll break even. But, of course you'll win a lot more if you are able to adapt your strategy to exploit any weaknesses that you have detected.
And detecting weaknesses in your opponent(s) in live play is something that humans will remain better at than AI for quite a while. Because it requires not just careful analysis of your opponent's actions but some contextual information as well. E.g., how old is your opponent, is he experienced or not, is he drunk, have you experienced similar players like hime before, etc.
To be more precise: you only need to replicate not mixed (pure) plays of the optimal strategy to not lose against it. Your frequencies for mixed actions can be completely off though.
How does the example you're responding to not win 50% of the time?
Rock v Rock = Tie
Rock v Paper = Loss
Rock v Scissors = Win
The optimal game theoretic play of randomly choosing rock paper scissors is inferior play against this particular opponent. All that game theoretic perfect play gets you is the benefit of getting at least a tie. Possibly more, but not always more for non-perfect play.
It's very easy to play crap in poker
If the optimal strategy is a 50-50 split against someone else playing the optimal strategy and at least a 50% win against someone who is not, it doesn't follow that if they other player plays crap every game that the optimal strategy is best against them.
Identification of the crap player means that you increase your bets vs what you would normally bet for perfect play. Perfect play is bullet proof, but it isn't necessarily maximal yield against a non-perfect player even if thats the rough trendline.
>Identification of the crap player means that you increase your bets.
That's not true. There are a lot of different ways to exploit bad players depending on how bad they are. Also a lot of weak players these days have much more subtle leaks (e.g. never bluff-raise the river) and it takes some skill to spot and exploit them.
> > You can only win in poker if you recognize how your opponent deviates from the optimal strategy and then play a strategy to exploit him. > This is not true.
You were right here when you said "This is not true." I still think there is more nuance than the discussion as a whole credits, but overall I think you were more right than wrong in our discussion and I was more wrong than right. Thanks for the chat.
I mean maybe this is a trivial thing to solve for you, in which case kudos, but its pretty neat for a lot of onlookers(and participants).
The statistics part anyone with basic math skills can figure in their head - that doesnt guarantee any win.
So I checked out this tool, and the team describes themselves as "programmers interested in algorithms"[0] ... what is the difference between A.I. and algorithms?
* Every computer program has algorithms, it's an extremely general term for "the idea behind how the computer will solve the problem". Advanced algorithms are typically those that took a lot of human effort to come up with.
* Machine Learning refers to a specific class of algorithms where the computer automatically figures out (part of) what it should do based on data.
* Deep learning is the subset of ML that uses deep neural networks.
* Artificial Intelligence is a marketing term, and is actually about the _problem_ being solved, not the technology being used to solve it. In particular, AI is any computer program that solves a problem people would previously expect can only be solved by a human applying creativity and/or intelligence, like playing a game, understanding natural language, or creating artwork. This definition is obviously a moving target as expectations change.
Deep learning is currently the most powerful and general toolkit for solving AI problems, so the concepts tend to get mixed together pretty frequently, but I personally like using the definitions above to keep things straight.
In the specific case of cruise control, I don't think maintaining speed accurately is typically viewed as "difficult" for a computer today, so it would be deceptive to refer to as AI. 100 years ago that may have been different.
Would you put statistical analysis "algorithms" in the same category as "Machine Learning"?
In modern vernacular it has become a synonym for machine learning or sometimes specifically neural networks.