Professional poker players know the optimal strategy but don't always use it
scientificamerican.com
scientificamerican.com
It's interesting how this affects play. If a player is bluffing slightly less than they should be, the adjustment is drastic, you should actually never call with hands that do not beat their value hands. If they are bluffing optimally, you are supposed to call using what is known as "minimum defense frequency". What's interesting about this is the minimum defense frequency is based on what the strongest hands you can possibly have in that situation are and the opponents possible hands do not even factor into it. It's required to prevent your opponent from bluffing with any two cards profitably.
To do the math, if you are on the river and the opponent bets 100 into 100, for this to be profitable for them they need to win 50% of the time or more. If your opponent is bluffing optimally, you need to call with 50% of the strongest hands you have in that specific situation (if you don't know what hands you have in a specific situation that's a problem) and sometimes they can be dogshit like King high.
But, very important to note, very few players actually bluff enough and if they bluff less than they should you should only ever call when your hand actually beats their range of possible value hands. (value vs bluff is kind of a difficult thing to communicate, generally it's value if you want your opponent to call)
Most players don't bluff enough as a result of most players calling too much! When they call too much you should obviously not bluff! This leads to very boring games of poker.
This is actually not quite true. MDF is purely a formula based on pot size and the bet size (pot size / pot size + bet size). The fact that it doesn't consider various ranges is why it's not really useful - it was a simplified formula used to try to understand the game before solvers existed.
There are situations where your opponent can bet any two cards profitably and you do have to fold - imagine they bet the size of the pot, but have the better hand 99% of the time, you're simply forced to let them bluff the 1% of the time they're bluffing. MDF is a pre-solver concept and not an especially useful concept in the modern game.
I know that before the river there are range advantages that make defending mdf a losing play.
If you open up the solver, and give one player only Ace-Ace as their starting range, and the other player a pair of twos, and the board Ace-Ace-Three-Three-Three, then the pair of twos will fold 100% on river and will not call at MDF.
So you can't make a ton of mistakes say "MDF" and call off, you have to have done the right things in previous streets to end up with a range that can call at MDF. That range (and those street actions) require an understanding of GTO (and the adjustments needed when someone isn't playing GTO).
It's true that "bluffing has minimal value against best play", but no human is in that situation. Even super-GMs will play "bluffs" if they are behind (or playing a lower-rated player and sure they can recover later. or just for fun if the stakes are low).
And that's not even mentioning optimal strategy under time pressure. For instance some of the Lichess tournaments are structured such that winning fast is more valuable than winning slowly because the resulting score comes from how many wins you can get in (or in other cases, how big of a streak you can get). So people will play in a way that optimizes for winning quickly by taking big bets / bluffing / creating chaos with un-calculated gambits, especially if they have a good reason to believe they're better than their opponents.
It wasn't the most accurate game of the tournament, but the most instructive as far as the psychology of chess goes
In live cash games, speed matters, you want to take all the available chips before the fish realise they're out-matched, but to protect yourself the optimal play, if you could memorise it, would be safer because it can't be exploited yet it does take the opponent's chips.
Poker players are gamblers. So "safer" isn't really what they were going for anyway.
This is not really what most people mean when they say “optimal strategy”. It’s true that exploitative play will make more money if you know exactly what your opponent is doing and if they keep doing it despite it not working. Neither of these will generally hold in an actual game.
The reason it’s called “optimal strategy” is because it works no matter what your opponent is doing. It will not make as much as a strategy tailored perfectly to your opponent but it will never lose to anyone under any circumstances, assuming infinite games (assuming infinite games just so we can ignore variance). The worse case the strategy has is break even to anyone else using it.
All of the good players identify the worst, riches player at the table, and they all take that players money.
Once they're out, you leave the table too
Ok, now that I (I think) understand your comment: I don't think you have to know exactly what non-Nash strategy someone is playing in order to exploit it. I don't think trying to estimate how someone is likely deviating from the Nash equilibrium, in order to try to exploit it, is necessarily a mistake. I think it could be feasible for someone to get better returns on average by noticing off-Nash patterns of play in other players, than playing Nash regardless would? (Not that I could win this way. I couldn't.)
If one player isn’t playing a strategy that is part of any Nash equilibrium, then the best response might also not be part of any Nash equilibrium.
If all other players are playing a strategy from a given Nash equilibrium, then you can’t do better (in expectation) than you would if you were to play the strategy for you in that Nash equilibrium.
(A game may have multiple Nash equilibria. Possibly one such equilibrium could be better for you (or for everyone) than another.)
These absolutely hold, 90% of casual poker players have the exact same strategy problems of not bluffing enough and calling too much.
What is your formal definition of "optimal strategy"? A Nash equilibrium is considered optimal in the sense that it's a state where no one can gain an advantage by deviating from the equilibrium.
Sure, if your opponents don't play the Nash equilibrium, there is room to exploit that deviation and potentially gain more than what you would get from playing the Nash equilibrium. However, you also make yourself exploitable in return, so I don't think you're presenting the whole picture here.
Assuming at least one of opponents is not playing Nash equilibrium (which is a very solid assumption), playing the Nash equilibrium becomes suboptimal as it doesn't exploit the exploitable as much.
It is not really related to your smartness or luck (doesn't apply to _everyone_ of course but I'd wager that the average HN reader is already smart enough for poker)
Any book/other resource recommendations for brushing up on this stuff?
[1] https://wasm-postflop.pages.dev/ [2] https://github.com/b-inary/desktop-postflop
I don't know, poker theory is all about optimal ranges and Nash equilibrium, but there's something satisfying (and very practically important, since if all your opponents even understand the phrase Nash equilibrium you should find a different game) about trying to make the most money against an opponent who calls or bluffs way too much.
I've actually been working slowly on https://github.com/krukah/robopoker, an open-source Rust implementation of Pluribus, the SOTA poker AI. What I've found interesting is the difference in how I approach actually playing poker versus how I approach building a solver. Playing the game naturally consists of reasoning about narratives and incorporating information like hand history, play style, live tells. Whereas solving the game is about evaluating tradeoffs between the guarantees of imperfect-information game theory and the constraints of Texas Hold'em, finding a balance between abstract and concrete reasoning.
Basically, play one level - exactly one level - beyond where you peg your opponents at.
Poker is not about playing cards. It's about playing people. Cards is just how we keep it civil and not too personal.
Human pressure yes. Imperfect information no.
When we talk about Nash equilibrium for a game like poker, it's already based on imperfect information.
Piosolver, the first public solver and the one mentioned in the article, has this feature.
However, what often happens is if you lock one node, then several other nodes in the game tree over-adjust in drastic ways, forcing you to lock all of the, which may be infeasiable. As a result, Piosolver recently introduced "incentives", which gives a player in the game an additional incentive to take a certain action . For example, you may suspect your opponent calls too much and doesn't raise enough, so you can just set that incentive and it will include that in its math equations and give you something similar to an exploitative solution with a much simpler UX.
This feature was literally just introduced a few months ago so it's still very much an active area of research, both for game theory nerds, and people trying to use the game theory nerd research to make money !
I still haven't seen an AI for a turn based strategy game. There's AlphaStar, but it wins via APM, not strategy.
The point of the optimal strategy is that it's unexploitable so you can disregard the other player's actions (in the game or outside it) entirely.
All exploitative strategies are in turn exploitable.
Whether the assumptions of the Nash equilibrium (or any of the others) make sense for your situation in a game of poker is an empirical question, right? It's not a given that playing a NE means you'll be "perfect" in the human sense of the word, or that you'll get the best possible outcome.
The best superhuman poker AIs at the moment do not play equilibriums either, for instance.
However the situation with an AI powered competitor which uses exploitative play is identical to a human, the GTO play will gradually take their chips at no risk.
It's not that they're optimal but that they've chosen not to be optimal and so that's why they lose money against GTO.
The AI is at least unemotional about this, humans with a "system" easily get tilted by GTO play and throw tantrums. How can it get there with KToff? What kind of idiot bluffs here with no clubs? Well the answer will usually be the one that's taking all your chips, be better. Humans used to seeing exploitable patterns in the play of other humans may mistake ordinary noise in the game for exploitable play in a GTO strategy and then get really angry when it's a mirage.
When there are more players, there can be multiple Nash equilibria, and (unlike the two player case) combinations of equilibrium strategies may no longer be an equilibrium strategy. So it's no longer true that you cannot be exploited, because that depends on other player's strategies too, and you cannot control those.
(See this paper for instance: https://webdocs.cs.ualberta.ca/~games/poker/publications/AAM...)
I really don't know anything about poker AIs but could it be you are referring to Libratus and/or Pluribus[0]?
on top of this, you could think about augmenting these systems to exploit weaknesses in opponent strategies. there is some work on this, but I don't think it's done much. The famous systems that played against professionals don't use it, they just try to get as close to GTO as possible and wait for opponents to screw up.
But it sounds like that must have been either misunderstanding or some other part of the bot's algorithm I guess.
Whenever I search this stuff I get practical poker strategy guides, but none of them seem to define the term haha
Common expression is "deviate from GTO" where you know what the solver would do but decide to play differently.
Life is a lot more complicated in multiplayer poker. There are Nash equilibria, but potentially many with different payoffs, and you can't force your opponents to choose the one you're aiming for. So in that case, it's not so obvious what "optimal" means.
As for CFR adapting to opponent play: CFR could bias its compute resources towards really finely optimizing strategies for the most likely scenarios facing certain players, and it seems like this has been done during poker tournaments.
But within those situations, it would still be trying to more perfectly approximate the Nash strategy, vs. more experimental approaches which actually choose a different strategy to exploit opponent weaknesses.
Any exploitative AI also needs the ability to adjust in real time to a different exploitative strategy, which also needs to be not easily predictable, etc.
Draft, Standard and Modern are for people who want a real game without having to worry about playing too well.
I love it. Love cycling as well.
https://www.livepokertheory.com
I do personally dislike that GTO became the nomenclature , as I prefer "theory-based", since it causes this confusion, but trying to fight it at this point is hopeless because GTO is the search term people are using. And when people say they "play GTO" they usually mean "equilibrium" rather than "optimal against my specific opponents" which is "exploitative".
If you actually watch what the top players advocate for, everyone suggests you want to play exploitatively. However, there's one equilibrium solution and effectively infinite exploitative solutions, so equilibrium is a reasonble starting point to develop a baseline understanding of the mechanics of the game. It's tough to know how much "too much" bluffing is unless you know a baseline.
Furthermore, if you "exploit" people by definition you are opening yourself up to being exploited so you need to be very careful your assumptions are true.
Also, with solvers like piosolver, you can "node lock" (tell a node in the game tree to play like your opponent, rather than an equilibrium way plays), but there's many pitfalls, such as the solver adjusting in very unnatural ways on other nodes to adjust, and it being impractical to "lock" a strategy every node in the tree. There's new ideas called "incentives" which gives the solver an "incentive" to play more like a human would (e.g. calling too much) but these are new ideas still being actively explored.
Rock paper scissors is frequently used to explain GTO but it's not the best example because equilibrium in rock paper scissors will break even against all opponents, but equilibrium poker strategy will actually beat most human poker players, albeit not as much as a maximally exploitative one.
There's two other huge pieces this article glosses over:
1) It's as impossible for a human to play like a computer in poker as in chess - in fact far more impossible, because in poker you need to implement mixed strategies. In chess there's usually a best move, but in poker the optimal solution often involves doing something 30% of the time and something else 70% of the time. The problem is that, not only are there too many situations to memorize all the solutions, but actually implementing the correct frequencies is impossible for a human. Some players like to use "randomizers" like dice at the table, or looking at a clock, but I find that somewhat silly since it still so unlikely you are anywhere near equilbrium.
2) Reading someone's "tells" live is still a thing. While solvers have led to online poker to decline due to widespread "real time assistance", live poker is booming (the 2024 World Series of Poker Main Event just broke the record yet again) , and in person in live poker, people still give off various information about their hand via body language. From the 70s to the early 2000s, people were somewhat obsessed with "tells" as a way to win at poker. Since computers have advanced so much, it's fallen out of favor, but the truth is, both are useful. It's totally mistaken to think that advancement in poker AI , GTO , and solvers have rendered live reads obsolete. In fact, in 2023, Tom Dwan won the biggest pot in televised poker history (3.1 million) and credited a live read to his decision, in a spot where the solver would randomize between a call and a fold.