How I Used Professional Poker to Become a Data Scientist
medium.springboard.com
medium.springboard.com
While there are C(52, 5) different hands you can have, having 4 diamonds and a heart is the exact same thing as 3 clubs and 2 hearts, etc. so it collapses down to ~2.5k hands.
Now the next observation we make is that for most hands, suit doesn't matter, but if it does we still need a way to distinguish it, but again, a flush in hearts is the same as a flush in spades in terms of card ranking. The trick we now pull out is the Fundamental Theorem of Arithmetic, which says that for some natural number N, N can be factored into primes in exactly one way, excluding permutations of the factors.
We take advantage of this and assign each card a prime number. 2=2 3=3 4=5 ... A=41
Now, by multiplying the prime numbers together, we get things like a 2 3 A K Q hand is 2 * 3 * 41 * 39 * 37, and if they are all the same suit, you multiply by 43. You can then take these 2.5k hashes, rank them against each other in a lookup table, and know in constant time for the number of hands how hands rank against each other by just multiplying their card's values together with either 1 or 43 if same suit and seeing where it falls in the list.
I realize expected value is more widely used as the basis for these types of systems, but I always thought that was a fun trick.
Edit: I missed
>>> rank them against each other in a lookup table
I'm just confused about the prime factors thing now. Why not just hash the hand itself? "3459To" where o stands for off suit?
Permutable hashes - you could just sort them and hash again, but the prime multiply hashing is pretty neat way of mapping all items in any order to the same hash code.
You need to either hash all the permutations (making your lookup table 5! times bigger) or sort the hand by card value. It is faster to multiply the 5 prime factors in linear time rather than the N log N of sorting.
For example how do these two hands compare?
A 2 3 4 5 2 3 4 5 6
Maybe there is a way we can extend this to work for straights.
Does anyone else on HN have more resources like this, applying data science to poker. I will google, but on forums like this one, I find personally recommended resources to be very helpful.
It is most useful for analysing your own game - you have access to your entire hand history so it is extremely valuable for finding out your weaknesses.
Databases of IRC poker matches hands: http://web.archive.org/web/20110205042259/http://www.outflop...
(Actually uses web-archive and handhq.com data.)
For reading: http://poker.cs.ualberta.ca/publications.html
Even for your own hands, where you see every result, there's so much noise on the individual hand level that it's hard to do good direct comparisons. You can easily have a million hand sample where, for example, you make more money with 22 than 88. Some people will (most likely erroneously) conclude that they there's something wrong with how they play 88. The more likely explanation is that even with a million hands, once you break out how often you get dealt 22, choose to play it preflop, flop some hand where you'll continue (most likely a set), have the other person in the pot have enough hand where they'll continue far enough to generate a big pot, and then play some sequence where you actually do generate a big pot, you're down to maybe a dozen instances, so a single outlier influences your results a ton.
tl;dr - don't worry about data science. If you play online, use one of the standard HUDs, look at VPIP, PFR, WTSD, and ignore everything else except overall win rate and standard deviation until you're damn sure you know what you're talking about.
You were meaning nail? Or was it intentional?
Does such a thing exist? Some preliminary searches didn't bring much up apart from one-off experiments. [EDIT found a recent article about this https://www.scientificamerican.com/article/time-to-fold-huma... which talks about http://www.computerpokercompetition.org/]
Another interesting one would be a table of 8 players. Half AI half professionals but no one knows who is who.
For a more "real-world" example -- AIs that were actually used in live online poker environments, competing against each other -- there was only one that I know of. It was hosted by the creator of the (apparently now defunct) WinHoldEm software, Ray E. Bornert II. It never ran again, only a small handful of the more experienced bot authors attended, and turnout was poor, with the highest max buy in at $100 ( http://robopoker.blogspot.com/2007/12/pokerbot-world-champio...). It was called the 2007 Poker Bot World Championship (PBWC).
600 hands per hour ... won $1,913.13 over 387,373 hands
That's less than $3/hour, switching to any other career would be an improvement.There are no sharks at those tables.
Also, at 25nl there are many who play for living, as $500 is a decent salary to them
This does not seem like a lot of money to earn for playing 387k hands!
Not a criticism just an observation - but surely this amount of effort could be better spent elsewhere? Purely from a money-making perspective, obviously there is an element of enjoyment here which might change the picture significantly.
Some people semi-derisively refer to these people as "rakeback pros", in that you can break even (or even lose a little bit) in your poker results, and still make a living wage from the rakeback alone.
Also, 387k hands probably takes less time to play than you think. Many online grinders play upwards of 1000 hands/hr - 20+ tables concurrently X 50-150 hands/hr.
Does a call count, or does it have to be an initial bet or raise?
Someone with Vpip: 35 and Pfr: 5 would be an extremely passive weak player.
me: You state on your CV that you are a data scientist
Candidate: Yes
me : What is your Phd in?
Candidate : I don't actually have a Phd
me : <sigh> right, no scientist.
I hate job interviews
As far as I understand, they were not data scientists in the contemporary meaning of the word. At least, they don't have the training, skills and experience I am recruiting for.
It does if you have a Phd in a data science related field.
Of course I know what qualifications they have. I don't know what "fault" you are talking about. Fault for what?
You'd be stupid to market yourself as a data specialist when every startup is looking for a data scientist.
There is for one of the positions I am recruiting for.
Even if there were, why would you exclude a B.S. or an M.S.?
Because I am looking for someone with a Phd in Data Science. How stupid of me, for excluding people that don't have the skills I need for my job.
You'd be stupid to market yourself as a data specialist when every startup is looking for a data scientist.
I'd be stupid for hiring a data specialist, when what I am looking for is a data scientist. I'd be even stupider for hiring someone that is trying to pass themselves off as a data scientist with Phd, when they don't actually have one. I recently interviewed someone with "12 years production hadoop experience" (seriously), claiming to have a "data science Phd" (doesn't exist as such) and when asked "can you please explain the relationship between hbase, hadoop, and hdfs" (yes, there is a reason for asking this question in this way) I had to listen to essentially a salespitch for Hortonworks, and to be told "hdfs has nothing to do with hadoop" <-- this shit, many times a week.
But really, it looks like you need to hire a new recruiter. If you're wasting your time interviewing a Hortonworks entry dev when you want a guy with a PhD, the recruiter is not doing their job.