Edit: he does make a good point later on that playing thousands of games is exhausting, which would affect the reliability of the comparison between the AI and the pros. I don't know anything about poker, but it sounds plausible. Presumably this means they should make the AI so good that it can demonstrate its greatness quicker than a human gets tired.
Unfortunately, some of the examples which he claims don't pass this significance threshold actually do. This does not help his credibility.
I don't know if Doug Polk understands that or not, but I agree with his criticism and I think his analogy with sports reporting is sound. While "statistical tie" is true in a technical sense, it's not usually how matches are reported, and it's somewhat disingenuous and self-serving to use that language. It would be more honest to say that it's very unlikely that bot is better than humans -- given the advantages the bot had (a gruelling 2-week schedule, and pressure on the humans to get through N hands per day) it still lost by a significant margin.
> It would be more honest to say that it's very unlikely that bot is better than humans
I don't think the analogy is sound. Consider the alternative: over a long match, the humans failed to beat the bot convincingly. But in any case, the examples he gives are so wide of the mark that it's clear this isn't the sort of argument he's making—he talks about "spin" and stuff.
Does that mean no-limit poker is volatile enough that results considered by insiders to be long-term don't achieve 95% confidence? Idk probably, I haven't looked at the math myself. Is 73 buy-ins over 80k hands a huge win? No doubt.