On the grid graph comparison of how different algorithms perform on different games, two questions:
1) What is the source data for that plot?
2) You specify "lighter = better", but how are they normalized across games and algorithms? How is better and worse quantified to get a "lightness"?
Edit: Found #2 in the second paper. Still don't know what 25 wins is white and 0 wins is black means? How do you "win" some of these?
Two papers are here: