Do you really need computer vision to detect when someone is in a game? There are certain static elements on the screen in every game that you could probably just take a current screenshot and match against known images?
One problem is that the elements themselves don't necessarily scale proportionally. For example, if you have a button with some text in it and you reduce the resolution, often the game will keep the text larger but reduce the padding around the text. This makes template matching tough.
We found blob detection works pretty reliably and uses minimally CPU, but we're not CV experts so if anyone has other ideas we'd love to hear them!
The blob detection does work surprisingly well though, doesn't take us very long to add a new game using it. Most of the time is spent making sure we understand the flow of games we don't play ourselves.