What are you training on using self play? Like alpha go? Curious what your setup is like .
I had to start with some heuristic-based bots that played the decks very simply just to get to the point where the was some signal to learn from. I did behavioral cloning on the bots as a foundation, then self-play.