I don’t know how much you find the use of numerical things or a tailored system an issue vs “here’s balatro let’s go” but this might be of interest
I'm a bit skeptical of their "2 seeds in a row!" boast. Last time I investigated a claim like that I found the seeds were cherry picked. This was way back in the OpenAI Gym days though (remember when OpenAI was open and just doing goofy research like OpenAI Gym?), their leader boards had some amazing claims about certain RL solutions, but when I ran them myself on new seeds they were far worse than claimed.