https://github.com/google-research/circuit_training
It pretty much gets the same results as found in the Nature paper.
The original codebase was heavily research-focused, used TF1, was impossible to run distributed training outside of Google's infra, and made it hard to try algorithms other than PPO. So it was reimplemented on top of TF2 and using some distributed training and collection technologies developed by the TF-Agents team at Google Brain and infra teams at DeepMind.
Everyone is welcome to poke at the training code and the model, and convince themselves that it does what it says on the box :)