Value-Based Deep RL Scales Predictably
arxiv.org
arxiv.org
For a different perspective, error vs compute, see
and comments
(I particularly liked the one about string theorists rediscovering a fundamental theorem in GR decades too late-- rediscovering how to integrate happens in every field, it's nothing to be ashamed of :)
It's exciting to see any progress in making accurate predictions about what settings will work for RL training. I hope that this research direction can be expanded in scope and that ultimately, people who want to do research in RL can become confident in their training recipes.
I am more excited about that, than about the dream of scaling compute per se.