Seems like persistent models like OpenAI's highly persistent internal model can become really effective over time. Those are the ones that drove most of the HF-OAI incident.
This training technique does not relate to how persistent a model is, at all really. They sample more parallel attempts at hard problems, to increase their chances of having at least one success to learn from.