We have experiments showing that agents at least can learn from the environment by overfitting on train. But we do not yet have full post train runs, mainly due to time. But follow our blog/X where we’ll regularly update our research
No comments yet.