Thanks! Do you know anything more applied, like fine-tuning of local models with RL? Reward function design, evals etc
That's a good starting point. I'm interested in hearing more about real-world applications and challenges in RL for LLMs. Any insights or experiences on that?