ParentFull threadeamag·Thanks! Do you know anything more applied, like fine-tuning of local models with RL? Reward function design, evals etcView on HN