Finetuning a Reasoning LLM with Supervised or Reinforcement Learning?discuss.huggingface.co2 points·verdverm··0 commentsOpen articleSaveView on HN