Recent reasoning research: GRPO tweaks, base model RL and data curationinterconnects.ai1 point·ydnyshhh··0 commentsOpen articleSaveView on HN