HNHacker News
TopNewBestAskShowJobs

anakin87

-1 karma · joined April 4, 2023

submissionscomments
anakin87··on Show HN: Hands-on course for building RL environments for LLMs
Hi HN, I've been spending some time lately trying to build Reinforcement Learning Environments and training small language models and wanted to share a little course I put together based on my experiments.

Over the past year, we've seen a shift in LLM Post-Training. Previously, Supervised Fine-Tuning was the most important part: making models imitate curated Question-Answer pairs. Now with RLVR and GRPO, we can make models learn through trial and error in dynamic environments, which are software artifacts.

But how to effectively build RL environments?

In the repo, I cover:

- Mapping core RL concepts (Agents, Environments) to the LLM domain.

- Using the Verifiers open-source library to construct single-turn, multi-turn, and tool-use environments.

- Hands-on: taking a small language model (LiquidAI's LFM2-2.6B) and turning it into a Tic-Tac-Toe master that beats GPT-5-mini. Build the game Environment, ese it to generate synthetic data for SFT warm-up, then Group-based Reinforcement Learning.

---

Links

Course: https://github.com/anakin87/llm-rl-environments-lil-course

Video walkthrough: https://www.youtube.com/watch?v=71V3fTaUp2Q

Play against the trained model: https://huggingface.co/spaces/anakin87/LFM2-2.6B-mr-tictacto...

Datasets and Models on HF: https://huggingface.co/collections/anakin87/lfm2-26b-mr-tic-...

---

I'm fascinated by the idea of building these "little worlds" where LLMs can learn, so I hope it's useful.

Feel free to share opinions...

anakin87··on Environments Hub: Your Language Model needs better (open) environments to learn
LLMs improve when they can practice and reason in interactive environments.

Recent work (DeepSeek-R1, GRPO) shows RL can teach models to prefer better outputs by giving rewards.

But most RL environments for LLMs are fragmented or closed. That makes it hard for the community to experiment or reproduce results.

Environments Hub is a new open platform by Prime Intellect where anyone can share RL environments for training or evaluating LLMs. Think of them as software packages: data, harness and scoring rules.

Agents today incorporate models and tools (from APIs to a terminal), so environments need to capture that complexity.

I wrote a hands-on walkthrough covering:

  - RL + LLM basics

  - Navigating the Environments Hub

  - Evaluating models and agents

  - GRPO-style training of a tiny model on an alphabetical sort task
If you want to experiment with RL for LLMs or just see how open environments can accelerate learning, this walkthrough is a practical starting point.
anakin87··on GRPO experiment: I trained a Language Model to schedule events
I experimented with GRPO lately, since I am fascinated by models learning from prompts and rewards - no example answers needed like in Supervised Fine-Tuning.

After the DeepSeek boom, everyone is trying GRPO with GSM8K or the Countdown Game, but I wanted a different challenge.

So I opted for teaching a model to create a schedule from a list of events and priorities.

Choosing an original problem forced me to think about the problem setting, generate data, choose the base model, design reward functions, and run multiple rounds of training, hoping that my model would learn something.

A fun and rewarding experience :-)

I learned a lot of things, that I want to share with you.

---

- Blog post: https://huggingface.co/blog/anakin87/qwen-scheduler-grpo

- Code: https://github.com/anakin87/qwen-scheduler-grpo

- Hugging Face collection (dataset and model): https://huggingface.co/collections/anakin87/qwen-scheduler-g...

---

Some hot takes from my experiment

- GRPO is cool for verifiable tasks, but is more about eliciting desired behaviors from the trained model than teaching completely new stuff to it. https://arxiv.org/abs/2504.13837

- Choosing the right base model (and size) matters.

- "Aha moment" might be over-hyped. https://oatllm.notion.site/oat-zero

- Reward functions design is crucial. If your rewards are not robust, you might experience reward hacking (as it happened to me).

- Unsloth is great for saving GPU, but beware of bugs.

anakin87··on I trained a Language Model to schedule events with GRPO
I experimented with GRPO lately, since I am fascinated by models learning from prompts and rewards - no example answers needed like in Supervised Fine-Tuning.

After the DeepSeek boom, everyone is trying GRPO with GSM8K or the Countdown Game, but I wanted a different challenge. So I opted for teaching a model to create a schedule from a list of events and priorities.

Choosing an original problem forced me to think about the problem setting, generate data, choose the base model, design reward functions, and run multiple rounds of training, hoping that my model would learn something.

A fun and rewarding experience. :-)

I learned a lot of things, that I want to share with you.

Blog post: https://huggingface.co/blog/anakin87/qwen-scheduler-grpo

Code: https://github.com/anakin87/qwen-scheduler-grpo

Hugging Face collection (dataset and model): https://huggingface.co/collections/anakin87/qwen-scheduler-g...

anakin87··on Llama2 + Haystack on Colab
I recently conducted some experiments with Llama2 and Haystack (https://github.com/deepset-ai/haystack), the NLP/LLM framework.

The notebook can be helpful for those trying to load Llama2 on Colab.

1) Installed Transformers from the main branch (and other libraries)

2) Loaded Llama-2-13b-chat-hf on Colab using 4-bit quantizazion, thanks to the material shared by Younes Belkada

3) Disabled Tensor Parallelism, which caused some issues

4) Installed a minimal version of Haystack

5) Found a hacky way to load the model in Haystack's PromptNode

6) Had a fun chat session with the model, discussing everything from David Guetta to Don Matteo (an Italian TV series)!

anakin87··on Open-Source Data Collection Platform for LLM Fine-Tuning and RLHF
In my experience, Argilla is a good open source platform for datacentric NLP. And these features are a great addition... Have you tried it?