Show HN: Llama-8B Teaches Itself Baby Steps to Deep Research Using RL
github.com
github.com
Questions:
-- Does a training over a body of data export into better performance over subsequent bodies of data - as you should also be training meta-skills?
-- Your benchmark revealed a growth from 23% to 53% after an hour: and after further training? If it plateaus, why does it?
I haven't tried to switch the dataset, but I am fairly certain the LLM is training meta-skills. It seems that the majority of what the model learns is to behave in a more reasonable way, and to stop hallucinating + improperly using tools. Not to memorize the data in the body of knowledge.
During the first hour of training, llama learns most of the low hanging fruit (stop messing up function calls and stop hallucinating). So after that, learning slows down.