The tide is shifting: 1.3B outperforms 7B Llama 2
arxiv.org
arxiv.org
Bingo. This (not the parameter count) is the amazing thing to me.
Garbage in garbage out, and there is a ton of garbage in the Falcon/Llama (and OpenAI?) datasets. It feels like such a waste of compute and parameter space.
do these llms get trained with something like a credibility weight on the training data? that was it seems they did in this paper, just manually curated that
Think common use cases. A lot of users are students, do I want it to write an essay like a linguist? Or solve my homework using the better but more advanced techniques and style?
If you have input data that includes a high rated reddit eli5 question, the content of that answer might be hard to verify, the style and way its delivered would be ideal to keep around in the training data. on the other side, technical in-depth answers have content that is worth keeping around, the style of its delivery would be very specific.
Keeping the entire internet around in your training data would still give you access to all these types of delivery still. hope that makes sense.
You do want the underlying model to be capable of the advanced answers, since if it is, it can be used to supply simple answers. You can't make that work the other way around in the same way.
This is true and most researchers understand this on some level (https://arxiv.org/abs/2305.07759) but understanding it doesn't really make the problem any easier. How do you curate a general purpose model's dataset without throwing the baby out with the bathwater ?
EDIT: This paper seems to take a good stab at the question. https://arxiv.org/abs/2309.04564
Time, money, and human work. The pruning process needs focused input from experts other than ML researchers and data scientists.
Maybe we could train an AI to do it ;)
"Perhaps achieving ChatGPT’s level of capability at the one billion parameters scale is actually achievable?"
Right now they're comparing to the 7B Llama 2, which is a shadow of 65B llama2, which is a couple steps off ChatGPT, which is a shadow of GPT4.
I'm still comfy saying yes because there's no reason to doubt it will follow the same logic as "eventually $phone will have the same FLOPS as $desktop"
I don't think finetuning phi-1 on good quality synthetic data will increase its accuracy as it is only trained on that.