Understanding Llama 2 and the New Code Llama LLMs
magazine.sebastianraschka.com
magazine.sebastianraschka.com
They state:
> The Code Llama models were trained on 500B tokens, whereas Llama 2 models were trained on 2T tokens. Since the Code Llama model was trained on 4x fewer tokens, maybe a CodeLlama 70B version did not perform well enough due to LLM scaling laws—there was not enough training data.
But if you read the paper, on page 1, it says:
> Our approach is based on gradually specializing and increasing the capabilities of Llama 2 models by applying a cascade of training and fine-tuning steps [...]
In fact, they show a diagram at the top of page 3 that details the process, starting with Llama 2 foundation models.
Llama 2 Foundation models (7B, 13B, 34B) -> Code training 500B -> Python / Long Context.
See the paper here: https://arxiv.org/abs/2308.12950
### off topic rants below
Somehow there are so many blogpost about these things, all trying to ask for your emails. Is it becoming easier to put more words together nowadays? I guess so.
I really wish there is a way to fact check all, instead of depending on good samaritans in a comment on HN to point these obvious misconceptions out.
You mean like reading original sources? Frequently, big research projects like this come with an official paper[1] and/or blog post[2] explaining what they did.
[1] https://ai.meta.com/research/publications/code-llama-open-fo...
[2] https://ai.meta.com/blog/code-llama-large-language-model-cod...
That's because Substack defaults to bothering people for their email, and lots of people are using Substack as their blogging platform these days.
they shouldn't. It's Medium all over again...
What I meant to say here was 500B domain-specific tokens. Maybe domain-specific is not the right word here, but tokens related to the problems that the LLM aims to solve.
EDIT: Updated the text to be more clear.
We know that for the original version of GPT-3.5, but my assumption was that Turbo was a distilled smaller model (which is why it uses OAI's new vocab & is so much faster).
If that's not the case, what could be the explanation for it being faster?
I'd say that it's probably a mix of all of the above (incl some distillation).
* For a lot of prompts every fine tuned model will make the same mistakes (they mostly share the same weights after all) and so you aren't getting nearly as much benefit as e.g. GPT-4 gets.
* It's going to be really expensive at inference time, since you have to run multiple models even though in most cases they won't help much
* Normally when people talk about hobbyists doing finetuning they mean <1M tokens, whereas Code Llama Python was finetuned on 100B tokens, way outside most people's price range. For the finetuning that you can afford, you can't teach it new knowledge, just show it how to apply the knowledge it already has.
Any LoRA approach is obviously going to be perform a little worse that a fully tuned model, but I guess the jury is still out on whether this approach will actually work well.
Exciting times!
Anyone else having a similar experience?
llama.cpp >> OpenAI translation server (included in llama git) >> Continue extension in vscode