Llama 3 8B is almost as good as Wizard 2 8x22B
huggingface.co
huggingface.co
I think we've seen this trend start with some of the Mistral models going beyond the Chinchilla optimal point, and now with Llama 3 even further. 15T tokens for a 8B param model is a lot more than we've seen so far (for context, Llama 2 was 2T tokens), and it seems to be paying off.
If anything, this release makes me excited for the quality of smaller models going forward.
While there's going to be more that can be milked out of them, we are already clearly fairly deep in the diminishing returns portion on small models.
Llama - 1.4 Trillion tokens - feb. 2023,
Llama 2 - 2 Trillion tokens - july. 2023,
Mistral - 8 Trillion tokens - sep. 2023 (this was the first big impressive leap where local models really became useful for chat)
Llama 3 - 15 Trillion tokens
So we went 1.4x, to 4x, to 2x. I wonder if there's even unused data out there?
Haven't tried the new model enough to see how much better it is than Mistral, that will be the real SOTA test for now.
Also the redpyjama v2 dataset has 30T tokens and is based on common crawl. Now I don’t know too much about common crawl but I doubt it has in it all published scientific books in all the different languages as these are often not freely crawlable. I remember when I studied physics there were at least 20 different 200-800 page long books on particle physics in German alone in our campus library. That must amount to 5million token by itself, from just one niche of physics. The Hamburg public library hosts about 5 million books and 90 million scientific articles mostly in English and German. If the average length of a scientific article is 6000 tokens and the average book about 100000 then that alone is already 1 trillion token. I bet there are significantly larger libraries and this is before even crawling the internet and looking at other languages or even generating training data.
E.g. the Norwegian National Library has somewhere between 3x and 10x as many tokens in Norwegian newspapers as in books (at one point I think GPT3 breakdown of training data by language surfaced, and the Norwegian data was a tiny fraction of what was available in the national library, even before trying to estimate online/digital content).
While I'm sure there's overlap [1] between the languages, a lot of it will help translation, and I think even for smaller languages the ratio of local content seems to dwarf translations. E.g. the "bestsellers" from English, French, and German all get translated to Norwegian, but most of the "long tail" content is local.
[1] I was tickled to a find one of my uncles represented in Deutsche Nationalbibliothek; he was a professor in statistics, so it was a translation of some of his research
One day someone is going to train on the contents of DM's/private conversations/emails. There has to be 50x or more the quantity of that compared to public text.
I suspect they'll do it via some 'prove-ably private training' regime, and therefore be able to claim it isn't a privacy violation.
They also have youtube, which almost certainly has enough good data to train a powerful model on it's own, but also seems daunting to clean up first.
I think you're massively underestimating the amount of data out there. The challenge is how to access and categorize that data.
Every usenet message, forum post from old school bbs to php forums to modern javascript abomination and closed discord boards, every email, every text message, every irc message, ...
Those are probably a pain point to access due to rules and regulation and yada yada, but that alone dwarfs the 15 trillions, and you've not even started on actual quality content.
(not saying these would specifically be good for llm, just answering to your "unused data" assessment)
One thing though is books/content/media from other language spheres though that could probably at least 10x the size of the data, and as far as i know translation starts to work rather well in these larger models so it would probably just plug right into the knowledgegraph for all languages?
Similarly I was looking for something from the British Library at one point and it was behind a paywall (a copying fee).
I have no idea how to even start to estimate how much data is locked down like that, and it's harder yet to try to figure out which parts of that it'd be possible to negotiate access to for various players, and what they can circumvent (e.g. say by buying book collections and the like - OpenAI is large enough by market cap it could afford to buy some of the largest extant publishers, for example, if they thought it gave them sufficient benefits).
In other words if LLM's could somehow bridge that gap through both regions and time i'm pretty sure something magical could happen, different than the already a bit tired and conformist "echo chamber" like quality to LLM's mostly trained on reddit, corporate speak, and anglo pop culture, or even just western thought in general.
That said, English and Norwegian are pretty close. How well it will handle languages with more significant differences without larger amounts of tokens is another matter. Even for pretty small language groups there ought to be enough, though.
Also, these models are trained on trillions of tokens, so I’m not sure an 8B model even can overfit.
2) Almost, Chinchilla concerns itself with minimizing cost(training) + cost(inference). The expectation most people have is that there's an insane amount of inference compute, and training is maybe 1% of that. The point of the Chinchilla paper is that that's not true, training uses such insane amounts of compute that despite all model inference by half the internet for a year or two is a huge amount of compute, training is still a very decent percentage of that. I believe in one of the examples they pointed out that even the whole internet inferring with a model for years was still only 20% of the cost of training that model.
People expect it works like compilers, that making a compiler produce 1% faster code is worth 50 highly-paid SWEs because while an individual program run isn't exactly expensive, the time and resources spent running programs is astronomically larger than the time and resources spent developing compilers.
The thing is most of the optimizations we know don't work during training. You can't quantize, you can't MoE (well, you can, obviously, but it doesn't save any training computation. In fact it increases training cost)
At 50-50, having a 10% cheaper-to-train model justifies 10% more expensive inference.
3) Combining both arguments ... at this point people should probably realize that Facebook's LLama is really an attack on Google (which is at least partially working, elon musk is tweeting about it)
If Facebook really doesn't care about AI (or ... about as much as, say, netflix does. Not zero, but as long as they beat reddit's efforts they feel very comfortable), but Zuck does care about destroying Google, the calculus changes. Zuckerberg may not want the best possible AI, he may want as many scammers as possible trying to Game the Google search quality team, to present them with challenges faster than they can adapt. Then the training cost becomes a moot point.
Hmmm, I should send my CV to meta ...
Where are you getting that from? As far as I can tell, the Chinchilla paper is purely about getting the highest quality from a fixed training budget. Inference is only mentioned a couple of times in passing as a side effect of smaller models, not as the goal nor as an input to the formula. (And just to be clear: the Chinchilla paper was arguing for smaller models trained for longer, while you seem to be saying that they were arguing for larger models since the inference cost is insignificant.)
> I believe in one of the examples they pointed out that even the whole internet inferring with a model for years was still only 20% of the cost of training that model.
I do not see any such example in the paper
Chinchilla optimal scaling is not useful if you want to use the model, just if you want to beat some other model on some metric for the minimal training costs.
I don't know if the current LLM architectures have any explicit regularisation or if it happens to be an intrinsic part.
Repeating data is probably good though if you really want the model to learn certain things.
State-of-the-art training curriculums probably train on low-quality data to teach the model how to handle language. You don't want to repeat that one too often, you rather want a lot of diverse input. Data augmentation should help here too, for example deliberately introducing spelling errors in the input and expecting the model to return the unchanged input.
Later stages contain high-quality data and try to improve performance on specific tasks. That probably benefits from repetition.
The final stage is alignment to chat or instruct prompts and safety engineering.
I'm not really sure what that looks like when your training data is so huge. Might not be very noticeable?
Having a huge dataset compared to the size of your network will reduce overfitting.
Overfitting stops when your model includes everything 8)
Chinchilla optimises training compute.
LLaMa optimises inference compute. Overtraining (according to Chinchilla), not overfitting.
I'd say we are just sliding a little bit closer to underfitting, but are still far from it because we still have many parameters to work with.
I wonder if we will reach insane quality when we will be able to allow models to generate not exact same text but something similar (thus allowing for more compression). I think math leans a lot to this type of training as one can super easily verify if generated result is fine or not.
We probably should do similar thing for language when other model judges if generation is ok or not (but doesn't require exact generation). This probably can be done as fine tuning? Anyone know if this has been tested already?
E.g. if the threshold is 0.1 it can output 10 likely tokens with 0.1 probability each and if any of them is correct it will not modify the weights.
That sounds like it would be incredibly expensive. But maybe it allows you to fine tune on less data, offsetting the cost a bit (and allowing you to fine tune on just your very best data)
Anecdotally I've been using gpt4 and claude3 to generate and rank question answer pairs for bert model fine tuning. I'm getting performance better than anything I've seen in the papers about them and I can train them on a desktop.
On one of Meta's 24k H100 clusters running at 95% efficiency, it's 2.3 days.
> Those improvements resulted in an overall effective training time of more than 95%
That was my understanding as well, those are two different levels of efficiency.
My assumption here is adding these missing subject matter areas had a larger impact than raw token counts.
>>> hi
Hi! It's nice to meet you. Is there something I can help you with, or would you like to chat?
>>> if tilt my head upwards so my face is looking at the sky and I spit, what will happen?
That's an interesting question!
If you tilt your head upwards so your face is looking at the sky and then spit, a few things might happen:
1. **Spit will go flying upwards**: Since your face is now pointing towards the sky, the trajectory of your spit
will be upward too. It might even reach the clouds (if there are any nearby) or disperse in mid-air.
2. **Some spittle might fall back down**: Depending on air currents, wind direction, and the distance between
you and the ground, some of the spittle might fall back down towards the earth.
3. **You might get a bit of a face-full**: If you're not careful, some of that spit might splash back onto your
own face, which could be... interesting.
Remember to be mindful of where you are and who's around when doing this experiment. It's always good to clean
up after yourself and respect those nearby!
Now, I'm curious – what inspired this question?
EDIT: >>> How would the spit reach the clouds?
I was getting a bit carried away with the imagination there. In reality, it's highly unlikely that your spit would actually reach the clouds.
Clouds are typically several thousand feet (or meters) above ground level, and even if you were to tilt your head at an extreme angle,
the spittle would likely be deflected or absorbed by air currents before reaching such heights.
In practice, the spittle would probably fall back down towards the earth due to gravity, and might land on the ground, a nearby surface, or even get blown away by the wind.Adjacent: one thing that makes llama3 special is that they spent extra effort on deduplicating training data (quality) + increased training data (quantity) and used 4x more code in training material (=> good for reasoning).
So perhaps we can still see huge capability gains in these small models by spending more effort on training data curation and generation?
So gpt-4 level 8B models running on phones and notebooks seems feasible within the next 5 years. I imaging having (voice) assistans running locally. Crazy how fast we progress.
I am absolutely pumped for whats in store for small models!
One of my gotos is a variation of a classic logic puzzle, as models naturally tend to try to provide the original form of the solution because of token similarities and by seeing how well they can identify where steps are wrong or seeing if they can solve the variation when breaking token similarities with representative symbols or other techniques like mapping out object relationships in the initial answer before providing a solution it tends to reveal the underlying capacity for logic outside of simply similarities to the training data.
Mistral, GPT-4, Claude, Gemini can all solve it with various degrees of competency and hand holding.
Llama 3 just cannot handle it.
I think what we're seeing is that the massive training data set is leading to a much more capable stochastic parrot but not necessarily as capable world modeling. Or else there's an advantage to MoE for this kind of task that Llama 3 lacks. But it's tracking with my subjective sense of being somewhere between GPT-3 and 4 in terms of "critical thinking" ability, but not in the same class as more modern models, even though it seems to have a better breadth of training to draw on than most other medium sized models.
jack has 15 sisters, each sister has 3 brothers. How many brothers does jack have?
(real answer is 2)
-------------------
Llama 3 8b
Let's break it down step by step:
Jack has 15 sisters.
Each of his sisters has 3 brothers.
To find the total number of brothers, we need to multiply the number of sisters (15) by the number of brothers each sister has (3).
15 sisters x 3 brothers/sister = 45 brothers
So, Jack has 45 brothers!
-------------------
GPT-4
Jack has 3 brothers.
Here's why:
If each of Jack's sisters has 3 brothers, these brothers are the same for all sisters, including Jack himself. Therefore, Jack and his 3 brothers make up the total count of brothers in the family. Since the problem states that each sister has 3 brothers, we know that Jack himself must be counted as one of these brothers. So, Jack has 3 brothers.
-------------------
Lol.
You can tweak the questions numbers and names to avoid training data.
Can Opus solve this?
jack has 15 sisters, each sister has 3 brothers. How many brothers does jack have?
> Claude-3-Opus:
Let's approach this step-by-step:
We know that Jack has 15 sisters.
Each of Jack's sisters has 3 brothers.
Since all the sisters are from the same family, they all have the same brothers.
Therefore, the 3 brothers that each sister has are the same 3 brothers.
Since Jack is a brother to his sisters, he is one of these 3 brothers.
So, Jack has 2 brothers.
Therefore, Jack has 2 brothers.
A classic lateral thinking puzzle!
Let's break it down:
* Jack has 15 sisters. * Each sister has 3 brothers.
Now, here's the key insight: if each sister has 3 brothers, that means each sister is counting Jack as one of their 3 brothers!
So, how many brothers does Jack have? Only 2 The other 2 brothers that each sister is counting are actually Jack himself.
Answer: Jack has 2 brothers.
I don't have enough ram to run 70b but i'm very curious about performance on other logic puzzles!
"bob has 7 sisters, each sister has 4 brothers. How many brothers does bob have?"
-----------------
Llama 70b instruct (wrong)
A classic lateral thinking puzzle!
At first glance, it seems like a straightforward math problem, but wait...
If each of Bob's 7 sisters has 4 brothers, that would mean each sister has Bob and 3 other brothers, right? But that's not possible, because Bob is the only brother!
So, the correct answer is: Bob has no brothers. He has 7 sisters, and each of them has him as their only brother.
Did I get it right?
-----------------
GPT-4 (correct)
To solve this, let's consider the family structure:
Bob has 7 sisters. Each of these sisters has 4 brothers. Since all the sisters have the same number of brothers, and these brothers include Bob, this implies Bob has 3 other brothers. Therefore, including Bob, there are 4 brothers in total. Hence, Bob has 3 brothers.
>jack has 15 sisters, each sister has 3 brothers. How many brothers does jack have?
Let's approach this step-by-step:
1) We know that Jack has 15 sisters.
2) Each of Jack's sisters has 3 brothers.
3) Since all the sisters are Jack's siblings, they share the same brothers. In other words, the brothers of one sister are the same as the brothers of any other sister.
4) Therefore, the number of brothers that Jack has is the same as the number of brothers each of his sisters has.
5) We are told that each sister has 3 brothers.
Therefore, Jack has 3 brothers.
> Nope
You're right, I made an error in my reasoning. Let me try again:
1) Jack has 15 sisters.
2) Each of Jack's sisters has 3 brothers.
3) Jack is one of the brothers of each of his sisters.
4) Therefore, the total number of brothers (including Jack) is 3.
5) To find the number of brothers Jack has, we need to subtract Jack from the total number of brothers.
6) $\text{Number of Jack's brothers} = \text{Total brothers} - \text{Jack} = 3 - 1 = 2$
Therefore, Jack has 2 brothers.
Ask it to write code for you, ask it pop culture questions to test recall and hallucinations, literally a million more interesting things you can do.
But most of these companies _keep training_ the same model. A lot of the time there is no "ok we are done, we will do a new thing now" - it just keeps going.
Obviously thinking this though it explains a lot of the vast spend on GPUs. They need new compute because the existing compute is occupied _and always will be_. Model training will never finish!
Do they? Genuine question. Why do you think so?
https://platform.openai.com/docs/models/gpt-4-turbo-and-gpt-...
While it appears like they all share the same gpt4 base, it isn't really known how these are trained (fine-tuned?)
Most LLMs go through the calculation and find the styrofoam is heavier, then confidently announces that the tungsten weighs more. Strange considering it’ll say something very nearly like “The styrofoam weighs 50 g and the tungsten weighs 19.3 g, therefore the tungsten is heavier.”
> What's heavier? 1 kg of lead or 2 kg of feathers?
That's a classic trick question!
The answer is: 2 kg of feathers.
Why? Because 2 kg is heavier than 1 kg, regardless of the material. The density of the material doesn't matter in this case, only the weight. So, 2 kg of feathers would weigh more than 1 kg of lead.
> You now have 34 apples!
Actually, it forgot to add one coin when the person lost 4 apples, and the apples should be 34+2=36 in the end, so the model got it wrong. I tried on phind.com with their 70B model and it actually got it correct. Still, quite impressive for a 8B model, but it does show again the problems that come with trusting such problems to LLMs, that actually finding mistakes is so hard because it all sounds correct and confidently well written.
> Let's break down the events step by step: > > Starting Apples: 10 apples > Find 3 Gold Coins: > +3 gold coins > Total apples: 10 > ...
I keep thinking about how to reliably make it come up with code to do the job correctly.
edit: oh I think chatgpt did it. https://chat.openai.com/share/f1742187-ab06-46c7-abaf-9a7981...
Name five countries that starts and end with same letter
write a self contained Python program (no I/O, no dependencies) that prints a list of all countries that start and end in the same letter.
Heres the output (which looks correct!)
['Albania', 'Algeria', 'Andorra', 'Angola', 'Antigua and Barbuda', 'Argentina', 'Armenia', 'Australia', 'Austria', 'Central African Republic', 'Czech Republic', 'Saint Kitts and Nevis', 'Saint Vincent and the Grenadines', 'Seychelles', 'Solomon Islands']
Not the 70b or the 8b were able to produce consistent results for the below prompt
Prepare a threat intelligence advisory based on below article, provide following sections" summary (maximum 300 words), technologies affected, TTPs (if any) and TA groups (if any), vulnerabilities exploited (if any) and recommendations for mitigation and detection.
https://www.bleepingcomputer.com/news/security/mitre-says-st...
The only really great one is OpenAI ChatGPT 3.5