Llama 3.1
llama.meta.com
llama.meta.com
Open source AI is the path forward - https://news.ycombinator.com/item?id=41046773 - July 2024 (278 comments)
Quick comparison with GPT-4o:
+----------------+-------+-------+
| Metric | GPT-4o| Llama |
| | | 3.1 |
| | | 405B |
+----------------+-------+-------+
| MMLU | 88.7 | 88.6 |
| GPQA | 53.6 | 51.1 |
| MATH | 76.6 | 73.8 |
| HumanEval | 90.2 | 89.0 |
| MGSM | 90.5 | 91.6 |
+----------------+-------+-------+Where could I get a mapping of token / time vs hardware?
The $10k figure is likely roughly the minimum amount of money/hardware that you'd need to run the model at acceptable speeds, as anything less requires you to compromise heavily on GPU cores (e.g. Tesla P40s also have 24GB of VRAM, for half the price or less, but are much slower than 3090s), or run on the CPU entirely, which I don't think will be viable for this model even with gobs of RAM and CPU cores, just due to its sheer size.
I would be curious to see relative failure rates over time of consumer vs Quadro cards as well.
I don't have any actual figures to back this up, but my gut tells me that the fact that enterprise GPUs are an order of magnitude (at least) more expensive than, say a, 3090, means that the payback period of them has got to be pretty long. I also wonder whether setting the max power on a 3090 to a lower than default value (as I suggest in my other post) has a significant effect on the average W/token.
Not necessarily saying that Quadros are cheaper, just that there's more to the calculation when trying to run 405B size models at home
I think it should work as-is with the components listed, but if you disagree please let me know!
I'm not necessarily saying that it's obviously better in terms of total cost, just that there are more factors to consider in a system of this size.
If inference is the only thing that is important to someone building this system, then used 3090s in x8 or even x4 bifurcation is probably the way to go. Things become more complicated if you want to add the ability to train/do other ML stuff, as you will really want to try to hit PCIE 4.0 x16 on every single card.
Will need more space, true.
Here are some TGI 405B benchmarks that I did with the different quantized models:
https://x.com/danieldekok/status/1815814357298577718
The 405B model is very useful outside direct use in inference though. E.g. for generating synthetic data for training smaller model:
I agree with you on open source in the original, home tinkerer sense.
I'd think of the 405B model as the equivalent to a big rig tractor trailer. It's not for home use. But also check out the benchmark improvements for the 70B and 8B models.
PCIE 5.0 x16 is 500 Gbit/s if I'm not mistaken, so using RAM is more viable an alternative in this case.
Edit: 3090 has 1 TB/s, not terabits
This is good enough for a lot of use cases... on a laptop. An expensive laptop, but hardware only gets better and cheaper over time.
(It also predicted that a MacBook Air in 2030 will be able to do the same, and that for smartphones to do the same might take around 20 years.)
A 405B LLM has 405 billion parameters. If you run it at full "prescision", each parameter takes up 2 bytes, which means you need 810GB of memory. If it does not fit in RAM or GPU memory it will swap to disc and be unusably slow.
You can run the model at reduced prescision to save memory, called quantisation, but this will degrade the quality of the response. The exact amount of degradation depends on the task, the specific model and its size. Larger models seem to suffer slightly less. 1 byte per parameter is pretty much as good as full precision. 4 bits per parameter is still good quality, 3 bits is noticeably worse and 2 bits is often bad to unusable.
With 128GB of RAM, zero overhead and a 405B model, you would have to quantize to about 2.5 bits, which would noticeably degrade the response quality.
There is also model pruning, which removes parameters completely, but this is much more experimental than quantisation, also degrades response quality, and I have not seen it used that widely.
The comment about the MPS PyTorch backend was related to performance, not whether the model would fit at all. I can't say whether it's accurate that the MPS backend has significant room for optimization, but it is still publicly listed as in beta.
I would be sceptical about increasing efficiency. I'm not that familiar with the subject, but as far as I know, LLMs for single users (i.e. with batch size 1) are practically always limited by the memory bandwidth. The whole LLM (if it is monolytic) has to be completely loaded from memory once for each new token (which is about 4 characters). With 400GB per second memory bandwidth and 4-bit quantisation, you are limited to 2 tokens per second, no matter how efficiently the software works. This is not unusable, but still quite slow compared to online services.
Maybe we'll get really good LLMs on local hardware when the hype has died down a bit, memory is cheaper and the models are more efficient.
For example here's the list of backends for Llama.cpp: https://github.com/ggerganov/llama.cpp?tab=readme-ov-file#su...
Ever priced out a four wheeler, a jet-ski, a filled gun safe, what a "car guy" loses in trade in values every two years, what a hobbyist day-trader is losing before they cut their losses or turn it around, or what a parent who lives vicariously through their child and drags them all over their nearby states for overnight trips so they can do football/soccer/ballet/whatever at 6am on Saturdays against all the other kids who also won't become pro athletes? What about the cost of a wingsuit or getting your pilots license? "Cruisers" or annual-Disney vacationers? If you bought a used CNC machine from a machine shop? But spend five grand on a laptop to play with LLMs and everyone gets real judgmental.
Not in the slightest. They even have a table of cloud providers where you can host the 405B model and the associated cost to do so on their website: https://llama.meta.com/ (Scroll down)
"Open Source" doesn't mean "You can run this on consumer hardware". It just means that it's open source. They also released 8B and 70B models for people to use on consumer gear.
I suppose language changes. I just prefer it changes towards being more precise, not less.
In terms of functional role, if we're to compare the models to open-sourced games, then all that's been open-sourced is the trivial[0] bit of code that does the inference.
Maybe a more adequate comparison would be a SoC running a Linux kernel with a big NVidia or Qualcomm binary blob in the middle of it? Sure, the Linux kernel is open source, but we wouldn't call the SoC "open source", because all that makes it what it is (software-side) is hidden in a proprietary binary.
--
[0] - In the sense that there's not much of it, and it's possible to reproduce from papers.
Not sure if the code is required under an open source license, but it's the same issue.
---
IMO, source is source and can be used for other datasets. Dataset isn't available, bring your own.
In this case, the source is there. The output is there, and not technically required. What isn't available is the ability to confirm the output comes from that source. That's not required under open source though.
What's disingenuous is the output being called 'open source'.
If some "opensource" model just have the model and training methods but no dataset, it’s like some repo which released an executable file with a detailed design doc. Where is the source code? Do it yourself, please.
NOTE: I understand the difficulty of open-sourcing datasets. I'm just saying that the term "opensource" is getting diluted.
The weights are, for all practical purposes, source code in their own right. The GPL defines "source code" as "the preferred form of the work for making modifications to it". Almost no one would be capable of reproducing them even if given the source + data. At the same time, the weights are exactly what you need for the one type of modification that's within reach of most people: fine-tuning. That they didn't release the surrounding code that produced this "source" isn't that much different than a company releasing a library but not their whole software stack.
I'd argue that "source" vs "weights" is a dangerous distraction from the far more insidious word in "open source" when used to refer to the Llama license: "open".
The Llama 3.1 license [0] specifically forbids its use by very large organizations, by militaries, and by nuclear industries. It also contains a long list of forbidden use cases. This specific list sounds very reasonable to me on its face, but having a list of specific groups of people or fields of endeavor who are banned from participating runs counter to the spirit of open source and opens up the possibility that new "open" licenses come out with different lists of forbidden uses that sound less reasonable.
To be clear, I'm totally fine with them having those terms in their license, but I'm uncomfortable with setting the precedent of embracing the word "open" for it.
Llama is "nearly-open source". That's good enough for me to be able to use it for what I want, but the word "open" is the one that should be called out. "Source" is fine.
[0] https://github.com/meta-llama/llama-models/blob/main/models/...
Fine-tuning and LoRAs and toying with the runtime are all directly equivalent to DLL injection[0], trainers[1], and various other techniques used to tweak a compiled binary before or at runtime, including plain taking at the executable with a hex editor. Just because that's all anyone except the model vendor is able to do, doesn't merit calling the models "open source", much like no one would call binary-only software "open source" just because reverse engineering is a thing.
No, the weights are just artifacts. The source is the dataset and the training code (and possibly the training parameters). This isn't fundamentally different from running an advanced solver for a year, to find a way to make your program 100 byes smaller so it can fit on a Tamagochi. The resulting binary is magic, can't be reproduced without spending $$$$ on compute for th solver, but it is not open source. The source code is the bit that (produced the original binary that) went into the optimizer.
Calling these models "open source" is a runaway misuse of the term, and in some cases, a sleigh of hand.
--
[0] - https://en.wikipedia.org/wiki/DLL_injection
[1] - https://en.wikipedia.org/wiki/Trainer_(games) - a type of programs popular some 20 years ago, used to cheat at, or mod, single-player games, by keeping track of and directly modifying the memory of the game process. Could be as simple as continuously resetting the ammo counter, or as complex as injecting assembly to add new UI elements.
No, because fine tuning is basically just a continuation of the same process that the original creators used to produce the weights in the first place, in the same way that modifying source code directly is in traditional open source. You pick up where they left off with new data and train it a little bit (or a lot!) more to adapt it to your use case.
The weights themselves are the computer program. There exists no corresponding source code. The code you're asking for corresponds not to the source code of a traditional program but to the programmers themselves and the processes used to write the code. Demanding the source code and data that produced the weights is equivalent to demanding a detailed engineering log documenting the process of building the library before you'll accept it as open source.
Just because you can't read it doesn't make it not source code. Once you have the weights, you are perfectly capable of modifying them following essentially the same processes the original authors did, which are well known and well documented in plenty of places with or without the actual source code that implements that process.
> Calling these models "open source" is a runaway misuse of the term, and in some cases, a sleigh of hand.
I agree wholeheartedly, but not because of "source". The sleight of hand is getting people to focus on that instead of the really problematic word.
No, because most video games aren't licensed in a way that makes that explicitly authorized, nor is modding the preferred form of the work for making modifications. The video game has source code that would be more useful, the model does not have source code that would be more useful than the weights.
When you require the same thing in software, namely the whole stack to run the software in question to be open source, we don't call the license open source.
Hell, in case of the models, "the whole stack to run the software" already is open source. Literally everything except the actual sources - the datasets and the build scripts (code doing the training) - is available openly. This is almost a literal inverse of "open source", thus shouldn't be called "open source".
This time, I just copy pasted the raw metrics I found and asked an LLM to format it as an ASCII table.
GPT-4o 30.7
GPT-4 turbo (2024-04-09) 29.7
Llama 3.1 405B Instruct 29.5
Claude 3.5 Sonnet 27.9
Claude 3 Opus 27.3
Llama 3.1 70B Instruct 26.4
Gemini Pro 1.5 0514 22.3
Gemma 2 27B Instruct 21.2
Mistral Large 17.7
Gemma 2 9B Instruct 16.3
Qwen 2 Instruct 72B 15.6
Gemini 1.5 Flash 15.3
GPT-4o mini 14.3
Llama 3.1 8B Instruct 14.0
DeepSeek-V2 Chat 236B (0628) 13.4
Nemotron-4 340B 12.7
Mixtral-8x22B Instruct 12.2
Yi Large 12.1
Command R Plus 11.1
Mistral Small 9.3
Reka Core-20240501 9.1
GLM-4 9.0
Qwen 1.5 Chat 32B 8.7
Phi-3 Small 8k 8.4
DBRX 8.0
If you want to learn more, there is a writeup at https://wow.groq.com/now-available-on-groq-the-largest-and-m....
(disclaimer, I am a Groq employee)
Free trial gets you 50 messages, no credit card required - https://double.bot
(disclaimer, I am the co-founder)
I gave a seminar about the overall approach recently, abstract: https://shorturl.at/E7TcA, recording: https://shorturl.at/zBcoL.
This two-part AMA has a lot more detail if you're already familiar with what we do:
I'm really impressed by what (&how) they're doing and would like to pay for a higher rate limit, or failing that at least know if "soon" means "weeks" or "months" or "eventually".
I remember TravisCI did something similar back in the day, and then Circle and GitHub ate their lunch.
Statement from Mark: https://about.fb.com/news/2024/07/open-source-ai-is-the-path...
Where the right hardware is 10x4090s even at 4 bits quantization. I'm hoping we'll see these models get smaller, but the GPT-4-competitive one isn't really accessible for home use yet.
Still amazing that it's available at all, of course!
https://about.fb.com/news/2024/07/open-source-ai-is-the-path...
[1]: https://opensource.org/blog/metas-llama-2-license-is-not-ope...
As I have stated time and again, it is perfectly fine for them to slap on whatever license they see fit as it is their work. But it would be nice if they used appropriate terms so as not to disrupt the discourse further than they have already done. I have written several walls of text why I as a researcher find Facebook's behaviour problematic so I will fall back on an old link [2] this time rather than writing it all over again.
Is it? Has there been a ruling on the enforceability of the license they attach to their models yet? Just because you say what you release can only be used for certain things doesn't actually mean what you say means anything.
It's "a Google and Apple can't use this model in production" clause that frankly we can all be relatively okay with.
You're only ok with it if you're not interested in having maximum freedom of movement vis-a-vis any potential exits.
It should also be noted (again) that the value of the terms open science and open source comes from the sacrifices and efforts of numerous academic, commercial, personal, etc. actors over several decades. They "paid" by sticking to the principles of these movements and Facebook is now cashing in on their efforts; solely for their own benefit. Not even Microsoft back in 2001 in the age of "fear uncertainty and doubt" were so dishonest as to label the source-available portions of their Shared Source Initiative as something it was not. Facebook has been called out again and again since the release of LLaMA 1 (which in its paper appropriated the term "open") and have shown no willingness to reconsider their open science and open source misuse. At this point, I can no longer give them the benefit of the doubt. The best defence I have heard is that they seek to "define open in the 'age of AI'", but if that was the case, where is their consensus building efforts akin to what we have seen numerous academics and OSI carry out? No, sadly the only logical conclusion is that it is cynical marketing on their part, both from their academics and business people.
[1]: https://en.wikipedia.org/wiki/Shared_Source_Initiative
In short. I think the correct response to Facebook is: "Thank you for the weights, we appreciate it. However, please stop calling your actions and releases something they clearly are not."
If you want a playground to test this model locally or want to quickly build some applications with it, you can try LLMStack (https://github.com/trypromptly/LLMStack). I wrote last week about how to configure and use Ollama with LLMStack at https://docs.trypromptly.com/guides/using-llama3-with-ollama.
Disclaimer: I'm the maintainer of LLMStack
In theory the benchmarks should be a pretty close proxy for quality, but that doesn't match my experience at all.
Examples: OpenAI's GPT 4o-mini is second only to 4o on LMSys Overall, but is 6.7 points behind 4o on MMLU. It's "punching above its weight" in real-world contexts. The Gemma series (9B and 27B) are similar, both beating the mean in terms of ELO per MMLU point. Microsoft's Phi series are all below the mean, meaning they have strong MMLU scores but aren't preferred in real-world contexts.
Llama 3 8B previously did substantially better than the mean on LMSys Overall, so hopefully Llama 3.1 8B will be even better! The 70B variant was interestingly right on the mean. Hopefully the 430B variant won't fall below!
For my use of the chat interface, I don't think lmsys is very useful. lmsys mainly evaluates relatively simple, low token count questions. Most (if not all) are single prompts, not conversations. The small models do well in this context. If that is what you are looking for, great. However, it does not test longer conversations with high token counts.
Just saying that all benchmarks, including lmsys, have issues and are focused on specific use cases.
Don't expect any meaningful score there before they wipe results.
Open source models are very exciting for self hosting, but the per-token hosted inference pricing hasn't been competitive with OpenAI and Anthropic, at least for a given tier of quality. (E.g.: Llama 3 70B costing between $1 and $10 per million tokens on various platforms, but Claude Sonnet 3.5 is $3 per million.)
[1]: https://github.com/meta-llama/llama-models/blob/main/models/...
[2]: https://github.com/meta-llama/llama-recipes/blob/main/recipe...
Have other major models explicitly communicated that they're trained on synthetic data?
We had a brief, abnormal, and special moment in time after the crypto wars ended in the mid-2000s where software products were truly global, and the internet was more or less unregulated and completely open (at least in most of the world). Sadly it seems that this era has come to a close, and people have not yet updated their understanding of the world to account for that fact.
People are also not great at thinking through the second order effects of the policies they advocate for (e.g. the GDPR), and are often surprised by the results.
Each new large regulation adds another category of company to the list of those who choose not to participate. Sure, you can always label them as companies who don't value principle X, but at some point it stops being the fault of the companies and you have to start looking at whether there are too many enormous regulations slowing down tech releases.
The word fault somehow implies that something’s wrong - from the eu regulator’s perspective, what’s happening is perfectly normal, and what they want : at some point, the advances in insert new tech are not worth the (social) cost to individuals, so they make things more complicated/ ask companies to behave differently.
Now I’m not saying the regulations are good, required, etc : just that depending on your goal, there are multiple points of view, with different landing zones.
I also suspect that what’s happening now ( meta, apple slowing down) is a power play : they’re just putting pressure on the eu, but I’m harboring doubts that this can work at all.
Why do you think he is surprised? I think very few are surprised.
Other than that, and GDPR (which is generally now regarded as a good thing), I'm not sure what requirements you've got in mind.
The only solution is a worldwide government that can impose laws in all countries at once, but that's unlikely to happen any time soon.
A Gibsonesque global Turing Police is a sure sign of Dystopia.
Let's hope the next moustached guy that tries to do this ends up dying in a bunker just like the last one.
https://aider.chat/docs/leaderboards/
77.4% claude-3.5-sonnet
75.2% DeepSeek Coder V2 (whole)
72.9% gpt-4o
69.9% DeepSeek Chat V2 0628
68.4% claude-3-opus-20240229
67.7% gpt-4-0613
66.2% llama-3.1-405b-instruct (whole) Llama 3 Training System
19.2 exaFLOPS
_____
/ \ Cluster 1 Cluster 2
/ \ 9.6 exaFLOPS 9.6 exaFLOPS
/ \ _______ _______
/ ___ \ / \ / \
,----' / \`. `-' 24000 `--' 24000 `----.
( _/ __) GPUs GPUs )
`---'( / ) 400+ TFLOPS 400+ TFLOPS ,'
\ ( / per GPU per GPU ,'
\ \/ ,'
\ \ TOTAL SYSTEM ,'
\ \ 19,200,000 TFLOPS ,'
\ \ 19.2 exaFLOPS ,'
\___\ ,'
`----------------'405B is hopelessly out of reach for running in a homelab without spending thousands of dollars. For most people wanting to try out the 405B model, the best option is to rent compute from a datacenter. Looking forward to seeing what it can accomplish.
On a related note, for those interested in experimenting with large language models locally, I've been working on an app called Msty [1]. It allows you to run models like this with just one click and features a clean, functional interface. Just added support for both 8B and 70B. Still in development, but I'd appreciate any feedback.
[1]: https://msty.app
Can you add GCP Vertex AI API support? Then one key would enable Claude, Llama herd, Gemini, Gemma etc
Let us know if you have other needs!
Too bad, too, I don't think my PC will fit 20 4090s (480GiB).
If you are running a commercial service that uses AI, you buy a few dozen A100s, spend a half million, and you are good for a while.
If you are running a commercial inferencing service, you spend tens of millions or get a cloud sponsor.
I have done experiments with 7B Llama3 Q8 models on a M3 MBP. They run faster than I can read, and only occasionally fall off the rails.
3B Phi-3 mini is almost instantaneous in simple responses on my MBP.
When I want longer context windows, I use a hosted service somewhere else, but if I only need 8000 tokens (99% of the time that is MORE than I need), any of my computers from the last 3 years are working just fine for it.
But also check out the 8B and 70B Llama-3.1 models which show improved benchmarks over the Llama-3 models released in April.
edit: If the AI bubble pops we will be swimming in GPUs... but no new models.
Open Source AI Is the Path Forward
https://about.fb.com/news/2024/07/open-source-ai-is-the-path...
Seems like the biggest GPU node they have is the p5.48xlarge @ 640GB (8xH100s). Routing between multiple nodes would be too slow unless there's an InfiniBand fabric you can leverage. Interested to know if anyone else is exploring this.
For home users 7B models (which can fit on an 8GB GPU) and 13B models (which can fit on a 16GB GPU) are in far more demand. If you're a researcher, you want a 70B model to get the best performance, and so your benchmarks are comparable to everyone else.
Such models will never top the number of downloads charts, or the community hype, as there’s just loads more people who can use the smaller models.
And if you can afford one 4090 you can probably afford two.
The perplexity per parameter is higher and the delta grows as it scales.
Not per bit, but per parameter.
Why this is happening really needs more attention and more consideration for pretrained model development right now.
A sleeping giant of a difference in a space where even marginal gains make headlines.
Times out.
And answer queries like:
Give all <myObject> which refer to <location> which refer to an Indo-European <language>.
https://github.com/meta-llama/llama-models/blob/main/models/...
Unless...
You have a couple hundred $k sitting around collecting dust... then all you need is a DGX or HGX level of vRAM, the power to run it, the power to keep it cool, and place for it to sit.
* You'll be running a Q5(ish) quantized model, not the full model
* You're OK with buying used hardware
* You have two separate 120v circuits available to plug it into (I assume you're in the US), or alternatively a single 240v dryer/oven/RV-style plug.
The build would look something like (approximate secondary market prices in parentheses):
* Asrock ROMED8-2T motherboard ($700)
* A used Epyc Rome CPU ($300-$1000 depending on how many cores you want)
* 256GB of DDR4, 8x 32GB modules ($550)
* nvme boot drive ($100)
* Ten RTX 3090 cards ($700 each, $7000 total)
* Two 1500 watt power supplies. One will power the mobo and four GPUs, and the other will power the remaining six GPUs ($500 total)
* An open frame case, the kind made for crypto miners ($100?)
* PCIe splitters, cables, screws, fans, other misc parts ($500)
Total is about $10k, give or take. You'll be limiting the GPUs (using `nvidia-smi` or similar) to run at 200-225W each, which drastically reduces their top-end power draw for a minimal drop in performance. Plug each power supply into a different AC circuit, or use a dual 120V adapter with a 240V outlet to effectively accomplish the same thing.
When actively running inference you'll likely be pulling ~2500-2800W from the wall, but at idle, the whole system should use about a tenth of that.
It will heat up the room it's in, especially if you use it frequently, but since it's in an open frame case there are lots of options for cooling.
I realize that this setup is still out of the reach of the "average Joe" but for a dedicated (high-end) hobbyist or someone who wants to build a business, this is a surprisingly reasonable cost.
Edit: the other cool thing is that if you use fast DDR4 and populate all 8 RAM slots as I recommend above, the memory bandwidth of this system is competitive with that of Apple silicon -- 204.8GB/sec, with DDR4-3200. Combined with a 32+ core Epyc, you could experiment with running many models completely on the CPU, though Lllama 405b will probably still be excruciatingly slow.
Assuming NUMA doesn't give you headaches (which it will) you would be looking at nearly 1 TB/s
Would love to hear your feedback!
Meta's goal from the start was to target OpenAI and the other proprietary model players with a "scorched earth" approach by releasing powerful open models to disrupt the competitive landscape.
Meta can likely outspend any other AI lab on compute and talent:
- OpenAI makes an estimated revenue of $2B and is likely unprofitable. Meta generated a revenue of $134B and profits of $39B in 2023.
- Meta's compute resources likely outrank OpenAI by now.
- Open source likely attracts better talent and researchers.
- One possible outcome could be the acquisition of OpenAI by Microsoft to catch up with Meta.
The big winners of this: devs and AI product startups
There is no defensible moat unless a player truly develops some secret sauce on training. As of now seems that the most meaningful techniques are already widely known and understood.
The money will be made on compute and on applications of the base model (that are sufficiently novel/differentiated).
Investors will lose big on OpenAI and competitors (outside of greater fool approach)
This is why Altman has gone all out pushing for regulation and playing up safety concerns while simultaneously pushing out the people in his company that actually deeply worry about safety. Altman doesn't care about safety, he just wants governments to build him a moat that doesn't naturally exist.
I work at OpenAI and used to work at meta. Almost every person from meta that I know has asked me for a referral to OpenAI. I don’t know anyone who left OpenAI to go to meta.
Classic strategy.
https://github.com/meta-llama/llama-models/blob/main/models/...