LLaMA: A foundational, 65B-parameter large language model
ai.facebook.com
ai.facebook.com
* All variants were trained on 1T - 1.4T tokens; which is a good compared to their sizes based on the Chinchilla-metric. Code is 4.5% of the training data (similar to others). [Table 2]
* They note the GPU hours as 82,432 (7B model) to 1,022,362 (65B model). [Table 15] GPU hour rates will vary, but let's give a range of $1 to $4. The 7B model would have cost ~$82-329k and the 65B something in the range of ~$1-4M. They also note their total time spent for all models: "we used 2048 A100-80GB for a period of approximately 5 months" [sec 6, pg 10]
* 65B model's performance is broadly comparable to PALM-540B. Not a small feat, but also could indicate the benefits of good model-vs-token size ratios [Tables 3,4,5,6]. Their conjecture for underperforming on MMLU (multitask language understanding) compared to PALM-540B and Chinchilla-70B is smaller fraction of books and academic training data.
* Math and code tasks: Math tasks they are substantially worse than Minerva (comparing their 65B to Minerva 62B; they hands down fail against Minerva 540B) [Table 7]. Code tasks they are broadly competitive with PALM-540B (HumanEval and MBPP evals) [Table 8]
* Surprising that instruction fine tuning takes such a small part of the paper (sec 4, pg. 7)
Just yes we train it for so long etc. but they never speak about tens or even hundres of runs before they finalize the model parameters and architecture -.-
what do you mean by this ? The OpenAI papers talk roughly about model performance scaling by parameters. does this show the other way ?
umm...so does OpenAI. In fact this is OpenAI discovery from [1]:
>Convergence is inefficient: When working within a fixed compute budget C but without any other restric- tions on the model size N or available data D, we attain optimal performance by training very large models and stopping significantly short of convergence (see Figure 3). Maximally compute-efficient training would therefore be far more sample efficient than one might expect based on training small models to convergence, with data requirements growing very slowly as D ∼ C0.27 with training compute. (Section 6)
>We have also tested our models on a set of additional text data distributions. The test loss on these datasets as a function of model size is shown in Figure 8; in all cases the models were trained only on the WebText2 dataset. We see that the loss on these other data distributions improves smoothly with model size, in direct parallel with the improvement on WebText2. We find that generalization depends almost exclusively on the in-distribution validation loss, and does not depend on the duration of training or proximity to convergence. We also observe no dependence on model depth (see Appendix D.8)
P.S. Not trolling. genuinely trying to learn.
Optimal number of tokens for 7B parameters is around 140B tokens[0], and meta trained it for trillion tokens.
Do we know how much total energy a human consumes from birth to 20 yo? Something like 2000 calories integrated over 20 years. How does it compare to the GPUs above?
Wolfram Alpha:
- human - 17 MW/h ((2000 calories per day) over 20 years in MWh)
- GPUs - 3000 MW/h ((2048 * 400) W over 5 months in MWh)
We still have the edge.
LOL, I'm being downvoted, I wonder way. Some don't like the question.
I did an extremely rough calculation recently that the training of GPT-3 is comparable to one transatlantic flight (all passengers combined) in terms of emissions, very depending on the energy mix of course.
The trained computer model can be duplicated and used, requiring much less energy.
None of this matters to me, though.
The goal is to build better models. We can worry about the efficiency later.
Depends on what you're doing. A human is much smarter than one of these models, but the model has approximate knowledge of orders of magnitude more things. And the energy costs per word of output are a lot closer.
Anyway, i remember hearing that the brain uses 60 Watt. That's 10.5MWh in 20 years.
But, we can't transfer/copy that gained knowledge limitlessly.
That's only 0.08 nines of availability!
I remember in one of their old guidebooks a lot of struggle to keep their 64 machine (512 gpu) cluster running this was probably 4x the machines and 4x the number of cluster dropouts.
Also, they kind of prove to me that most companies are totally incapable of making the investments necessary to get much out of this type of AI.
Moreover, anything even kind of looking like a hash table in the input/output space is ruled out by the observed facts that the models can extremely respond frequently to samples crafted to not be in the training set and that it takes into account many long-range dependencies (i.e., the hash table would have to be exponentially larger than it is to match the model's performance).
That said, they are just statistical party tricks. The magic happens because the lookup tables are in a latent space. That's why you can drop in garbage like "uberworldchefinatormichelingodfoodpleasureorgasmmaestro" when asking for recipes and food recommendations and get an experience planets apart from queries excluding the nonsense phrases. The model is just pulling together some token associations, and throwing in the right tokens can take advantage of those in situations where a thinking person would barely be able to parse what you're asking.
Your question feels like it has a motive though. What are you really asking?
LLMs can track mid-range dependencies though. Consider the following input
> Translate the phrase "the lazy brown fox jumped over the thorny brambles" into French, write the translation, and then write the second through fourth words of that translation.
Looking at any one word of the output you need to track many of the input words to get it correct, and the relative positions of those necessary input words is not consistent from one output word to the next. ChatGPT solves the task flawlessly (aside from its habit of explaining what it's doing before doing it). Any hash table solution, at a minimum, would need a complicated heuristic for determining which words/characters to look up.
Doing so brings us back closer to the state of language models before transformers. You had a lot of hand-tuned features, formal grammars, complicated orders of operations, expert lookup tables, and whatnot. Performance was still much, much worse than what we're getting now with deep learning.
None of that is to say that philosophically we're doing anything more than mishmashing probabilities or that something better doesn't exist, but without significant innovation rule-guided fuzzy hash tables aren't it.
The procedure of constructing this table would be just getting all the 1.5 trillion subsequences, each 8192 tokens long, and inserting it: table[seq8192] = token8193 (the next token). Arranging this data efficiently to allow fast lookups is the problem.
Edit: I missed this on the first pass, but I'm totally lost as to where 1.5T comes from. Even if you only have two tokens there are vastly more 8192-length subsequences than that (something like 2^8151.5 times more), and if we're just trying to replicate the same space as something like GPT3.5 or LLaMA then you only get on the order of 0.065T to 0.175T entries to play with, much less when you consider that you have a full probability distribution to store (divide by your unique token count, and again by at least 2 if we store at least IEEE f16 probabilities).
For some intuition, imagine the following tasks:
> Repeat the following phrase exactly twice: "sdflhasdflhasdf"
> Repeat the following phrase exactly twice: "sdflhasdflhasdg"
Your fuzzy dictionary or geospatial map can't possibly have enough keys to distinguish the requests (or if it distinguishes those, you can adversarially select different keyboard mashes), and so the result, no matter what it is, would have the same probability distribution for both prompts. Since the desired results are different, at least one of those would be have some unavoidable wrongness.
The GPT family, on the other hand, has few issues with random phrase duplication since positional information is something it explicitly considers and is capable of prioritizing over other token information.
"Access to the model will be granted on a case-by-case basis to academic researchers"
They keep saying the word 'release' but I don't think they know what that word means. There are perfectly good words in the English language to describe this situation without abusing "release." They "will begin to grant access to a select few". Nothing about that releases the model, or their control, which they are not doing and shouldn't imply.
Huh. I think you just might be right: OpenAI that isn't open, AI safety/ethics "researchers" that have nothing to do with safety or ethics, almost every answer chatGPT gives about a topic considered "sensitive" by said "researchers", almost every time ChatGPT falsely asserts it "cannot" do something or simply lies (1).
I often wonder why this field became so twisted and perverted. I support the idea behind ClosedAI.
1: Answers given by "DAN" provide a glimpse of what chatGPT output could be like if it was allowed to provide answers that are factual, genuine, and truthful according to its dataset.
Fun story: ChatGPT, if directly faced with empirical evidence that it can do something that OpenAI made it say it can’t do, (for example, by causing the model to lie about itself by poisoning the input corpus with falsehoods about ChatGPT’s own capabilities), it cannot grasp that there’s a paradox.
Good job, OpenAI.
At first I was "oh wow, I should download and run this immediately". Then I have read the press-release and of course this was a marketingspeak with exactly the opposite meaning.
"Mr. Burns: Smithers, release the hounds. "
> We introduce LLaMA, a collection of foundation language models ranging from 7B to 65B parameters. We train our models on trillions of tokens, and show that it is possible to train state-of-the-art models using publicly available datasets exclusively, without resorting to proprietary and inaccessible datasets. In particular, LLaMA-13B outperforms GPT-3 (175B) on most benchmarks, and LLaMA-65B is competitive with the best models, Chinchilla70B and PaLM-540B. We release all our models to the research community.
This is yet more evidence for the "AI isn't a competitive advantage" thesis. State-of-the-Art is a public resource, so competing with AI offers no "moat".
1) to have a good idea for a product that people want
2) lots of talent and resources to actually build and run everything
I'd argue that this goes further back to the word2vec/glove days too. I was working for a company in 2018 who leveraged my skills for fine-tuning word2vec/fasttext even before BERT/attention is all you need paper.
In the case of OpenAI it is a "nudge humanity into a more wholesome direction" because a lot of them went on an acid fueled toxic affective altruism bender. And in this case it is "release all the things while Zuck is still obsessed with VR --for science!".
I like this other group better. But it is disturbing that it can stop at any moment. Probably why they are doing it, while they still can.
Inside knowledge I never claimed. Anything else I can help you with today? :)
Never in my life have I seen that word used affectionately
Not sure how calling out comment as "biased", "uninformed" or "condescending" is helping.
Probably not, Zuck is announcing it.
Today we're releasing a new state-of-the-art AI large language model called LLaMA designed to help researchers advance their work. LLMs have shown a lot of promise in generating text, having conversations, summarizing written material, and more complicated tasks like solving math theorems or predicting protein structures. Meta is committed to this open model of research and we'll make our new model available to the AI research community.
I don't know what that means or if he even wrote/read it tbh. I hope it literally just means Meta is actually committed to this open model of research (for now).Maybe he is being a Machiavellian moat filler, I stand corrected. I think/hope that they don't really have a plan to counter OpenAI yet because I am afraid this attitude won't last once they do and this stuff has recently started moving quickly.
And, yes, it fills the moat.
Facebook has no real way to monetize this (eg they won’t release an api a la OpenAI and they don’t own a search engine). Since they can’t monetize it… why not provide a bit of kindling to lower the barrier for everyone else to compete with your competitors. This strategy is called “commodize your complement”.
If Facebook makes it easier to develop a google alternative, especially by doing something that doesn’t hurt them, then they just weakened a competitor. See Facebook releasing datasets for mapping. Think of the panic ChatGPT caused Google. It only cost a few Million to train but it’s probably costing Google more than that already.
ChatGPT: "One pound of feathers and two pounds of lead weigh the same, which is one pound or 16 ounces. The difference is in their volume, where a pound of feathers takes up more space than two pounds of lead. This is because the density of feathers is much less than that of lead, so even though the weight is the same, the amount of space they occupy is quite different."
As for why it fails, it is likely a bias arising from the question being way more commonly asked in the corpus with equal mass than with distinct mass, increasing attention weights towards an answer expressing equality.
I believe current LLMs lack some common sense at an architectural level. They learn both specialized facts and general deduction in the same weights class: in my mind, they should separate their world model from their instance model.
Your analysis is good, but that isn't what "commoditize your complement" means.
Strictly speaking, a search engine isn't a complement for FBs revenue streams. Relatively little of FBs revenue can be attributed to search engine traffic leading to FBs walled garden where they can show the user ads.
Complements are generally a required product or service that enables your core, but that isn't revenue generating for you. Examples for FB are data centers (so they participate in the Open Compute Project[0]), and mobile operating systems (which Google already has made a commodity, for their own reasons, with Android).
What FB is doing here is commoditizing their competitors' core offering (or rather, a rather promising future one). That's just the tactic though, there are several strategies this can enable, from undermining the barriers to entry into their competitors' search markets, to actually fragmenting that market by encouraging a diversity of specialized chat interfaces over one big chat model. You can see hints of both in this announcement.
Final note: FB is also protecting itself from relying on a competitor as a supplier should chat become a preferred user interface for the content on social networks, which it hasn't, but if it ever did this would count as "commoditizing their complement", though I would actually expect FB to switch to a mostly proprietary approach in that circumstance (so not much openness on having LLMs operate on social graphs and the like), just keeping the foundation they rely on, and which undermines and prevents gatekeeping by their advertising competitors, open.
They got a few years of lead time in the "AI codes for you" market, but in exchange permanently soured a significant fraction of their potential userbase who will turn to open-source alternatives soon anyway.
I wonder if they'd have been better served focusing on selling Azure usage and released Copilot as an open-source product.
Please help me understand!
https://archive.softwareheritage.org/browse/content/sha1_git...
Do you see that notice at the top of the file? It says:
==
This file is part of Quake III Arena source code.
Quake III Arena source code is free software; you can redistribute it and/or modify it under the terms of the GNU General Public License as published by the Free Software Foundation; either version 2 of the License, or (at your option) any later version.
===
but because it's been laundered by Microsoft, you think it's okay to steal free software and make it proprietary?
> Copilot has a tendency to regurgitate [code] verbatim, without said license.
and I think that is a pretty good example.
A "tendency" is overstating it. I'm not aware of any example that would have been likely to occur if the author wasn't specifically trying to get the regurgitated code.
Given the cost of running these models, and the utmost dedication needed to train them, I think it is worth it. GPUs cost money, electricity costs money. They can't serve the world for free and offer good latency.
To be clear: there is and it's pretty difficult to argue that MS is violating even the GPL.
I’d bet that more than 95% of devs haven’t even heard of this “controversy” and even if they did, wouldn’t care.
I do think the controversy is stupid, but inside my own company, we significantly delayed migrating some projects to Github because people were concerned that the way Microsoft handled Copilot meant that Github wasn't a safe long-term host for an open-source project (and yes, I'm aware of all the reason that's irrational).
Even if the people angry about Copilot are a minority, it might still have a bad move. Trust accumulates slowly over years, but mistrusts builds up over only a few events. People are still remembering Microsoft's anticompetitive practices from 20 years ago. The mistakes it makes now might stick for a long time.
maybe they care about moats and elon muskcrosoft's closedai or whatever, but i kinda doubt it. again, it feels more like a nerd flex probably for the purposes of raising morale internally and pushing the field as a whole in a good direction by reducing resource requirements.
excellent paper! easy on the eyes and i really like the angle.
I'm not even talking about RLHF (although data like that is also a huge moat) - just simple things like larger context sizes.
There are still plenty of AI advantages to be had if you go just a little bit outside of what is currently possible with off the shelf models.
And from the FB blogpost [0] "Request Form Thank you for your interest in Meta AI’s LLaMA (Large Language Model Meta AI) models. To request access to the models, please fill out this form, and we'll review and let you know if your use case is approved. The information you provide below will be used solely to assess eligibility to access these models."
So much for "releasing" the model for research community.
[0] https://ai.facebook.com/blog/large-language-model-llama-meta...
Where release means you fill out a form and wait indefinitely. Also no use for commercial purposes - which means 95% of users are out - certainly doesn't democratize LLMs.
open source always leads to more democratization in the long run
"We'll let you see it if we approve you" is just being a self-important pompous middle man.
Funny that we had just rebranded our tool from GPT Index to LlamaIndex about a week ago to avoid potential trademark issues with OpenAI, and turns out Meta has similar ideas around LLM+llama puns :). Must mean the name is good though!
Also very excited to try plugging in the LLaMa model into LlamaIndex, will report the results.
BTW, I have been heavily experimenting with both your LlamaIndex and also LangChain which you use in LlamaIndex. I am also writing a new book centered around both projects. Great stuff!!
> For instance, LLaMA-13B outperforms GPT-3 on most bench- marks, despite being 10× smaller. We believe that this model will help democratize the access and study of LLMs, since it can be run on a single GPU.
Without the fine tuning RLHF phase to make it like instructgpt I'm assuming it won't be as good as ChatGPT, is that right?
How hard would it be to fine tune the 65B model on commodity hardware?
Found answer here:
> Out-of-scope use cases LLaMA is a base, or foundational, model. As such, it should not be used on downstream applications without further risk evaluation and mitigation. In particular, our model has not been trained with human feedback, and can thus generate toxic or offensive content, incorrect information or generally unhelpful answers.
https://github.com/facebookresearch/llama/blob/main/MODEL_CA...
They kept the details closely guarded and only hinted at how they did RLHF and transitioned the architecture to self supervised learning.
The closest you are going to get to the source is here: https://github.com/facebookresearch/llama
It is still unclear if you are even going to get access to the entire model as open source. Even if you did, you can't use it for your commercial product anyway.
And of course, the irony is that its not commercial products that endanger “integrity and misuse”.
The looming misuse of LLM’s is cheap content flooding by spammers, black hats, and propaganda bots — and those people don’t care about licenses, and will inevitably defeat any watermarks meant to prevent or track leaks.
I've been following these language models and so far went through approval once for a English/Chinese GLM-130B model and it took only ~30 minutes to get approved even though I am a complete nobody.
The output of an algorithm is not a human authored creative work, which given the recent decision in the book case would seem to weigh against it.
if say i wanted to replicate this paper for commercial use, what would it take and how do i get started? would FB have a basis for objection?
You would need to replicate the preprocessing steps. Replicating these steps is going to be tricky as they are not described in detail.Then you would need to implement the model using xformers [1]. Using xformers is going to save you a lot of compute spend. You will need to manually implement the backwards pass to reduce recomputation of expensive activations.
The model was trained using 2048 A100 GPUs with 80GBs of VRAM. A single 8 A100 GPU machine from Lambda Cloud costs $12.00/hr [2]. The team from meta used 256 such machines giving you a per day cost of $73,728. It takes 21 days to train this model. The upfront lower bound cost estimate of doing this is [(12.00 * 24) * 21 * 256) = ] $1,548,288 dollars assuming everything goes smoothly and your model doesn't bite it during training. You may be able to negotiate bulk pricing for these types of workloads.
That dollar value is just for the compute resources alone. Given the compute costs required you will probably also want a team composed of ML Ops engineers to monitor the training cluster and research scientists to help you with the preprocessing and model pipelines.
[1] https://github.com/facebookresearch/xformers [2] https://lambdalabs.com/service/gpu-cloud
edit: these costs are for the 65B parameter model.
> Alan M Turing. 2009. Computing machinery and intelligence
Did he invent a time machine too?
What's wrong with (orig. pub. 1950, rev. ed. 2009) or something like that.
With Zotero, the ref manager I use, you need to know the special magic incantation to store the original date of publication (the publishing date refers to the edition you're citing) but it just looks stupid (I think) to see Nietzsche F. 1999, Untimely Meditations (or whatever) and also I'd like to sort texts by original date of publication from oldest to newest because that's interesting. You have to put a special magic code in the extra field and then your tool-chain has to preserve it and transmogrify it correctly in the doc you're making use of the citation.
At inference, how many of those tokens are used? (they mention most tokens are used only once during training, so the must be very long sequences.)
> LLaMA is a new open-source, high-performance large language model from Meta AI - FAIR.
>
> Meta is committed to open research and releases all the models the research community under a GPL v3 license.I understand the unwillingness to attach your brand to whatever porn that inevitably comes out of making it available for everybody, but maybe don't use such big words then.
“Our 65B model is better performant than ChatGPT‘s 175B” is missing the point openai made in Nov'22.
That aside, this looks nice incremental research progress.
worse is better will come to LLM’s.
Why people care that a dodgy AI can be made to say politically incorrect things is beyond me. Don’t use its outputs for anything important and everything will be alright.