XGen-7B, a new 7B foundational model trained on up to 8K length for 1.5T tokens
blog.salesforceairesearch.com
blog.salesforceairesearch.com
This is huge.
MPT and Falcon are cool, but the inference runtimes and various tooling is mostly optimized for LLaMA. If this is a drop-in replacement for 7B, it's going to catch on much faster than any other small model.
XGen-7B is probably the superior 7B model, it's trained on more tokens and a longer default sequence length (although both presumably can adopt SuperHOT (Position Interpolation) to extend context), but larger models still probably perform better on an absolute basis.
What use cases do people have for these smaller LLM's?
You don't have to go to openai.
https://www.reddit.com/r/LocalLLaMA/wiki/models/ gives you a list of VRAM requirements to load the model into GPU VRAM. the more VRAM the computer has, the larger the model you can load in, thus making 3090s the current consumer grade king due to price to max VRAM.
This being said however most models are LLAMA based which all fall under that specific research license.
So following the rules, you would be limited to a subset of models which are foundational models which allow for commercial use
I can barely run 33B, but anything more than 800 context and I oom. But it would run very comfortably on a bigger GPU or a 24GB+ laptop.
Theoretically some phones can comfortably handle 13B on mlc-llm though in practice its not really implemented yet.
None. Training a functionally useless model and releasing it is a great way to demonstrate that your company is hip and current. That way when prospective clients ask about AI you can vaguely gesture at some model that you released and say you employ cutting edge AI experts.
If it was good, then they’d charge for it.
…what they (and everyone) is gonna do is play with smallish models to iterate on the process for relatively small expense and earn karma.
Then pay big $$$ to make a really good model for internal use and/or an api that people have to pay for.
Tldr; it’s free. It’s by sales force. You should expect it to be a) crippled and b) a loss leader for a paid product.
Not judging; it’s a fair strategy. Just saying: salesforce is not a company that just gives hundreds of thousands of dollars away for nothing.
If you want a good free open model, you’re kidding yourself if you think a corporate giant is going to kiss you on the head and give it to your for free.
Yep! That makes sense!
I would love to be the CEO of the company that does give away an actually useful model and little forehead kisses though. The amount of goodwill that one would generate from that would be astronomical and training costs are getting so low that nearly any company with enough cash could do it.
I look forward to waking up and hearing that the Nabisco/Canadian Tire/A&W usefully-tuned model is revolutionizing the economy and seeing the infinite amount of good press that it would generate.
The problems with chatGPT are many - dependence on third party, privacy, externally imposed ideology and rules, cost, and most importantly - prompting is context-size limited and token-expensive, you can't pack much data into it.
Fine-tuning is a more powerful approach where you can actually fix the model problems instead of futzing around with the prompt and demonstrations. Yes, you got to work on your dataset. But if you don't already have it you can bootstrap with GPT-4 for a small sum.
Meta provided the training wheels - LLaMA, every company tried fine-tuning it for their purposes, but could not proceed for lack of a commercial base model. Salesforce XGen and a few other open-small-LLMs (funny how that sounds!) open the flood gates.
So the recipe is: use an existing dataset, or make one with regular GPT-4 prompting and a bit of curation. Then fine-tune a small open model. You can get it to be better than stock GPT-4, cheaper, faster and private. If you use LoRA's you can save each skill in a separate diff model just 1% the size of the base model and use a single GPU to fine-tune it, in a single day.
These free models are both a defensive move against behemoths, and kindling to rapid business development.
If you're relying on prompting for the 7B models IMO you're gonna have a bad time — they're mostly toys at that: interesting output but not consistently useful. But finetuning gets better results, and it's cheap to finetune.
The metrics are good though, perhaps placing this closer to 13B.
And 8K context is huge. When you can stuff that much example text in, it gives the model more to "latch onto," and its also the point where you would start worrying about RAM/VRAM consumption for a ~13B model.
QLORA is the gold standard for more affordable training.
As for datasets, just look at the open datasets the best-in-class models are using, like Vicuna or https://huggingface.co/NousResearch/Nous-Hermes-13b
Some model datasets like Manticore, Chronos or the infamous Pygmalion are more "secretive," but you can find the dataset gathering scripts on Github or in community chats.
You can easily finetune 7B or 15B LORA model with that on consumer GPUs.
It's a rabbithole, and unfortunately there's no good shortcuts.
There are startups that do finetuning on your own data, but with zero hints on how to preprocess your data and absurd costs (both upfront training and GPUs for serving inference) that's it's extremely difficult to argue from a customer business perspective compared to just using an API.
Vapourware GPT startup inc is valued at $2bn the afternoon after you form the company and buy your first macbook.
Actual usage of Ai, fine tuning etc. I can offer you $100,000 for 30% of your company if you can demonstrate a fully working product.
If all you have is an M1 or whatever, ya, you need a real workstation and depending on your use ChatGPT might be cheaper/better.
You still need the memory to be able to go that high, but it's totally doable.
But I assumed full training would give better perplexity for large contexts, and perhaps this method would be more effective at 16K+ with an 8K model to start with.
1) 7B foundational model
2) 8K length
3) 1.5T tokens
2. It can handle upto 8k tokens. Tokens are usually some representation for a word. If your tokens are characters then, "h", "e", "y" represent 3 tokens for hey. Most of the algos use byte pair encoding. For example "hand-le" has two tokens "hand" and "le". This is a very crud example which is enough to give the gist but is not accurate. You can look into byte pair encoding for more details.
3. The token size 1.5T token means they have huge variations for input and output. Simply put, it was trained on large data corpus.
I hope this simplifies it. You can research further if you are interested! Hope it helps!
This one doesn't even make any sense. Of course it doesn't have 7B parameters _per_ neuron.
I hope this clarifies the answer now.
Now that is done I am quite curious on how you came up with the idea it was written by ChatGPT? I just wanted to simplify as best as I could. It’s funny you thought it that way.
What could I have done so that it didn’t sound like response from ChatGPT? I am asking it to prevent future misunderstandings. I thought my grammatical errors would be enough to show it wasn’t a ChatGPT response.
Looking forward to your reply!
- 7B means 7 billions parameters.
- 8K length means the size of input/output is 8K tokens.
- 1.5T tokens mean the training set has 1.5T tokens.
A: What's a parameter?
Q: More parameters your model has, more complex relationship it can represent. For example let's say you have a function f(x). This is a 2-parameter model:
f(x) = ax + b
This is a 4 parameter model:
f(x) = ax^3 + bx^2 + cx + d
As you can see as the number of parameters grows, the function is able to represent more complex relationship between f(x) and x.
A: What's a token?
Token is a way to encode text, like ASCII or Unicode. Unlike Unicode, tokenizor usually favors common combinations of alphabets. For example, "the" is a single token for GPT-3 tokenizor, but "eht" is two tokens (e and ht).
* Note that the number of parameters is more like an "upper limit" of the model's capabilities. If your a, b, c, d are just random shit, it's still a 4-parameter model, but it's still useless. The whole concept of "training" is just "finding the best parameters".
"7B" refers to the number of parameters or weights for a model. For a specific model, the versions with more parameters take more compute power to train and perform better.
A foundational model is the part of a ML model that is "pretrained" on a massive data set (and usually is the bulk of the compute cost). This is usually considered the "raw" model after which it is fine-tuned for specific tasks (turned into a chatbot).
"8K length" refers to the Context Window length (in tokens). This is basically an LLM's short term memory - you can think of it as its attention span and what it can generate reasonable output for.
"1.5T tokens" refers to the size of the corpus of the training set.
In general Wikipedia (or I suppose ChatGPT 4/Bing Chat with Web Browsing) is a decent enough place to start reading/asking basic questions. I'd recommend starting here: https://en.wikipedia.org/wiki/Large_language_model and finding the related concepts.
For those going deeper, there are lot of general resources lists like https://github.com/Hannibal046/Awesome-LLM or https://github.com/Mooler0410/LLMsPracticalGuide or one I like, https://sebastianraschka.com/blog/2023/llm-reading-list.html (there are a bajillion of these and you'll find more once you get a grasp on the terms you want to surf for). Almost everything is published on arXiv, and most is fairly readable even as a layman.
For non-ML programmers looking to get up to speed, I feel like Karpathy's Zero to Hero/nanoGPT or Jay Mody's picoGPT https://jaykmody.com/blog/gpt-from-scratch/ are alternative/maybe a better way to understand the basic concepts on a practical level.
2) Currently every model that can run locally was trained with a 2K context size. It's a hard limit on prompt length. There have been recent advances with [A] position interpolation, but those methods explore fine-tuning/loras. This base model was trained with 8k sequences.
3. 1.5T tokens is the size of the total training corpus. Training cost and time increases with training size. [B]
A. https://arxiv.org/abs/2306.15595
B. https://www.semianalysis.com/p/the-ai-brick-wall-a-practical... (Jan 2023)
OTOH, there ought to be a construction in the form of an web app that can pinch out nonrepetitive, coherent ~100 page trashy romance novels in the style of any author given name with open source or specific text(s) or transcripts with enough original input volume: Churchill, The Unabomber, psycho happy kindergarten child development IEP manual writer, The Dude, Walter (agro gun nut), Bob Ross, Grace Hopper, Ayn Rand, LBJ, The Dalai LaMa%, Hitler, Kanye (Ye), Bhad Bhabie, and the King James Bible. Ethical and generational safety features be damned; it'd be generating fucking^2 art for hilarious entertainment purposes. How does one stretch training input to something that might involve human/computer output validation to discard sticking on repetitive nonsense?
% He never saw that one coming Ow^(3 + i).
With this model (and they say this in the blog post), they were testing the hypothesis that training on a longer context size would provide more performance at the same parameter count/inference FLOPs. From a quick perusal of their post, it looks like this was true, and we should train all future models with as long of a context size as we can afford.
https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderb...
salesforce air esearch ;)
Many researchers are improving very fast, and I would bet that soon we will see more efficient LLMs.