1) 7B foundational model
2) 8K length
3) 1.5T tokens
1) 7B foundational model
2) 8K length
3) 1.5T tokens
- 7B means 7 billions parameters.
- 8K length means the size of input/output is 8K tokens.
- 1.5T tokens mean the training set has 1.5T tokens.
A: What's a parameter?
Q: More parameters your model has, more complex relationship it can represent. For example let's say you have a function f(x). This is a 2-parameter model:
f(x) = ax + b
This is a 4 parameter model:
f(x) = ax^3 + bx^2 + cx + d
As you can see as the number of parameters grows, the function is able to represent more complex relationship between f(x) and x.
A: What's a token?
Token is a way to encode text, like ASCII or Unicode. Unlike Unicode, tokenizor usually favors common combinations of alphabets. For example, "the" is a single token for GPT-3 tokenizor, but "eht" is two tokens (e and ht).
* Note that the number of parameters is more like an "upper limit" of the model's capabilities. If your a, b, c, d are just random shit, it's still a 4-parameter model, but it's still useless. The whole concept of "training" is just "finding the best parameters".
2) Currently every model that can run locally was trained with a 2K context size. It's a hard limit on prompt length. There have been recent advances with [A] position interpolation, but those methods explore fine-tuning/loras. This base model was trained with 8k sequences.
3. 1.5T tokens is the size of the total training corpus. Training cost and time increases with training size. [B]
A. https://arxiv.org/abs/2306.15595
B. https://www.semianalysis.com/p/the-ai-brick-wall-a-practical... (Jan 2023)
"7B" refers to the number of parameters or weights for a model. For a specific model, the versions with more parameters take more compute power to train and perform better.
A foundational model is the part of a ML model that is "pretrained" on a massive data set (and usually is the bulk of the compute cost). This is usually considered the "raw" model after which it is fine-tuned for specific tasks (turned into a chatbot).
"8K length" refers to the Context Window length (in tokens). This is basically an LLM's short term memory - you can think of it as its attention span and what it can generate reasonable output for.
"1.5T tokens" refers to the size of the corpus of the training set.
In general Wikipedia (or I suppose ChatGPT 4/Bing Chat with Web Browsing) is a decent enough place to start reading/asking basic questions. I'd recommend starting here: https://en.wikipedia.org/wiki/Large_language_model and finding the related concepts.
For those going deeper, there are lot of general resources lists like https://github.com/Hannibal046/Awesome-LLM or https://github.com/Mooler0410/LLMsPracticalGuide or one I like, https://sebastianraschka.com/blog/2023/llm-reading-list.html (there are a bajillion of these and you'll find more once you get a grasp on the terms you want to surf for). Almost everything is published on arXiv, and most is fairly readable even as a layman.
For non-ML programmers looking to get up to speed, I feel like Karpathy's Zero to Hero/nanoGPT or Jay Mody's picoGPT https://jaykmody.com/blog/gpt-from-scratch/ are alternative/maybe a better way to understand the basic concepts on a practical level.
2. It can handle upto 8k tokens. Tokens are usually some representation for a word. If your tokens are characters then, "h", "e", "y" represent 3 tokens for hey. Most of the algos use byte pair encoding. For example "hand-le" has two tokens "hand" and "le". This is a very crud example which is enough to give the gist but is not accurate. You can look into byte pair encoding for more details.
3. The token size 1.5T token means they have huge variations for input and output. Simply put, it was trained on large data corpus.
I hope this simplifies it. You can research further if you are interested! Hope it helps!
This one doesn't even make any sense. Of course it doesn't have 7B parameters _per_ neuron.
I hope this clarifies the answer now.
Now that is done I am quite curious on how you came up with the idea it was written by ChatGPT? I just wanted to simplify as best as I could. It’s funny you thought it that way.
What could I have done so that it didn’t sound like response from ChatGPT? I am asking it to prevent future misunderstandings. I thought my grammatical errors would be enough to show it wasn’t a ChatGPT response.
Looking forward to your reply!