Building LLMs from the Ground Up: A 3-Hour Coding Workshop
magazine.sebastianraschka.com
magazine.sebastianraschka.com
Anyway I will watch it tonight before bed. Thank you for sharing.
https://www.manning.com/books/build-a-large-language-model-f...
(I've taken it from the footnotes on the article)
how Llama and OpenAI could be cleaning and structuring their training data
If you're interested in this, there are several sections in the Llama paper you will likely enjoy:https://ai.meta.com/research/publications/the-llama-3-herd-o...
edit: grammar
What turned GPT into chatGPT was a lot of structured training with human feedback.
It's a fine PyTorch tutorial but let's not pretend it's something low level.
To someone who already knows excel, from scratch with excel sheets instead of python may work with them.
Although if your claim is then that most programmers do not care about being effective, that I would tend to agree with given the 64 gigs of ram my basic text editors need these days.
While I agree it's good to know how your collections work. "Efficient key-value store" may be enough to use it effectively 80% of the time for somebody dabbling in Python.
Sadly I've met enough people that call themselves programmers that didn't even have such a surface level understanding of it.
Calling that from scratch is like saying "Just go to the store and tell them what you want" in a series called: "How to make sausage from scratch".
When I want to know how to do X from scratch I am not interested in "how to get X the fastest way possible", to be frank I am not even interested in "How to get X in the way others typically get it", what I am interested in is learning how to do all the stuff that is normally hidden away in dependencies or frameworks myself — or, you know, from scratch. And considering the comments here I am not alone in that reading.
For example, building a web server from scratch - I’d probably assume the presence of a sockets library or at the very least networking card driver support. For logging and configuration I’d assume standard I/o support.
It probably comes down to what you think makes LLMs interesting as programs.
E.g. when a title says it shows you how to do a thing in vanilla javascript from scratch bringing in jquery in the first step makes that tile a lie. If you bring in a hefty dependency on step 1 and run three imported function the vanilla javascript part might be fine, but the from scratch starts to do some heavy lifting.
It’s definitely out there and in productive use.
Which engines in particular? I never found especially flexible ones.
Unfortunately, your argument is a well known fallacy and carries no weight.
Never had a piano, not even a fortepiano .. though reportedly he played one once.
E.g.
1. Programming basics
2. How to manipulate text using programs (reading, writing, tokenization, counting words, randomization, case conversion, ...)
3. How to extract statistical properties from texts (ngrams, etc, ...)
4. How to generate crude text using markov chains
5. Improving on markov chains and thinking about/trying out different topologies
Etc.
Sure markov chains are not exactly LLMS, but they are a good starting point to byild a intuition how programs can extract statistical properties from text and generate new text based on that. Also it gives you a feeling how programes can work on text.
If you start directly with a framework there is some essential understanding missing.
Like driving a car, you don't need to understand what's under the hood you start driving, but eventually understanding it makes you a better driver.
But it doesn’t start there. It uses top-down pedagogy, instead of bottom up.
NVDIA value lies only in pytorch and cuda optimizations with respect with pure c implementation, so saying that you need go lower level than cuda or pytorch means simply reinventing Nvidia. Good luck with that
2. I didn't say CUDA wouldn't be ground up or low level (please re-read) (I say in another comment about a no-code guide with CUDA, but it's obviously a joke)
3. And finally, I think your comment comes out as holier than thou and finger pointing and making a huge deal out of a minor semantic observation.
Pytorch is low level enough to understand and interpret each and every passage. In pytorch, you can use builtin transformers, or code them yourself down to the "lowest" level in which there's still a theoretical meaning. So pytorch is just a tool and your comment was just pompous and empty.
At least when i saw the "Building LLMs from the Ground Up" what i expected was someone to open vim, emacs or their favorite text editor and start writing some C code (or something around that level) to implement, well, everything from the "ground" (the operating system's user space which in most OSes is around the overall level of C) and "up".
To a statistician or a practitioner approaching machine learning from a mathematical perspective, the computational details are a distraction.
Yes, these models would not be possible without automatic differentiation and massively parallel computing. But there is a lot of rich detail to consider in building up the model from first mathematical principles, motivating design choices with prior art from natural language processing, various topics related to how input data is represented and loss is evaluated, data processing considerations, putting things into context of machine, learning more broadly, etc. You could fill half a book chapter with that kind of content (and people do), without ever talking about computational details beyond a passing mention.
In my personal opinion, fussing over manual memory management is far afield from anything useful unless you want to actually work on hardware or core library implementations like Pytorch. Nobody else in industry is doing that.
People are looking at the ground up for a clear picture of what the thing is actually doing, so masking the important part of what is actually happening, then calling it “ground up” is disingenuous.
If you are interested in how the model works conceptually, how training works, how it represents text semantically, etc., then I maintain that computational details are an irrelevant distraction, not an essential foundation.
How about another analogy? Is SICP not a good foundation for learning about language design because it uses Scheme and not assembly or C?
But if all is relative and depends on your PoV that implies that there isn't actually a problem here, right? :-P
I don't think there is anything wrong with "building up the model from first mathematical principles" as you wrote, it just wasn't what i personally had in mind with the "from the ground up" part.
And FWIW i'm not that stuck up on the "vim and C" aspect, i used those as an example that i expected most would understand and leave little room for misinterpretation in what you'd have to work with (i.e. very very little) and have to implement yourself (pretty much everything) - personally i'd consider it from "the ground up" even if it was in C#, D, Java, JavaScript or even Python, as long as the implementation was done in a way that didn't rely on 3rd party libraries so that whatever is implemented in, say, Java could also be implementable in C#, D, JavaScript or Python with just whatever is available out of the box in those languages or even C, if one doesn't mind writing the extra bookkeeping functionality themselves.
https://16x.engineer/2023/12/29/nanoGPT-azure-T4-ubuntu-guid...
What sort of things could you do with it? How do you train it on current events?
Beyond learning how it all works and demo, there is not much practical usage. You can train it on current events if you feed that corpus during training instead of just OpenWebText. Shouldn't be hard.
I also dislike software development as it reminds me of developing a photograhic negative – like "oh let's check out how the software we developed came out".
It should be software engineering and it should be held to a similar standard as other engineering fields if it isn't done in a non-professional context.
Wrong angle. There is a problem, your consideration of the problem, the refinement of your solution to the problem: the solution gradually unfolds - it is developed.
I do agree that a "coder" creates code, and a programmer creates programs. I expect more of a complete program than of a bunch of code. If a text says "coder", it does set an expectation about the professionalism of the text. And I expect even more from a software solution created by a software engineer. At least a specification!
Still, I, a professional software engineer and programmer, also write "code" for throwaway scripts, or just for myself, or that never gets completed. Or for fun. I will read articles by and for coders too.
The word is a signal. It's neither good nor bad, but If that's not the signal the author wants to send, they should work on their communication.
You can't use a language that will be taken by everyone the same way. The public is heterogeneous - its subsets will use different "codes".
Also in my ears coder always sounded cooler than programmer and it wasn't until a few years ago i first heard that to some people it has negative connotations. Too late to change though, it still sounds cooler to me :-P.
Now: "code" is something you establish - as the content of the codex medium (see https://en.wikipedia.org/wiki/Codex for its history); from the field of law, a set of rules, exported in use to other domains since at least the mid XVI century in English.
"Program" is something you publish, with the implied content of a set of intentions ("first we play Bach then Mozart" - the use postdates "code"-as-"set of rules" by centuries).
"Develop" is something you unfold - good, but it does not imply "rules" or "[sequential] process" like the other two terms.
ChatGPT doesn't know either.
https://developer.nvidia.com/cuda-downloads?target_os=Linux&...
Your resource is really bad.
"We'll then load the trained GPT-2 model weights released by OpenAI into our implementation and generate some text."
What a bad take. That resource is awesome. Sure, it is about inference, not training, but why is that a bad thing?
The GPT from scratch post explains, from the ground up, ground being numpy, what calculations take place inside a GPT model.
Next-token prediction: https://github.com/bennyschmidt/next-token-prediction
Good for auto-complete, spellcheck, etc.
AI chatbot: https://github.com/bennyschmidt/llimo
Good for domain-specific conversational chat with instant responses that doesn't hallucinate.
However, currently there is some language-specific stuff in Transformer that should be moved to Language :) I'm focusing first on language models, and getting into image generation next.
Even in CSS a matrix "transform" is the same concept - the word "transform" is not unique to language models, more a reference to how 1 set of data becomes another by way of computation.
Same with tile engines / game dev. Say I wanted to rotate a map, this could be a simple 2D tic-tac-toe board or a 3D MMO tile map, anything in between:
Input
[
[0, 0, 1],
[0, 0, 0],
[0, 0, 0]
]Output
[
[0, 0, 0],
[0, 0, 0],
[0, 0, 1]
]The method that takes the input and gives that output is called a "transformer" because it is not looking up some rule that says where to put the new values, it's performing math on the data structure whose result determines the new values.
It's not unique to language models. If anything vector word embeddings are much later to this concept than math and game dev.
An example of use of word "Transformer" outside language models in JavaScript is Three.js' https://threejs.org/docs/#examples/en/controls/TransformCont...
I used Three.js to build https://www.playshadowvane.com/ - built the engine from scratch and recall working with vectors (e.g. THREE Vector3 for XYZ stuff) years before they were being popularized by LLMs.
If you'd please review https://news.ycombinator.com/newsguidelines.html and stick to the rules when posting here, we'd appreciate it.
If you'd please review https://news.ycombinator.com/newsguidelines.html and stick to the rules when posting here, we'd appreciate it.
> Unless I'm missing something.
Only that I said "without taking the LLM approach" meaning tokens aren't scored in high-dimensional vectors, just as far simpler JSON bigrams. I don't think that disqualifies using the term "transformer" - I didn't want to call it a "computer" or a "completer". Have a better word?
> JSON instead of vectors
I did experiment with a low-dimensional vector approach from scratch, you can paste this into your browser console: https://gist.github.com/bennyschmidt/ba79ba64faa5ba18334b4ae...
But the n-gram approach is better, I don't think vectors start to pull away on accuracy until they are capturing a lot more contextual information (where there is already a lot of context inferred from the structure of an n-gram).
The idea of tokenizing words and producing completions is not unique to the original transformers, it's a basic idea from NLP. So I'm not sure why you think it should be called a transformer just because it uses tokenized inputs and produces completions as well. It's like saying your new programming language has a "Java-based architecture" simply because they both have classes (and nothing else in common otherwise).
>I didn't want to call it a "computer" or a "completer". Have a better word?
I've seen projects which also use Markov chains + additional rules ontop, for example there's quite a few projects called "Markov chains with POS tagging":
https://github.com/26medias/context-aware-markov-chains
>not from lookups or assembling based on rules.
Not quite sure about "it's not based on rules" when your code has things like:
const MATCH_FIRST_MODAL = new RegExp(/IS|AM|ARE|WAS|HAS|HAVE|HAD|MUST|MAY|MIGHT|WERE|WILL|SHALL|CAN|COULD|WOULD|SHOULD|OUGHT|DOES|DID/);
or const properNoun = `${part.value} `;
if (isPrevNNP) {
result += prependArticle(query, properNoun);
}
Pretty sure your examples in the video are also cherry-picked. The very first example is you asking "where is Paris?" What really happens is, one of the hardcoded regexps transforms it to "Paris is" and then the bigram model repeats the second sentence in the Paris dataset verbatim.A CSS matrix "transform" is the same concept.
Same with tile engines & game dev. Say I wanted to rotate a map:
Input
[
[0, 0, 1],
[0, 0, 0],
[0, 0, 0]
]
Output[
[0, 0, 0],
[0, 0, 0],
[0, 0, 1]
]The function is a "transformer" because it is not looking up some rule that says where to put the new values, it's performing math on the data structure whose result determines the new values.
> Not quite sure about "it's not based on rules" when your code has things like: > > const MATCH_FIRST_MODAL
Totally irrelevant to the topic. This is the chat interface itself which mostly just parses questions into cursors to be completed. You would be a fool to think ChatGPT has no NLP or parts-of-speech analysis. text-ada-embedding itself uses POS.
> Pretty sure your examples in the video are also cherry-picked
Fantastic detective work, you caught me. But just to confirm - why not just use it yourself? npm i next-token-prediction
Here is an example you can run very easily in Chrome, so you don't have to rely solely on your amazing bullshit detector: https://github.com/bennyschmidt/next-token-prediction/tree/m...
Don't forget to log the completions to prove that they aren't broken down by token, and instead just doing key/val lookups or text searches as you said.
> What really happens is, one of the hardcoded regexps transforms it to "Paris is"
The only thing you got right - that questions are transformed into sentences using conventional NLP in order to complete them. This functionality is what makes it a chat bot that you can ask questions.
It's still misleading to call it a transformer in the context of NLP. It doesn't matter what it means in other, non-NLP areas (linear algebra, CSS or gamedev).
It's like creating a procedural language and calling it "functional" because it has functions. Sure the concept of functions existed long before compsci but it would be very misleading because "functional programming" is a well-established term.
>You would be a fool to think ChatGPT has no NLP or parts-of-speech analysis
Pretty sure it doesn't. At least it's not required to. I've run lots of local models and it's just model weights without hardcoded regexps. In fact, I was able to feed grammar rules of an invented language into Claude Sonnet and it was able to construct proper sentences.
>text-ada-embedding itself uses POS
Do you have a link?
> it's not required to. I've run lots of models
Then you must know about skip-gram and how embeddings are trained: https://medium.com/@corymaklin/word2vec-skip-gram-904775613b...
What is meant by "sliding window" or "skip gram" is bigram mapping (or other n-gram).
This is ML 101.
It's the same training methodology and data structure used in my next-token-prediction lib, and is widely used for training for LLMs. Ask your local AI to explain the basics, or see examples like: https://www.kaggle.com/code/hamishdickson/training-and-plott...
> ChatGPT doesn't use parts-of-speech
Yes it does, there's not only a huge business in tagging data (both POS and NER) adjacent to AI, but OpenAI specifically famously used African workers on very low wages to tag a bunch of data. ChatGPT uses text-embedding-ada, you'll have to put 2 and 2 together as they don't open source that part.
Mistral says:
"The preprocessing stage of Text-Embedding-ADA-002 involves applying POS tags to the input text using a separate POS tagger like Spacy or Stanford NLP. These POS tags can be useful for segmenting sentences into individual words or tokens."
> I use Claude to make new languages
Cool story, has nothing to do with the topic
I didn't say that it's unique to LMs. My argument is that saying "my LM is a transformer" is misleading because "transformer" in the context of LMs means a very specific architecture. You're deliberately misusing terms, probably to draw attention to your project.
>OpenAI specifically famously used African workers on very low wages to tag a bunch of data
Did they tag Polish parts of speech too? Or Ancient Greek? ChatGPT constructs grammatically correct Ancient Greek. I thought they tagged "harmful/non-harmful", not parts of speech?
>ChatGPT uses text-embedding-ada
[Citation needed]
NanoGPT, for example, learns embeddings together with the rest of the network so, as I said, manual tagging is not required.
Anyway, looking forward to hearing news about your image generation project. Any news?
The wikipedia on deep learning transformers:
All transformers have the same primary components:
- Tokenizers, which convert text into tokens.
- Embedding layer, which converts tokens and positions of the tokens into vector representations.
- Transformer layers, which carry out repeated transformations on the vector representations, extracting more and more linguistic information. These consist of alternating attention and feedforward layers. There are two major types of transformer layers: encoder layers and decoder layers, with further variants.
- Un-embedding layer, which converts the final vector representations back to a probability distribution over the tokens.
Where does it say bigrams can't be used for next-token prediction? Or that you can't tag data? Note "...which converts tokens and positions of the tokens..."> You're deliberately misusing terms, probably to draw attention to your project.
Haha well since I have like 30 followers and the npm is free/MIT whatever scheme you think I'm up to it's not working. Anyway a text autocomplete library is not exactly viral material. Jokes aside, no I am trying to use accurate terms that make sense for the project.
Could just make it anonymous - `export default () => {}` - and call the file `model.js`. What would you call it?
> Did they tag Polish parts of speech too? Or Ancient Greek?
Yes, all the foreign words with special characters were tokenized and trained on. An LLM doesn't "know any language". If it never trained on any Polish word sequences it would not be able to output very good Polish sequences anymore that it could output good JavaScript. It's not that has to train on Polish to translate Polish per se, but it does has to have the language coverage at the token level to be able to perform such vector transformations - which is probably most easily accomplished by training on Polish-specific data.
See https://huggingface.co/pranaydeeps/Ancient-Greek-BERT
> The model was initialised from AUEB NLP Group's Greek BERT and subsequently trained on monolingual data from the First1KGreek Project, Perseus Digital Library, PROIEL Treebank and Gorman's Treebank
First1KGreek Project
> The goal of this project is to collect at least one edition of every Greek work composed between Homer and 250CE
> Citation needed
https://openai.com/index/new-and-improved-embedding-model/
> The new model, text-embedding-ada-002, replaces five separate models for text search, text similarity, and code search, and outperforms our previous most capable model, Davinci, at most tasks, while being priced 99.8% lower.
https://platform.openai.com/docs/guides/embeddings/embedding...
Scroll to embedding models
> Anyway, looking forward to hearing news about your image generation project. Any news?
Not yet! Feel free to follow on GitHub or even help out if you're really interested in it. Would be cool to have pixel prediction as snappy as text autocomplete.
https://github.com/bennyschmidt/next-token-prediction
^ If you look at this GitHub repo, should be obvious it's a token prediction library - the video of the browser demo shown there clearly shows it being used with an <input /> to autocomplete text based on your domain-specific data. Is THAT a Markov chain, nothing more? What a strange question, the answer is an obvious "No" - it's a front-end library for predicting text and pixels (AKA tokens).
https://github.com/bennyschmidt/llimo
This project, which uses the aforementioned library is a chat bot. There's an added NLP layer that uses parts-of-speech analysis to transform your inputs into a cursor that is completed (AKA "answered"). See the video where I am chatting with the bot about Paris? Is that nothing more than a standard Markov chain? Nothing else going on? Again the answer is an obvious "No" it's a chat bot - what about the NLP work, or the chat interface, etc. makes you ask if it's nothing more than a standard [insert vague philosophical idea]?
To me, your question is like when people were asking if jQuery "is just a monad"? I don't understand the significance of the question - jQuery is a library for web development. Maybe there are some similarities to this philosophical concept "monad"? See: https://stackoverflow.com/questions/10496932/is-jquery-a-mon...
It's like saying "I looked at your website and have concluded it is nothing more than an Array."
If you google "Is ChatGPT just a glorified Markov chain?" you will amazingly get pages of results of people asking this question, just like "Is jQuery just a glorified monad?" as if to reduce something novel down to useless, mere philosophy that "we've had" for thousands of years. Imagine suggesting using a state management library in React to improve FE dev and getting the retort: "Isn't that just a state machine?" in a discounting manner, and imagine the rest of the team actually nodding their head in agreement like a scene in Idiocracy - welcome to Hacker News.
For smart people, the answer to any question like this is "No". Google is not a glorified Array. Bitcoin is not a glorified LinkedList. Language models are not glorified Markov chains. To even ask that is so reductionist and incorrect that any answer obfuscates what they actually are.
Here's a gist you can paste into your browser that shows how both n-grams and conventional NLP (parts-of-speech analysis) are used to derive vector embeddings in the first place: https://gist.github.com/bennyschmidt/ba79ba64faa5ba18334b4ae... (following in the style of text-embedding-ada-002 albeit much tinier)
They are not mutually exclusive concepts to begin with. Never have been. None of these comments even deserve these lengthy replies (I am likely responding to a mix of 12- and 24-year-olds who don't care that much anyway, just want to "win"), yet I feel compelled to explain.
edit: I do appreciate your work and explanation, btw.
Thanks.
So Markov chains