Llama3 implemented from scratch
github.com
github.com
Like a roadmap, do A, do B And finally go through this in the end.
words i understand but don't really get: weights, feed forward, layers, tensors, embeddings, normalization, transformers, attention, positioning, vector
There's "programming" in the plumbing sense where you move data around through files/sockets and then there's this... somebody without a math background/education... very unlikely you'll understand it. it's just skimming python and not understand the math/library calls it makes
It helps me to think in terms of levels of abstraction rather than complexity. My education stopped at a 4 year degree, but AI is mostly postgraduate still. So I have to translate to what I know because I haven't internalized the lingo.
Here's the most approachable teaching of neural nets (NNs) and large language models (LLMs) that I've seen so far:
https://news.ycombinator.com/item?id=40213292 (Alice’s Adventures in a differentiable wonderland)
https://arxiv.org/pdf/2404.17625 (pdf)
https://news.ycombinator.com/item?id=40215592 (tensor and NN layer breadcrumbs)
II A strange land 105
7 Convolutional layers 107
..
7.1.3 Translational equivariant layers 112
..
9 Scaling up the models 143
..
9.3 Dropout and normalization 151
9.3.1 Regularization via dropout 152
9.3.2 Batch (and layer) normalization 156
III Down the rabbit-hole 167
10 Transformer models 169
10.1 Introduction 169
10.1.1 Handling long-range and sparse dependencies 170
10.1.2 The attention layer 172
10.1.3 Multi-head attention 174
10.2 Positional embeddings 177
10.2.1 Permutation equivariance of the MHA layer 177
10.2.2 Absolute positional embeddings 179
10.2.3 Relative positional embeddings 182
10.3 Building the transformer model 182
10.3.1 The transformer block and model 182
10.3.2 Class tokens and register tokens 184
11 Transformers in practice 187
11.1 Encoder-decoder transformers 187
11.1.1 Causal multi-head attention 188
11.1.2 Cross-attention 189
11.1.3 The complete encoder-decoder transformer 190
11.2 Computational considerations 191
11.2.1 Time complexity and linear-time transformers 191
11.2.2 Memory complexity and the online softmax 192
11.2.3 The KV cache 194
11.2.4 Transformers for images and audio 194
11.3 Variants of the transformer block 197That's exactly where I am at. Despite watching Karpathy's tutorial videos, I quickly got lost. My highest level of math education is Calculus 3 which I barely passed. This probably means that I will only ever understand LLMs at a high level.
Dive into Deep Learning is another.
Both have free PDF versions available.
The math isn't difficult. The notation is a little foreign, and you have to take your time reading and rereading the equations.
The only downside is that in 2024, you are probably going to use PyTorch and not Keras + Tensorflow as shown in the book.
What is the name of the course? This? https://ocw.mit.edu/courses/res-6-007-signals-and-systems-sp...
Google and find the examples where someone does it in a spreadsheet. It's much more approachable that way.
You are going to find it's not that complicated.
This was posted on HN a while ago and led to some great discussion. Myself and others agreed that this type of stateful visualization was _way_ more effective at conceptualizing how an LLM works than reading code or stepping through a debugger.
Of course, this is not all there is to a modern LLM, it would probably take another thousand lines or two to implement training, and many more than that to make it fast on all the major CPU and GPU architectures. If you want a flexible framework that lets a developer define any model you want and still goes as fast as it can, the complexity spirals.
Most programmers have an intuition that duplicating a large software project from scratch, like Linux or Chromium for example, would require incredible amounts of expertise, manpower and time. It's not something that a small team can achieve in a few months. You're limited by talent, not hardware.
LLMs are very different. THe code isn't that complicated, you could probably implement training and inference for a single model architecture, from scratch, on a single kind of GPU, with reasonable performance, as an individual with a background in programming and who still remembers their calculus and linear algebra, with a year or so of self study. What makes LLMs difficult is getting access to all the hardware to train them, getting the data, and being able to preprocess that data.
https://arstechnica.com/information-technology/2024/03/once-...
I suspect GPT-4's "secret sauce" in terms of edging out competitors is that OpenAI is better about managing data contractors than the other folks. Of course it's a haze of NDAs to learn specifics, and clearly the contractors are severely underpaid compared to OpenAI employees/executives. But a lone genius with a platinum credit card can't create a new world-class LLM without help from others.
… built on the back of a disposable workforce…
There is something grim and dystopian, thinking about the countless small hands feeding the machine.
Dystopian indeed, this is pretty much how Manhattan Project and CERN were done, with many independent contractors doing different parts, and only a few has the overview. A page out of corporate management book, it very much allows concentration of power in the hands of a few.
https://twitter.com/emollick/status/1786213463456448900
$22 in 2008 -> $33 today https://data.bls.gov/cgi-bin/cpicalc.pl?cost1=22&year1=20080...
Comparing LLMs to the Manhattan Project based on budget alone is stupid and arrogant. The comparison only "makes itself" because Ethan Mollick is a childish and unscientific person.
Just want to clarify. The comparison to Manhattan Project or CERN is referencing "the countless small hands feeding the machine." In projects such as these, roles and jobs are divided into small parts, that people who are working on it don't really see the forest from the tree, and that only a few that has the picture of the whole project.
I am wondering if "CERN was pushed on the masses by the few" is an oblique reference to public fears that the LHC would destroy the world.
Yes, that's my opinion too. GAOs (Grassroots AI Organisations) are constrained by access to data and the hardware needed to process the data and train the model on it. I look forward to a future where GAOs will crowdsource their computations in the same way many science labs borrow computing power from people around the world.
But only for the same reasons. Linux runs on very nearly every piece of hardware ever made. The APIs you have to implement in order to run "Linux programs" are large and full of old complexity that exists for compatibility. Chromium is full of code to try to make pages render even though they were designed for Internet Explorer 6.
Conversely, some university programs have students create a basic operating system from scratch. It's definitely something a small team can do as long as you don't care about broad hardware support or compatibility with existing applications. In principle a basic web browser is even simpler.
It actually teaches you how to build llama iteratively, test, debug and interpret the training loss rather than just desribing the code.
I have implemented inference of Whisper https://github.com/Const-me/Whisper and Mistral https://github.com/Const-me/Cgml/tree/master/Mistral/Mistral... models on all GPUs which support Direct3D 11.0 API. The performance is IMO very reasonable.
A year might be required when the only input is the research articles. In practice, we also have reference Python implementations of these models. Possible to test different functions or compute shaders against the corresponding pieces from the reference implementations, by comparing saved output tensors between the reference and the newly built implementation. Due to that simple trick, I think I have spent less than 1 month part-time for each of these two projects.
Great overview. One gap I've been working on (daily) since October is the math working towards MA's Mathematics for Machine Learning course (https://mathacademy.com/courses/mathematics-for-machine-lear...).
I wrote about my progress (http://gmays.com/math) if anyone else is interested in a similar path. I recently crossed 200 days of doing math daily (at least a lesson a day). It's definitely taking longer than I want, but I also have limited time (young kids + startup + investing).
The 'year of self study' definitely depends on where you're starting from and how much time you have, but it's very doable if you can dedicate an hour or two a day.
This is an indication that we’re at the infancy of this field.
"the code for AGI will be simple" - John Deremetrius Carmack
I'm kind of shocked. I thought there would be more dynamism by now and I stopped dabbling in like 2018.
There's still some room for experimenting if you care about memory/power efficiency, like MoE models, but they're not as well understood yet.
There seems to be no rhyme or reason, no scientific insight, no analysis. They just try a million different permutations, and whatever scores the highest on the benchmarks gets published.
E.g. "In-context Learning and Induction Heads" is an excellent paper.
Another paper ("ROME") https://arxiv.org/abs/2202.05262 formulates hypothesis over how these models store information, and provide experimental evidence.
The thing is, a 3-layer MLP is basically an associative memory + a bit of compute. People understand that if you stack enough of them you can compute or memorize pretty much anything.
Attention provides information routing. Again, that is pretty well-understood.
The rest is basically finding an optimal trade-off. These trade-off are based on insights based on experimental data.
So this architecture is not so much accidental as it is general.
Specific representations used by MLPs are poorly understood, but there's definitely a progress on understanding them from first principles by building specialized models.
This particular (tock) is still playing out. The next (tick) does not feel imminent and will likely depend on when we discover the limits of the transformers when it comes to solving for long tail of use-cases.
My $0.02.
There are certainly tradeoffs to both, the general transformer motif scales very well on a number of axis, so that may be the dominant algorithm for a while to come, though almost certainly it will change and evolve as time goes along (and who knows? something else may come along as well <3 :')))) ).
I think it's dominance is not going to substantially change any time soon. Dont you know, the solution to all leetcode interviews is a hash table?
Furthermore, I think a replacement will require that we _understand_ what the current crop of models are doing mechanically. Some of it was motivated in [1].
[1] https://openaipublic.blob.core.windows.net/neuron-explainer/...
Personally, I think going linear instead of quadratic for a core operation that a system needs to do is by definition an optimization.
My bet will be on something else than gradient descent and backprop but really I don't wish any company or country to reach agi or any sophisticated ai ...
If a 100x improvement in performance is left on the table, then surely even lower priority optimizations won't be implemented any time soon.
Consider this: a lot of clever attention optimizations rely on some initial pass to narrow the important tokens down and discarding them from the KV cache. If this was actually possible, then how come the first few layers of the LLM don't already do this numerically to focus their attention? Here is the shocker: they already do, but since you're passing the full 8k context to the next layer anyway, you're wasting it on mostly... Nothing.
I repeat: Does the 80th layer really need the ability to perform attention over all the previous 8k outputs of the 79th layer? The first layer? Definitely. The last? No. What happens if you only perform attention over 10% of the outputs of layer 79? What speedup does this give you?
Notice how the model has already learned the most optimal attention scheme. You just need to give it less stuff to do and it will get faster automatically.
Step change, then optimization of that step change
Kind of like a grand father clock with a huge pendulum swinging to one side, then another(commonly used metaphor).
So until the hardware allows for comparable (say with 2-4x) thoroughput of samples per second I expect model architecture to mostly be static for most effective models and dynamic architectures to be an interesting side area.
I wonder shouldn't AI be the best tool to optimize itself?
Serious question: assuming this is true, if an incumbent-challenger like OpenAI wants to win, how do they effectively compete against current services such as Meta and Google product offerings which can be AI enhanced in a snap?
Their task now is to maintain and exploit those advantages as best they can while they build up a more stable long term moat: lots of companies having their tech deeply integrated into their operations.
Really? Most of our testing now has Gemini Pro on par or better (though we haven't tested omni/Ultra)
It really seems like the major models have all topped out / are comparable
gpt, claude, gemini, even llama and mistral, all tend to produce the same nauseating slop, easily-recognizable by anyone familiar with LLMs - these days, I cringe when I read 'It is important to remember' even when I see it in some ancient, pre-slop writings.
creativity - one of the very few applications generative AI can truly excel at - is currently impossible. it could revolutionize entertainment, but it isn't allowed to. the models are only allowed to produce inoffensive, positivity-biased, sterile slop that no human being finds attractive.
What's really funny is they all have "jailbreaks" that you can use to make then say anything anyway. So for "corporate" uses, the method you propose is already mandatory. The whole thing (censoring base models) is a misguided combination of ideology and (over the top) risk aversion.
My written Chinese is limited 一二三 and that from Mahjong tiles, and I keep getting 四 and 五 mixed up.
https://www.google.com/search?q=gemini+german+soldier
prompt-injected mandatory diversity has led to the most hilarious shit I've seen generative AI do so far.
but, yes, of course, other instances of 'I reject your reality and substitute my own' - like depicting medieval Europe to be as diverse, vibrant and culturally enriched as American inner cities - those are doubleplusgood.
* What exactly are the current ones doing that makes them generate 'black Vikings'?
* How would you change it so that it doesn't do that but will also generate things that aren't only representative of the statistical majority results of large amount of training data it used?
* Would you be happy if every model output just represented 'the majority opinion' it has gained from its training data?
* Or, if you don't want it to always represented whatever the majority opinion at the time it was trained was, how do you account for that?
* How would your method be different from how it is currently done except for your reflecting your own biases instead of those you don't like?
There is presumably a system prompt or similar that mandates diverse representation and is included even when inappropriate to the context.
> How would you change it so that it doesn't do that but will also generate things that aren't only representative of the statistical majority results of large amount of training data it used?
Allow the user to put it into the prompt as appropriate.
> Would you be happy if every model output just represented 'the majority opinion' it has gained from its training data?
There is no "majority opinion" without context. The context is the prompt. Have you tried using these things? You can give it two prompts where the words are nominally synonyms for each other and the results will be very different, because those words are more often present in different contexts. If you want a particular context, you use the words that create that context, and the image reflects the difference.
> How would your method be different from how it is currently done except for your reflecting your own biases instead of those you don't like?
It's chosen by the user based on the context instead of the corporation as an imposed universal constant.
The models you encounter are going to be fine tuned, where they take the base and train it again on question and answer sets and chat conversations and also have a layer of 'alignment' where they have sets of questions like 'q: how do I be a giant meanie to nice people who don't deserve it' and answers 'a: you shouldn't do that because nice people don't deserve to be treated mean' etc. This is the layer that is the most difficult to get right because you need to have it but anything you choose is going to bias it in some way just by nature of the fact that everyone is biased. If we go forward in history or to a different place in the world we will find radically different viewpoints than we hold now, because most of them are cultural and arbitrary.
Wait, why do you need to have it? You could just have a model that will answer the question the user asks without being paternalistic or moralizing. This is often useful for entirely legitimate reasons, e.g. if you're writing fiction then the villains are going to behave badly and they're supposed to.
This is why people so hate the concept of "alignment" -- aligned with what? The premise is claimed to be something like the interests of humanity and then it immediately devolves into the political biases of the masterminds. And the latter is worse than nothing.
Imagine if search engines adopted this same sort of moral totalitarian mindset and if you happened to search for the 'wrong' thing, the engine would instead start offering you a patronizing and blathering lecture, and refuse to search. And 'wrong' in this case would be an ever-encroaching window on anything that happened to run contrary to the biases of the small handful of people engaged, on a directorial level, with developing said search engines.
Your leap to "thou shalt not search this" is missing the possible middle ground
I have no idea how to make it happen, but the talk about biases, safeguards, etc should be made between many different people and not just within a private company.
I saw this issue working at Tinder too. One day they announced how they will be removing ethnicity filters at the height of the BLM movement across all the apps to weed out racists. Nevermind that many ethnical minorities prefer or even insist on dating within their own ethnicity and this was most likely hurting them and not racists.
That really pissed me off and opened my eyes to how much power these corporations have over dictating culture, not just toward their own cultural biasis but that of money.
without those '''safeguards''' implemented to appease the aforementioned 0.01%, things could be very different - some big models, particularly Claude, can be tard wrangled into producing decent prose, if you prefill the prompt with a few thousand token jailbreak. my own attempts to get various LLMs to assist in writing videogame dialogue only made me angry and bitter - big models often give me refusals on the very first attempt to prompt them, spotting some wrongthink in the context I provide for the dialogue, despite the only adult themes present being mild, not particularly graphic violence that nobody except 0.01% neo-puritan extremits would really bat an eye at. and even if the model can be jailbroken, still, the output is slop.
Have you played around with base models? If you haven't yet, I'm sure you'll be happy to find that most base models are delightfully unslopped and uncensored.
I highly recommend trying a base model like davinci-002[1] in OpenAI's "legacy" Completions API playground. That's probably the most accessible, but if you're technically inclined, you can pair a base model like Llama3-70B[2] with an interface like Mikupad[3] and do some brilliant creative writing. Llama3 models can be run locally with something like Ollama[4], or if you don't have the compute for it, via an LLM-as-a-service platform like OpenRouter[5].
[1] https://platform.openai.com/docs/models/gpt-base
[2] https://huggingface.co/meta-llama/Meta-Llama-3-70B
[3] https://github.com/lmg-anon/mikupad
> Further, in developing these models, we took great care to optimize helpfulness and safety.
The model you linked to isn't a base model (those are rarely if ever made available to the general public nowadays), it is already fine-tuned at least for instruction following, and most likely what some in this game would call 'censored'. That isn't to say there couldn't be made 'uncensored' models based on this in the future, by doing, you guessed it, moar fine-tuning.
Does grok do this, given where it came out of?
And no, Twitter is no excuse to type like an illiterate teenager.
And I will bet you someone edits his blogs to not look like that.
Do you think using capitals at the beginning of a sentence aids comprehension?
I view punctuation and spelling rules as a way to maximize comprehension (akin to having a linting standard). In non formal writing, I don't see any harm in avoiding capitalization (at least it doesn't seem to me to help understanding / reading speed, etc at all).
One would expect Altman to know how to use the SHIFT key when running a massive business, but, hey - once you achieve escape velocity from society, you don't have to live by its norms or grammar rules.
I can assure you that it would cost most people people here a promotion or a raise if they did this at work.
Then, I guess, he decided, he was too important to follow grammar rules.
I personally don't really mind that bit of capitalization that English does. German is much worse.
You misspelled 'better'.
And they are not alone.
> d' 'ou 'xp'ct h'br'w sp''k'rs t' wr't' 'n 'ngl'sh l'k' th's?
But the Greeks added vowels to the alphabet because Indo-European languages rely a lot on vowels (as opposed to Semitic languages which are easy to understand without vowels).
A just question.
such as your comment and my comment!
Also it looks more casual and authentic, less LLM generated
The post looks informative I hope to learn something from it later tonight. Thx
I mean wtf is this. https://kubernetes.io/blog/2024/04/17/kubernetes-v1-30-relea...
Anime/waifu shit, furries and all becoming commonly accepted as of late? 10-15 years ago you'd be exiled. Now it seems like it's whatever
Why? I don't know. Video games may be a common denominator. Also, Japan was really big into tech in the 90s, and they still are to a lesser extent.
It's also very much offtopic since it generates repetitive thread-gobbling tangents, like this one is threatening to. Mentioned in the site docs a couple of different ways:
Please don't pick the most provocative thing in an article or post to complain about in the thread. Find something interesting to respond to instead.
Please don't complain about tangential annoyances—e.g. article or website formats, name collisions, or back-button breakage. They're too common to be interesting.
I'm starting to understand that there is a much deeper social conflict going on with whatever is happening on this topic, so I got my answer and I"m just going to move on.
It has become so bad that moderators will not ban these people even if they explicitly try to justify molesting children. Some of them are moderators themselves. And even have calls to genocide in their bio. This is most prevalent in the ArchLinux community. Specifically, their Telegram channels.
Is there something particularly different about this one?
Edit - guess not?
You might as well look at llama.cpp for a serious and production grade implementation to learn from. Otherwise, nothing to see here.
> Is there something particularly different about this one?
Other than the immature lowercase, anime BS, etc, then…
No.
Anyway, I'll take a look at this too, not sure if it has inference and training. Having just inference would be a disappointment.