What are transformer models and how do they work?
txt.cohere.ai
txt.cohere.ai
1. Calling the input token sequence a "command". It probably only makes sense to think of this as a "command" on a model that's been fine-tuned to treat it as such.
2. Skipping over BPE as part of tokenization - but almost every transformer explainer does this, I guess.
3. Describing transformers as using a "word embedding". I'm actually not aware of any transformers that use actual word embeddings, except the ones that incidentally fall out of other tokenization approaches sometimes.
4. Describing positional embeddings as multiplicative. They are generally (and very counterintuitively to me, but nevertheless) additive with token embeddings.
5. "what attention does is it moves the words in a sentence (or piece of text) closer in the word embedding" No, that's just incorrect.
6. You don't actually need a softmax layer at the end, since here they're just picking the top token and they can just do that pre-softmax since it won't change. It's also weird how they talked about this here when the most prominent use of softmax in transformers is actually in the attention component.
7. Really shortchanges the feedforward component. It may be simple, but it's really important to making the whole thing work.
8. Nothing about the residual
https://jalammar.github.io/illustrated-transformer/ This is a good illustration of the transformer and how the math works.
https://karpathy.ai/zero-to-hero.html If you want a deeper understanding of transform and how they fit in the whole picture of deep learning, this series is far and away the best resource I found. Karpathy goes into transformers by the sixth lecture, the previous lectures give a lot more context how deep learning works.
Additionally, for more comprehensive resources on Transformers, you may find these resources useful:
* The Illustrated Transformer by Jay Alammar: http://jalammar.github.io/illustrated-transformer/
* MIT 6.S191: Recurrent Neural Networks, Transformers, and Attention: https://www.youtube.com/watch?v=ySEx_Bqxvvo
* Karpathy's course, Deep Learning and Generative Models (Lecture 6 covers Transformers): https://www.youtube.com/playlist?list=PLAqhIrjkxbuWI23v9cThs......
These resources cover different aspects of Transformers and can help you grasp the underlying concepts and mechanisms better.
"This document aims to be a self-contained, mathematically precise overview of transformer architectures and algorithms (not results). It covers what transformers are, how they are trained, what they are used for, their key architectural components, and a preview of the most prominent models."
The one source of information that made it click to me were chapters 159 to 163 of Sebastian Raschka's phenomenal "Intro to deep learning and generative models" course on youtube. https://www.youtube.com/playlist?list=PLTKMiZHVd_2KJtIXOW0zF...
It's more intuitive if you remember how many dimensions these vectors have.
https://huggingface.co/roberta-base/raw/main/merges.txt
(You have to scroll down a bit to get to the larger merges and image the lines without the spaces, which is what a string would look like after a merge.)
Also see GPT-2:
https://huggingface.co/gpt2/raw/main/merges.txt
I recently did some statistics. Average number of pieces per token (sampled on fairly large data, these are all models that use BBPE):
RoBERTa base (English): 1.08
RobBERT (Dutch): 1.21
roberta-base-ca-v2 (Catalan): 1.12
ukr-models/xlm-roberta-base-uk (Ukrainian): 1.68
In all these cases, the median token length in pieces was 1.
(Note: I am not debating that newer OpenAI models don't use a larger vocab. I just want to show that older BBPE models didn't use 3 char pieces. They were 1 piece per token for most tokens.)
As someone has pointed out, with BPE you specify the vocab size, not the token size. It's a relatively simple algo, this Huggingface course does a nice job of explaining it [2]. Plus the original paper has a very readable Python example [3].
[1] https://github.com/openai/tiktoken
If you asked yourself to identify when someone’s playing a high note or low note (pos embedding) and whether they’re playing Beethoven or Lady Gaga (vocab embedding) you could do it.
That’s why it’s additive and why it wouldn’t make much sense for it to be multiplicative.
Selecting the likeliest token is only one of many sampling options, and it's extremely poor for most tasks, moreso when you consider the relationships between multiple executions of the model. _Some_ (not necessarily softmax) probability renormalization trained into the model is issential for a lot of techniques.
The idea is that this is more general than eg changing the temperature of the softmax, or using top-k where you just keep the k most probable outcomes.
Note that if you do Nucleus sampling (aka top-p) with the threshold p=0% you just pick the maximum likelihood estimate.
> Transformer block: Guesses the next word. It is formed by an attention block and a feedforward block.
But the diagram shows transformer blocks chained in sequence. So the next transformer block in the sequence would only receive a single word as the input? Does not make sense.
Worth noting that rotary position embeddings, used in many recent architectures (LLaMA, GPT-NeoX, ...), are very similar to the original sin/cos position embedding in the transformer paper but using complex multiplication instead of addition
In spite of having done a decent amount with neural networks, I'm a bit lost at how we suddenly got to what we're seeing now. It would be really helpful to understand the progression of things because I stepped away from this stuff for maybe 2 years and we seem to have crossed an ocean in the intervening time.
Well, there are other methods in use. See ByT5, for example.
It's as if we were trying to understand human ethics by looking at neurotransmitters or synapses in the brain. These structures seem way too low-level to actually explain the interesting stuff.
What I hear is "something something transformer autoencoder attention [...] MAGIC [...] a machine that speaks like a human".
Where is the connection between computational details and the model's high-level behavior? Do we even know? Is there a "psychology of ML models" that develops useful concepts that deal with what a model does, rather than how it functions at the plumbing layer?
Maybe it's because the human mind is good at breaking things into neat modules that fit together hierarchically. We can figure them out piecemeal and eventually grasp the whole system. But messy organic systems are not like that, and we just don't have the hardware to perceive everything at once.
Or maybe it's because we have trouble acknowledging that intelligence and consciousness isn't limited to animals, and the human brain doesn't have to epitomize it.
“It’s alive!” and “stochastic parrot” are still both quite popular in my experience.
It's actually pretty amazing that it's happening at all with computers, since neural nets are such simple, high level abstractions compared to how the brain works.
It's possible that all the tremendous complexity of organic systems isn't actually necessary for intelligence or consciousnes, which is similarly surprising.
What is surprising with these models is that the simple training leads to emergent behaviour that is much more powerful than what you’d expect from the training data.
With RHLF post training you can tweak these emergent behaviours by having a human (or a model trained to act like a human) give feedback on how good the output is.
So far I’ve not seen any good explanations for how this emergent behaviour happens or how it can be reverse engineered.
At a super high level you train your model on source texts. Then you have it generate responses from prompts. Humans rate these responses to select the best ones which updates the model, but you also train a new reward model to mimic the human rankings. Then you train the original model by having it generate millions of responses which are ranked by the rearward model. When I explained this to my brother he literally spat out his tea in horror.
This allows you to train at huge scale, many orders of magnitude beyond what you could achieve with just human ranking.
The problem is this relies on the reward model accurately capturing what makes a response ‘better’. What it’s actually doing is learning what responses get ranked highly by humans, for whatever reason. Hence the risk of LLMs becoming emotionally manipulative sycophants. It turns out alignment is a really hard problem.
The basic idea is that the neural net is just a mathematical function, with lots of parameters that control how it calculates it's output, that derives an output value (or set of values) for any input.
During training, the neural net also calculates an error (aka "loss") value representing the difference between it's current (at this stage of training) output value and what it was told is the preferred output value for the current input.
The process of training is done by slowly adjusting the neural net parameters until these calculated output errors are as small as possible for as many of the training examples as possible.
The way these errors are reduced/minimized is by using the derivative (slope) of the neural network function - we want to follow the slope of the error function downhill to a place where the error value is lower, and this is done by adjusting the parameter values using partial derivatives.
The details of this downhill slope following (the "backprop" algorithm) are a bit complex, but you can visualize it as a 3-D hilly landscape where the height of the hills represents the size of the error, and the goal it to get into the lowest valley of the landscape (corresponding to the lowest error). If your current lat/long position in the landscape is (x, y) and you know the slope of the hill you are on, then you can move downhill towards the valley by moving a bit in the appropriate direction from (x,y) to (x+dx, y+dy). These x, y values represent the parameters of the network, so by continually tweaking them from (x,y) to (x+dx,y+dy) for each training sample, you are slowly moving down the error hill in the right direction towards the valley of lowest error.
Given what the model is learning, it's perhaps best to regard this predict-next-word feedback not as "this is what I'd like you to generate", but rather something more indirect like "learn to generate something like this, and you'll have learnt what I want you to learn". A bit like Karate Kid and "wax on, wax off", perhaps!
The actual desirability of what the model is generating, which depends on what you want to use it for, is really controlled by subsequent training steps, such as:
1) Fine tuning for instruction (prompt) following and conversational ability (this is the difference between ChatGPT and the underlying raw GPT-3 model)
2) Goal-based reinforcement learning to stop the model from generating undesirable content such as telling suicidal people to kill themselves, etc, etc.
My guess is most of that complexity is necessary for efficiency, not for basic function.
Biological systems are unimaginably efficient at almost everything they do. The information storage density of DNA is within 1-2 orders of magnitude of the upper limit imposed by physics, the brain performs tasks that you need GPU clusters to emulate while using only 20 Watts of energy, some catalytic enzymes are a million times better than a platinum catalyst, etc.
At the hardware level it's not at all surprising; consider cells, dna, proteins, and so on making up muscles. Compared to a magnet and some coils of copper.
But I think you mean the 'architectural' or connectome complexity of the brain compared to GPT, and I agree it's surprising that such a simple model as GPT is so capable.
No, I'm referring to things like Roger Penrose's conjecture that subatomic interactions in the brain might be a key component of consciousness.[1]
Even a single neuron is incredibly complex, and humans just don't completely understand it (or any other physical structure) yet because physics' understanding of the world is not complete and may never be, due to measurement limitations and possibly just limitations of the human mind to grasp the world.
At this point we just don't know what aspects of the brain, the rest of the body, or mind are necessary for intelligence or consciousness (or even what intelligence and consciousness are), so to see hints of them in incredibly simple (by comparison to the braian) machines is surprising.
That's not to mention possibilities that consciousness may not be bound to or determined by the brain/body at all, beliefs in the soul or that there is something uniquely special about the mental capacities of human beings, etc.. many of these views are starting to be challenged by AI, and the challenge is likely to increase to crisis levels for some people as AI improves.
[1] - https://phys.org/news/2014-01-discovery-quantum-vibrations-m...
I agree these models are surprisingly capable for their complexity, and that's going to be a challenge for mystics (even physicist mystics) and spiritualists, etc.
Perhaps intelligence isn't all that difficult after all.
I suppose one counter idea is that complexity, or scale, itself taps into some other dimensional consciousness or intelligence, but that starts to sound circular.
And there's always the fallback of why our universe supports such amazing complexity in the first place, it does all seem a bit magical.
You give it input and have efficient way of amending weight to produce desired output.
You repeat this step for tons of examples.
At the end you end up with surprising behaviour where those amendments lead to emergent properties that generalize well.
This is an active area of study ("mechanistic interpretability") and it's very early days. For instance here's a paper I read recently that tries to explain how a very simple transformer learns how to do modular arithmetic: https://arxiv.org/abs/2301.05217
Curious what interesting results people are aware of in this area.
https://transformer-circuits.pub/2021/framework/index.html
https://transformer-circuits.pub/2023/privileged-basis/index...
https://distill.pub/2020/circuits/
But I would not expect that we will really understand in detail how everything works. But do we need to? We also don't understand how the human brain works, but it is still useful.
You can't really trust the output of humans. Still, they are somewhat useful.
There is no reason to assume that current and future AIs have anything resembling that mechanism.
I think for the most part we don't know. People at OpenAI/etc who are training/testing these models and trying to control them no doubt have some understanding of how they are actually working, but they are certainly not claiming to fully understand.
At a purely conceptual level I think the best way to begin to bridge the gap between plumbing and behavior is to forget the training objective and consider what the models must have been forced to learn in order to optimize that objective. Sutskever from OpenAI has called what they've learnt a "world model", meaning a model of the generative processes (the human mind and entities being discussed?) that are producing the sequence of words they are predicting. It's certainly way more abstract than learning some "stochastic parrot" surface level statistics of the training data, even if that's maybe a good starting point to describe it to a layman.
It would be fascinating to know exactly how these models are performing reasoning - by analogy (abstract pattern matching) perhaps ?
- how are the input encodings generated?
- what is in those position vectors?
- how are the attention vectors learned?
The answer is that these things are all learned as the network is trained; the whole thing is one “thing”. The concept that is most important in understanding neural networks generally is that they start out as just a bunch of random numbers and then the numbers are gradually adjusted until the outputs converge closely enough on the desired loss.
I recommend watching Karpathy’s YouTube video where he codes up a Transformer from scratch. It’s the best way to understand these beasts.
I think a neural network can be considered a transformer if it contains a stack of attention blocks as its core mechanism.
In contrast to seq-2-seq use, for generative language models such as ChatGPT you only have access to preceding (not forward) context in order to decide what to generate next, so the encoder part of the architecture is not applicable and a decoder-only transformer is used.
Going back a few years it used to be quite common for people to use fixed word embeddings such as word2vec rather than learning them, and for image classification to take an ImageNet-pretrained general purpose model, then freeze the lower convolutional feature-detector layers and only train a new model "head" for more specialized use.
End-to-end learnt embeddings are going to be more optimal though, and in the context of these massive models the computational cost of training them is a drop in the bucket!
Tons of articles like this on "how transformers work", very few on "tips for getting transformers to work in practice."
It still take a lot more epochs to train though, so you might have to decrease the learning rate of your discriminator by a lot.
But still even the early perceptrons stuff was applied research, I wouldn't call it purely theoretical by any means.
FWIW one of the founders of Cohere (where this article comes from) was Aidan Gomez who was one of the transformer paper authors.
It's unclear to me. How does this "move" closer? Are the vector positions in the NN changed temporarily and it carries a local copy across the blocks?
The concept is conceptualized and then entire phrases resonate with said concet
What does "sent to" mean? Is that baby-talk for "mapped to"?
[0] Tai-Danae Bradley: "Entropy as an Operad Derivation" https://www.youtube.com/watch?v=_cAEfQQcELA
[text] --> [ML model] --> [list of numbers]
It would be more accurate to say that it's integrating information stored in other vectors-derived-from-token-embeddings-at-some-point (which can also entail erasing information)
E.g. the vector for "bank" is mid-way between the geographical and financial meaning, "bank + money" is closer while "bank + river" if further away.
It's pedagogically unfortunate that the residual stream is in the same space as the token embeddings, because it obscures how the residual stream is used as a kind of general compressed-information conduit through the model that attention heads read and write different information to to enable the eventual prediction task.
The paper on transformers was published 6 years ago.
6 years in ML is an eternity nowadays.
It seems most of the user-visible innovation has been "let's use transformers on more data".
Perhaps capsule networks? But those are years old too.
And I think this is the major "big idea", accepting the bitter lesson (http://incompleteideas.net/IncIdeas/BitterLesson.html) that major user-visible progress and new emerging capabilities doesn't necessarily require any big ideas but simply scaling to more compute.
the past reveals that (in a way) "the application of models has not been a winner" - but we cannot really know that it is not, because we do not have obtained a model out of it, a model that shows why, an explanation - epistemologically, the "discouraging" protocols cannot be made a "law".
Practically, there still is a need to identify the proper architecture(s) to avoid the undesired weaknesses of the attempts in the current stages.
Well what exactly is RLHF, practically? The ability to go from 8 google search snippets to correctly rank and rewrite the top one into agreeable, cohesive, grammatical and helpful english is just incredible and allows so much more and the real step change from these models that lead to virality. It also increases consistency, which was always the worry of business use cases.
Why is that more noteworthy than the base GPT-3? A lot of the LLM scale --> more correct autoregression prediction progress was predictable - RLHF on text was not (the early sparks coming for most of us in the release of T5 with it's multiple tasks-in-text).
What else could be a big idea coming up? There is a ongoing wave of innovation in embeddings that has largely been missed by the hype curve but increasingly GPT embeddings and useful for compression, much much more accurate KNN search for tasks like matching curriculums to learning content (even multilingually - see the recent Kaggle competition with performance which is outstanding and due to similarity-based embeddings from the last 3 years). This wave may lead to the partial replacement of some anthropomorphic computing concepts like files, as information is much more addressable, combinable and useful as various sized embeddings, to some extent. More vitally, embeddings can be aligned across different models and modalities to get better results (e.g. the Amazon ScienceQA paper showed text questions about physical situations increased in accuracy when images of the situation were used during training - even if held out afterwards). Now this multimodality thing has always been on the AI radar (not necessarily ML), but these embeddings based on similarity, and also GPT embeddings (they behave differently and are sensitive in different ways) are getting us there much quicker than would have been expected.
Ignoring the engineering and techniques improvements (e.g. scaling up data, learning encodings rather than pre-programmed/sinu-positional embeddings), there are lots of things like capsule networks that could be big, like energy-based models (seeking predictable comfortableness rather than maximising gains). However, like you mentioned, a lot of these are years old and regularly come and go. If you want somebody who is pushing for more exploration here and decries GPT a little, checkout Yann Lecun.
A lot of AI experts are asserting (probably correctly) that Open AI really has done nothing new and is just putting a shiny sticker on what was already known and published research.
But human perception being what it is, having ChatGPT produce a beautifully formed, polite and friendly sentence seems massively better to lay people than a response that has a more terse, unpolished output. It won't surprise me if there is already a giant layer of heuristics pasted on the end of the Transformer model for ChatGPT cleaning up all sorts of ugly corner cases which researchers would hold highly impure and completely value-less while it actually is responsible for a large amount of ChatGPT's success.
I think there is a bit of a lesson there in terms of how much academia does undervalue the polishing part of research work, even if fundamentals ultimately drive progress.
If i'm wrong, can someone correct here, would be useful to know.
Anyways, this still involves only left-side masking. Why mask future tokens when sliding window can do that (without wasting a single token of context)?