Attention Is Off By One
evanmiller.org
evanmiller.org
The author is suggesting that we add 1 to the denominator of the softmax that is used within attention mechanisms (not the final output softmax).
The softmax inside an attention unit allows it to see key/query matches as probabilities; those probabilities support a continuous-valued version of a key-value lookup (instead of 1/0 output of a lookup, we get weights where a high weight = the desired key-value lookup).
Adding 1 to the denominator would change an attention unit by no longer working with a true probability vector of weights, but rather working with weights that add up to less than 1. The motivation is that the network can learn to provide high weights so that the adjusted softmax is very close to a probability vector; and it has a new option to provide all-low weights which give all-low output weights, meaning it can opt out of having high confidence in anything.
(switching to opinion mode)
2. How can we tell if this is good?
2a. We should just try it out: Train an LLM with this, see if it works.
2b. There are two reasons I suspect it won't make a big difference.
First, if an attention node has low confidence, it can already assign similar scores pre-softmax. Then we get what looks like a uniform distribution as output. Then we're basically taking an average of a bunch of vectors (vs a weighted average that is more like choosing one of them). Statistically, we expect that averaged vector to be close to zero. In other words, the node already has a way to effectively opt-out by providing a near-zero output vector.
Second, in a transformer, each attention unit has many other learned weights that can support the ability to opt out. Both the V matrix and the feed-forward layer after the attention unit give that module a way to provide low values to the activation function after the feed-forward layer, which would result in a value as small as you like — again, a way to opt out.
3. I appreciate the non-academic tone of the article and the willingness to play around with fundamental ideas. Although I'm not totally convinced by the note, I'd love to read more stuff like this.
Disagree here, I think neural nets are quite bad at implicitly learning low entropy transforms, similar to how they struggle to model the identity function, necessitating residual connections. In both cases the change doesn't increase expressivity, but it does bake these needle-in-a-haystack transformations into the model that may be hard to access with gradient descent.
Can't speak to how useful it is though.
The equations all seem to be matrix operations with a fixed number of rows / columns (you can take me as a real layman here). Unless you change that, I don't understand _how_ you can reduce memory needs. Granted, I'm probably putting my foot in my mouth not understanding transformers.
During quantization we find that values in the network vary from 0->5000, but 95% of values are <100. Quantizing this to 8bits would mean that our values would be in increments of about 20. Remembering that 95% of our values are below 100, we would only have about 5 discrete values for 95% of our values - so we would be losing a lot of "resolution" (entropy/information). For example (assuming rounding is used), an original value of 19 would be quantized to 20 and 30 would be quantized to 40. The original values differ by 11, but the quantized values differ by 20!
This is where exotic encodings come into play. We might try to use a logarithmic scheme, for example. This would result in higher value densities at lower values - but we would probably still waste bits and it would require more APU cycles.
Now switch to the softmax1 network:
The range of values is less important than the distribution - instead of 95% of the values falling in a small range, we would see the values more evenly spread out. Assuming that the range is now 105 (so the 5% outlying neurons from the softmax network are still >100), we would have 243 values to represent everything under 100. The same example with 19 and 30 would result in 19.27 and 30.34 respectively, a difference of 11.07 - which is very close to the unquantized difference of 11. We have retained more information in the quantized version of the network.
Information is lost either way, but what's important is how much information is lost.
The reason that the large values appear is because the heads attempt to "scream really loud" when they are certain that they are right. This is an emergent behavior due to softmax - it ironically sucks at paying attention to a few of the heads: it boosts the volume of the heads that are trying to abstain, and mutes the volume of the heads that are trying to vote.
Instead of using an 8bit integer with even step size quantification, wouldn't they still use an 8bit float?
Either way you would still only have 256 discrete values.
Example: multiplying a bunch of float16s together gives you a float16. That is passed on to the next layer of float16s. Why should forcing the output of the first step to be float8 confer any advantage here? The only way I can see this argument working is if you make all the layers float8 too, and the reason you can do that is that the output of the first step can be faithfully represented as float8 because it doesn't ever blow up. If that's what the author is saying, it wasn't very clear.
This is striking. If true, why not try to ignore whitespace and puctuation?
In old Latin, scripto continua [1] was a way to write continuously, for the exact same reason: to save space. Other modern languages still do that, and are no less parseable.
Granted, it's unlikely a commercial LLM would become popular if it produced output without spaces or punctuation; but an open source one that promised to be much more compressible, and therefore work on smaller machines, might be super useful.
It's not hard for a human to add spaces afterwards. It used to be a job for beginning journalists at the time of telex machines: press releases were sent in all caps without spaces, and interns were tasked with adding slashes between words. In French it was called "bâtonner les dépêches" (literally: add sticks to press releases -- not sure about the idiomatic English translation).
That's not to say a pipeline couldn't be effective.
It is initially, but thinking about it some more, there's a lot of information packed in whitespace and punctuation choice.
Scripto continua may have worked because the few readers who lived back then expected it to encode some form of legal or religious prose, but even then they could learn things from the overall shape of the document. LLMs are working in a much richer domain of document types, but the only thing they can "see" is a stream of tokens. There's no spatial or geometric data attached there. So whitespace and punctuation are the only thing an LLM has to make inferences about otherwise textually identical inputs. Such as:
(see: other) -- vs -- {see: other}
One being likely a text fragment, the other likely a piece of code.Or how spacing may imply Markdown or YAML being used. Or how it may imply a list. Or a poem. Or a song. Or specific writing style, such as "lol im a casual who not care bout comms" vs. "I am a distinguished professor, about to retire. Elites like us put two spaces after full stop."
The Latin literature was extremely rich, from Cicero to Tacitus, and was certainly not limited to legal information.
Here's part of your comment with white space and punctuation stripped:
scriptocontinuamayhaveworkedbecausethefewreaderswholivedbackthenexpectedittoencodesomeformoflegalorreligiousprosebuteventhentheycouldlearnthingsfromtheoverallshapeofthedocumentllmsareworkinginamuchricherdomainofdocumenttypesbuttheonlythingtheycanseeisastreamoftokenstheresnospatialorgeometricdataattachedtheresowhitespaceandpunctuationaretheonlythinganllmhastomakeinferencesaboutotherwisetextuallyidenticalinputs
It's a little hard to read, but not that hard. I think one would get used to it.
Also, for creative use of LLM, it may be a feature, as trying to find the words could be inspiring.
I think it would be worth a try.
string.replace(/[\s\.\*<>\!\?,;:\-–\|"'\[\]\(\)]/g, '')This post's author has a different solution, and one that theoretically could avoid causing large outliers that prevent efficient quantization. These large outliers seem to be an unfortunate side-effect of the models' learned solution.
So getting rid of spaces would do nothing to solve the problem, and would instead force the models to learn a new solution, one that presumably isn't as optimal.
> Originally I wanted to call this function ghostmax, as you can think of there being an extra zero-valued entry in x (as exp(0)=1), as well as a zero vector in the V matrix that attenuates the result.
Don't think of this as weighting the options so that some of the time none of them is chosen. ("Weights that add up to less than 1.") Instead, think of this as forcing the consideration of the option "do nothing" whenever any set of options is otherwise considered. It's the difference between "when all you have is a hammer, everything looks like a nail [and gets hammered]" and "when all you have is a hammer, nails get hammered and non-nails get ignored".
I like this framing because, as an example, it bothers me that our speech-to-text systems use this method:
1. A human predetermines what language the input will use.
2. Audio in that language is fed to transcribing software.
3. You get, with modern technology, a pretty decent transcription.
3(a). ...if the audio sample was really in the language chosen in step 1.
If you ignore the choice of language and feed French audio to an English transcriber, you get gibberish. This is wildly at odds with how humans do transcription, where absolutely the first thing that a system that only knows how to transcribe English will do, when given French audio, is object "hey, this is definitely not English".
a) train two identical models on a large dataset, one with the +1 in the denominator for the softmax steps of the attention modules, one without
b) show that they have similar performance (doubt the +1 will make performance better, but we need to show it doesn't make things worse)
c) show that there are less "blowups" in the model with +1, and therefore they are more effectively quantized.
So you might want to re-search the hyper params for a fair shot.
Yes but how much would this cost?
Would it be possible to build a small dataset that produces known outlier values, and test on that?
One of the most important keys to the success of deep learning in the last couple years has been the fact that emergent features exist after certain scales, so I wouldn't be too quick to dismiss things that don't help at smaller scales, nor would I be certain that all the tricks that help in small data/parameter regimes will necessarily help in larger models. Unfortunately!
[1] https://timdettmers.com/2022/08/17/llm-int8-and-emergent-fea...
So while you obviously wouldn't be able to conclusively prove the idea fixes the issue in larger models, if you know what you are looking for you should be able to validate that the method works in general down to very small models.
That said, consumer grade cards should be able to train an 8B model with quantization, so you might as well train the whole thing.
I'm not sure that's the case, especially in high dimensions.
The expected value of the absolute value n random variables, uniform [-1,1], grows with n. I'm pretty sure it's proportional to the sqrt of n.
Also, random walks in high dimension return to zero with probability zero, so the sum of random variables in high dimensions going close to zero seems unlikely as well.
Mathematically, we can write v_out = V * w,
where v_out is the vector of output from the attention unit, w is the probability vector from the softmax, and V is the set of input vectors, where each column is an input vector.
For a moment, pretend that the columns of V are orthonormal to each other. This might not be true, but it's an interesting case.
When the model wants the output to be small, it can set w = 1/n, meaning all coordinates of vector w are 1/n. (n = the number of columns in V)
In that case, the length ||v_out|| will be 1/sqrt(n) exactly, which is small compared to the input lengths of 1 (since we're pretending they were orthonormal).
Now if we stop pretending they are orthonormal, the worst case is that they're all the same vector, in which case the weights w can't change anything. But that's a mighty weird case, and in high dimensions, if you have any randomness at all to a set of vectors, they tend to point in wildly different directions with dot products close to zero, in which case the same intuition for the orthonormal case applies, and we'd expect a uniform distribution coming out of the softmax to give us a vector that's much smaller than any of the input vectors.
In particular, I think the feedforward point you list in your "Second" is actually wrong. Replacing a softmax with 0, as the OP wants to do, is tantamount to passing the information unchanged, because the attention block is within a residual (skip) connection. If it's set to zero, the next output is identical to the previous layer output. There is no way to recover this effect with the feedforward layer.
The part that you can set V to zero is true, but somehow a different idea: the Q and K should be able to set to 0 if no token wants to be "close" to some other token, in some sense. But the V layer shouldn't "know" about this, because it can't look at other tokens. This is of course only how we think of transformers, which might or might not (more likely, the latter) be how it actually works. But nevertheless, having a 0 value coming out of the K.Q^T part only would be very meaningful.
Your "first" point is technically true (albeit logically false): if you have a sequence of length 32k, like GPT4-32k, and your softmax logits all predict the same value, the result will be an average of the V layer, divided by 32k, which is effectively close to zero. However, calibrating "exactly the same value" is extremely hard for a neural network, and there is no "default value" it can predict to make sure that's the case - even if you push all the values to one side, the result doesn't change, because softmax is translation invariant. Plus, if you have a short sentence, that's not true anymore. If you only have two tokens, one of them must be activated, or both with only a 0.5 factor. Surely if you have very few tokens there's much more contamination between Q, K, and V, so in that case V can indeed take a 0 value, but it's non-trivial and requires more layers.
All in all, adding that "+1" isn't quite meaningless, I think. Nevertheless, I believe it won't change much: these very big models have ways to get around any kind of smart small modification you do. If the intuition is very right, it might be that you can squeeze 1% out more accuracy in a handful of tests, after you carefully optimize all other parameters, which would be enough to get you a paper in a top conference. And it might also be implemented as a standard from them on (because, in this case, it basically doesn't cost any more computations, so it's "free"). But I would bet it won't be a major revolution.
That said, as you say, the only way to know would be to train a few models with this option and check the actual quality of them (certainly not GPT-style, nor GPT4-size, models, to begin with, but something quicker to train and easier to test in a fully automated way; old "boring" models like those in the BERT family would be a good point to start testing). But to do that effectively, you'd need somebody skilled in training this kind of models, with the cleaned data ready at hand, etc. (and a small compute budget, of course, but nothing revolutionary, a few thousand $ in GPU credits could be enough)
OP calling it a "bug that's been overlooked for 8+ years" is click bait.
Could anyone kindly point me to this as I can't find it.
[1]: https://pytorch.org/docs/stable/generated/torch.nn.Multihead... [2]: https://github.com/google/flaxformer/blob/main/flaxformer/co...
For example, see this snippet from an old Google repo: https://github.com/google/flaxformer/blob/ee62754ebe5a5eeb11...
Technically softmax is not implemented as presented but through exp(x_i-max(x)), and summing over it in the denom. But maybe I am missing something.
Furthermore, the residuals are used exactly because the networks cant learn the identity function; but they can learn zero; at which point the residual is `f(x): x+g(x)` with being `g:x ~> 0` (ie approximately 0).
It is also the case that `f(x): x+g(x)` makes it easier for gradients to flow through.
Regardless of numerical stability tricks (e.g. exp(x_i-max(x))), you are still simply normalizing the logits such that the probabilities sum to 1.
The blog adds an additional hidden logit (equal to 0) to allow for softmax(x) = 0 when x -> -inf.
Enough weights don't fall under that "nearly" that we require more bits per weight to cover those edge cases. If we were able to delete the "nearly" we would need fewer bits (smaller models).
I certainly don't think it will help at all with stability. Things like Q/K layernorm are better tricks for softmax stability when scaling: https://arxiv.org/pdf/2302.05442.pdf
How would you have known if the trick actually reduces the outliers in the weights? Even if the transformer quality does not improve overall, having less outliers as a result is very beneficial for more accurate quantization of the data
The "how" is pretty straightforward.
>> because no one has yet looked at whether the trick helps reducing outliers in very large models
Given a softmax version doing exactly as the blog post says is baked into a google library (see this thread), and you can set it as a parameter in a pytorch model (see this thread), this claim seems off. "Let's try X, oh, X doesn't do much, let's not write a paper about it" is extremely common for many X.
Agree - the "how" is straightforward
An example here with an actual algorithm, although it's been a couple of years so my explanation might be a bit wrong in places. and/or i might have gotten the completely wrong end of the stick with the current thread.
--
The CTC (Connectionist Temporal Classification [0]) algorithm maps a sequence x with length X -> sequence y with length Y.
i.e. in speech to text we might have some audio features that correspond to the following class predictions (post softmax classification)
x -> hellllloooooooooo wwwooorrrllld
we want to get this as the output y -> hello world
we have the alphabet as classes we try to predict for each sequence item in x.we could just removed all the duplicate in the first long sequence, but we would end up with `helo world` ... we need to preserve one of the early `l` characters in `hello` somehow
CTC uses a blank token (aka dummy) token to handle potentially deliberately repeated items in sequence x.
By adding the blank token to the classes predictions, we can get the model to predict something like this (post softmax classification)
y* -> hel~l~~oooo~~~~~~ w~~o~~r~~l~~d
The CTC decoder (non-ML decoding algo) heuristically removes repeated tokens. Turning the above into ... y -> hello world
... the duplicate `o` and `~` characters are removed.It was a decent enough algorithm for speech-to-text prior to attention/transformers etc.
However, it makes CTC vulnerable to well designed adversarial example attacks because there is a massive bias within models to predict the blank token -- meaning it's very easy to modify input sequence x to switch the output sequence y to include blank tokens for nefarious purposes (the subject of my unfinished phd).
[0]: www.cs.toronto.edu/~graves/preprint.pdf
This is a great solution. Though that's a dummy token in the output rather than the input. I guess you could do something inverse to do text to speech, but it might be hard to say where to insert the dummy tokens in that case.
Adding 1 to the denominator can be useful if you have softmax with just a few options. Not in self-attention where you have thousands.
Except with the +1 denominator, it might be that the model trains all of the inputs to become very negative so softmax chucks out close to zeros, whereas it wouldn't bother before because making one prob bigger makes another smaller.
It still can't do this because of L2 regularization / weight decay. If two vectors are norm 1, their inner product is at least -1, so with 2000 vectors that's still 2000 * e^(-1) =~ 735.
Not saying it's theoretically impossible that it could happen. But you would have to try _really_ hard to make it happen.
Of course it might have some other serious repercussions.
Doesn't describe the implications even briefly. If they add just your second sentence to that description, it'll immediately become so much more useful.
In 2011, I wanted to copy the reddit ranking algorithm in a project of my own, so I went to source code to look at it... the algorithm in the source code I found wasn't doing anything at all sensible with negative-sum voted posts.
I thought I discovered the error, some terms swapped in the simple equation, the sign for positive/negative was misapplied.
I blogged it, and [posted it to reddit](https://www.reddit.com/r/programming/comments/td4tz/reddits_...), only to have MANY people, including reddit employees, tell me I am definitely definitely wrong, and the algorithm was working as intended. And that I was in fact not the first to notice what I thought I noticed, and point it out, and be told by everyone I was wrong.
OK, I didn't really understand what was going on, I couldn't make sense of the algorithm if it wasn't wrong, but so be it. I updated my blog post to say that people smarter than me said there was no error in the reddit algorithm, all I can say is this variation makes more sense to me.
Then, three years later in 2014, a commit was made to the reddit source code with exactly the correction I (and others before me) had suggested all along. The one that everyone piled on to tell me how dare I have the temerity to suggest reddit source code is wrong.
https://github.com/reddit-archive/reddit/commit/50d35de04b92...
¯\_(ツ)_/¯
Open source means there are lots of eyes that can find bugs, but sometimes they can't convince anyone they've found a bug. (And of course, then reddit close-sourced their code in 2017).
I never did end up using the ranking feature in my own project, that I had wanted to copy from reddit. I didn't end adding "vote" features to the app.
You can make a long, impactful career by just being "the guy who adds log statements throughout the codebase and reasons through it", doing this at even a simplistic level has always shown me an astonishing fix to some long-standing issue.
n.b. It also attracts a ton of political fun. People's first order reaction is denial, and it only gets worse from there. Absolutely no one except 1-2 colleagues will see it as "oh we should fix that", and at least one person will make sure your boss' boss' boss is CCd on an email with a nice version of "no he's just insufficiently concerned about {concurrency, memory management, take your pick}" Just wait it out quietly when that happens, do not engage or complain. If nothing happens and you're never asked about it by leadership, but your peers ask, make plans to move onto another team.
I've been at big tech companies for most of my career and I've never seen anyone deny the existence of a technical bug. I've seen plenty of teams mark a bug as lower priority and never fix it because other things are higher priority. But denying that the bug exists, especially after a detailed explanation? That doesn't resonate with my experiences.
It used to be writing the outputs from the C/C++ preprocessor (.i files) to disk took forever (5+ minutes IIRC) with Microsoft's compilers. I asked one of the lead compiler developers why, and he waved me away saying it was just really complicated. Around that time a bunch of tools existed for GCC that worked with .i files, but none existed in the Microsoft ecosystem likely because writing .i files was so slow.
I was on the compiler test team at the time and we did lots of stuff with .i files, our tests were distributed across a large cluster of test machines (see my post about that https://meanderingthoughts.hashnode.dev/how-microsoft-tested...) so it wasn't a big deal, but it still annoyed me.
One day I decided to find out what was going on, so I loaded up process monitor while outputting a .i file and watched what was happening. Much to my surprise, only 1 byte was being written at a time! No wonder writes were taking forever.
A quick dive into the source code revealed a comment above the file write call that read to the effect
// to work around a bug in windows 98
So anyway I opened a bug against the compiler saying we should probably fix that. :)
Of course the lead developer waved you off. You wondered why things took forever, and the lead developer knew it was a complicated system and figured it wasn't worth their time investigating. It happened to be incorrect, but the lead developer wasn't in denial. They just filtered the issue out because they can't afford to go down every rabbit-hole they come across. I'm sure once you found the actual bug, it was later fixed.
The person I was responding to seems to think a large number of people are in denial when a bug is filed against them. That doesn't make sense, and isn't something I see. It'd be as if when you pointed out the actual bug, the lead developer continued to say it wasn't actually a bug (which is of course ridiculous and I bet didn't happen).
I know it's not so confusing as to get that sort of interpretation, because of the score on the comment, and comments like the above that explain to you how this happens.
As a result, I don't feel comfortable providing more detail publicly about my situation. That far off the mark tends to indicate an aggressive rather than curious interlocutor.
I am comfortable building on their example. The particulars of the issue are quite similar in a very helpful way.
I did the investigation, did a fix, worked it up to my manager and my managers manager. Elated, we work diligently for a couple weeks to document concisely, 3 page tech doc, briefest code deltas possible, one page + slides withs simple diagrams.
It gets bogged down at managers managers coleads submanager for the platform team implicated. They basically say "reading the single byte at a time means its provably serial and thus has no concurrency bugs.", as indicated in my original comment.
I did learn an important lesson that more senior people are not always right, and as someone who's usually more senior than my colleagues now I try to remember it daily.
If it weren’t for the torturous gaslighting, this is borderline hilarious. Appeal-to-authority types have a way of submitting so effortlessly when a grander poobah comes around. Spine made of jelly.
And the response was just like "We disagree, we think it makes sense the way it is and the product is correct"
That's kind of the end of the argument, there's nothing more one can say!
It didn't help that I came in assuming that of course everyone would see which version was correct (as you just did! although i didn't find it obvious, it took me lots of study to figure out), instead of producing a narrative designed to gently persuade them that. (That's on me -- I think I've learned something about technical communications around bugs and disagreements since then, although I'm still far from perfect).
The real answer, I think was given in one of the reddit thread comments -- the way it's broken for the most part _doesn't matter_ in the usual operations of reddit, it matters only in edge cases, and not very important ones, so really people mostly don't notice and we don't care.
Fair enough, I guess? But they did fix it three years later? I forget how I even found out they had fixed it; I can't at this point find any context for _why_ they fixed it, or who with power finally noticed/agreed it could use fixing why.
(And if it had happened three years after that, it would not have been in public source, and my gloating satisfaction would have been stolen!)
* Outlier values are used to prune values. * Transformers seem to undergo a "phase shift" in how outlier features are treated around 6.7B parameters. This could complicate research on removing them.
Maybe you and Tim Dettmers would have a lot to talk about :)
If the author is reading any of these comments, though, I would urge them to expand on their claim that “I’m 99.44% sure that it will resolve the outlier feedback loop”. As it stands, that’s the only explanation we get of how the outliers might be related to softmax!
And because the effects of the problem are subtle. Supposing the diagnosis is correct, full-precision LLMs still avoid the issue through large attention weights given to meaningless tokens to give harmless attention outputs. The problem only matters when quantizing weights, and quantized performance isn't really the goal of recent cutting-edge LLM development.
For such a simple change to the softmax it wouldn't take long to verify. It's really embarrassing to not do that before publishing.
I think there might be some curse of the auto-didact here, hinging on the meaning of publish: it would be embarrassing if he was capital-P publishing, as in a scientific paper.
The blog goes to great lengths to point out it is _not_ capital-P publishing.
To validate the idea the author has, it would be required to train a LLM from zero. If the author is right, you would get similar results to the current generation of LLMs, but with (a lot) less space required for the intermediate layers.
The time to achieve that is still measured in kilo- to mega-dollars, why is it wrong to put that idea in the open to substantially criticize or adopt?
And yes I do disregard his research effort. There are hundreds of well-justified and well-researched "clever tricks" for improving Transformers, and almost all of them don't work. I'll believe it when I see the results.
Train a Transformer based model with and without the modified Softmax (Suggestions: GPT-2 or nanoGPT)
Measure performance - I'd probably start with Perplexity and see if there is any difference (we'd expect little difference).
Quantize both models with different quantization strategies.
Measure the perplexity of the quantized models of different sizes. We'd expect the performance to drop off quicker for the non-modified model than the modified one if this is working.
In any case, that was an lmgtfy-level question. Here's what I found: https://til.simonwillison.net/llms/training-nanogpt-on-my-bl...
I shall try that soon.
I did a writeup like this. (Not as nicely as Simon though) where I modal.com (cloud GPU, containers, quick starts, free $30/m spend) to use their GPUs (e.g. T4, A100).
https://martincapodici.com/2023/07/15/no-local-gpu-no-proble...
T4 I think was good enough for the job, not much need for the A100.
Since this post I am working on an easy way to do this with a script called lob.py that requires no code changes to the nanoGPT repo (or whatever repo you are using) and runs in modal.com. The script exists but gets refined as I use it. Once it is battle tested a bit more I will do a post.
(It is named lob.py as it "lobs the code over to the server" where lob is UK slang for throw)
Watch this space.
BERT 109M, testing perplexity
OPT 125M, testing perplexity
ViT 22M, testing on ImageNet top-1.
I doubt that is true. Softmax is extremely well understood within the ML community. It's a very common trick, these properties are well-known as well. It feels very unlikely that nobody has thought of this before. That said, it's also plausible that the current softmax convention was chosen by accident and the author is right to identify this drawback.
So it turns out someone did. Specifically google did. This exact same idea has been in flaxformers since at least November 2021.
https://github.com/google/flaxformer/blame/ee62754ebe5a5eeb1...
Specifically to save people a click it says:
> """Softmax function with an additional virtual logit equal to zero.
For compatibility with some previously trained models.
This is equivalent to adding one to the denominator.
In the context of attention, it allows you to attend to nothing.
And creates the exact same modified softmax as this essay. I suppose only time will tell why it was ignored publicly before, maybe it doesn't do much, maybe it just fell through the cracks, maybe google just didnt push it, who knowsMaybe quantization wasn't as hot back then than it is now?
The post could have probably gotten the point across in less than 1/4 of the overall length (probably even less than 1/8th), instead the author wrapped the the post into lots of informalisms and a thinly veiled complained about academic publishing.
The result of this is reflected in the discussion here, nobody actually writes about the result/idea behind the post, instead we have ~200 comments discussing the merits of academic publishing vs blog posts and formal vs informal writing.
So I guess if you want to get your blog-post on the front page of HN it's a good writing style. If you want someone to consider and discuss the merits of your idea, maybe not so much.
This plants the seed for the info explosion (those 200 bikeshedding comments or those 6 billion videos on how to boil an egg).
To counter it we have rankings of comments and links and news feeds from google to fb to hn. But its just another layer of bullshit cause most of the pool of what is being ranked is bullshit.
We are yet to design Information systems that take into account what Goldhaber said about Attention 3-4 decades ago.
The point may be to entertain as well as inform. Many humans enjoy the unfocused discussion around the main point, and perhaps the author prefers it to the clinical and formal tone an academic paper tends to take.
edit to add details in case anyone is interested
I didn't add one to the softmax denom. I added a learned parameter (the attention sink) that would be appended to the beginning of QK but would be removed after softmax, so when multiplying by V the totals wouldn't sum to one. I tried variants that included looking at the current pos and not, and also variants that predicted used an ffn to generate the sink per position instead of a learned param. In my setting neither approach really made much of a difference. But I also had a bunch of other weird stuff in there too, so it may be worth trying again.
Open to being wrong here, but wouldn't it be functionally similar to adding a constant to the softmax denom? the function could sort of learn a specific position to have sink and q multiply to one, then removing it before multipling with v would be exactly identical?
Now it’s possible that softmax should be replaced wholesale, but it’s worked pretty well for the most part, except for this one wee little bug that prevents attention heads from saying nothing. So I propose a very small tweak on which I am willing to stake all future Internet claims to being correct. The tweak is so small, yet so obvious, and it’s been sitting here under everyone’s noses ever since attention was invented (2014).
I didn't test for outliers, but I don't think this will lead to a large improvement in attention overall/it will fix a lurking bug.
The problem with using softmax is that it forces each attention head to make an annotation, even if it has no information to add to the output vector. Using softmax to choose among discrete alternatives is great; using it for optional annotation (i.e. as input into addition) is, like, not cool, man.
I'm saying it uses the current position to do this, that if it was a significant error I would expect it to improve the training loss. I sort of interpreted the blog post as being a bit more positive on the idea than just being about improving the quantization
Then you don't know if the approach he is advocating actually improves what he is aiming for
I am, however, of the similar opinion that there could be better attention formulations. A paper from 2020 https://arxiv.org/abs/2005.09561 helped a lot in one of the transformers model I trained (not a vanilla LM but a specialised multi-modal graph problem).
It proposes normalised attention which if I'm not wrong should help with the quantisation problem too.
I’m surprised by the pompousness in the OP. Especially about something that most people who do transformer research understand. I’m also surprised that so many in the replies are taking the position of “this is what research should look like” when this is clearly an example of why research doesn’t work like this. Peer review is good for many things and one of those things is saving yourself some embarrassment.
You are reading some of the more ambiguous self-deprecation as genuine claims.
TL;DR on why this is important and he's sharing: it's a sort of niche thing that really only matters if you're trying to run pale imitations of ChatGPT on constrained hardware. That's why it's entirely possible the big guns didn't see it as important, they're not trying to run LLMs on a 3090
He’s writing in a colloquial and self-deprecating and humorous tone. I can’t speak to the merits, but I can follow the reasoning perfectly fine. It’d be hard to find something further from pompous.
> saving yourself some embarrassment
Implying of course that being wrong, or not the first one to discover this, is embarrassing. And that’s not pompous?
The motivation follows from the same problem the author points out in the original softmax formulation that it always "forces a choice" when it may be more useful to put a "Not Applicable" option into the model itself.
https://link.springer.com/article/10.1007/s10260-021-00578-2
The author mentions that he would maybe have written this as a scientific paper:
> I tried writing a serious-looking research paper about the bug and my proposed fix, but I lost a series of pitched battles against Pytorch and biblatex, so I figured I’d just write a blog post instead. (History is written by the winners; blogs are written by…)
Honestly, thank god he didn't. This paper is so much more readable and approachable than what gets published in "serious" journals. The tone is self-effacing, it does not have an "ego" the way scientific papers tend to have. If all science read like this, and if we were "allowed" to cite research that reads like this, I think we would be much better off. This reads like a conversational, approachable textbook, not like an impenetrable wall.
Is it because I don't understand attention at a PhD level that I hold this opinion? Maybe. Could he be writing like this because he's a layman and utterly wrong about the topic, unlike those Serious Science Authors? Maybe, I don't know.
But my god, wouldn't it be nice to be allowed to write like this?
But this is still scientific communication. It's really nice that it's legible!
> Even though softmax1 is facially quite boring, I’m 99.44% sure that it will resolve the outlier feedback loop that’s making quantization the subject of cascades of research. If you want to run some experiments and prove me right, DM me on Twitter and we’ll get a paper going.
I'm guessing that in the stodgy world of science, a communication like this might happen over lunch at a conference, limited to a small clique of researchers who are zealously guarding their next paper. Who could blame them, publish or perish!
But someone will probably test this theory out (after my read, it will probably happen in llama.cpp with preliminary results on GPT-2 by next week) and achieve results, and it will happen quickly and legibly to the outside world, because this was published openly and without all of the pretension that formal science (tm) has. If it works, it works. Stuff like this is the soul of the internet. Sharing knowledge and making it legible for all.
Workshop submissions often don't need evidence. They just need a small kernel to spur discussion.
Without experiments, there is no hope of publishing this in anything more than a workshop. Nor should there be.
Any reason to believe this? The author never mentioned it, and I can’t think of any other a priori reason why it should be true.
https://arxiv.org/pdf/2208.07339.pdf
Outliers appear at model size 6.7B and are not present at 2.7B
We call him an "independent AI researcher" because his google scholar is "bland" compared to many academics who play the academia game - https://scholar.google.com/citations?user=yk1QMowAAAAJ&hl=en
It's not a paper. It's an idea that sounds plausible, presented in a highly entertaining form.
What should be encouraged is for academics to blog about their research as well. It would even help when recruiting and onboarding new members. Right now the sociological and economical incentives don't promote this at all.
Equations are perfectly clear. I was able to follow his reasoning perfectly well.
I cannot say the same for so many papers (tm) that I've read. Mostly in a similarly computational (though non- deeplearning) applied math domain.
>What should be encouraged is for academics to blog about their research as well. It would even help when recruiting and onboarding new members. Right now the sociological and economical incentives don't promote this at all.
I will add onto this that a lot of journals have been pushing for video abstracts and "plain English" abstracts. For the most part I don't see these too often but when they're there they're appreciated, and I vaguely recall that someone found that citations go up when they're used (specifically plain English, I don't think anything has been on video abstracts).
There are a lot of good blogs for computational academic subjects (ml, bioinformatics, comp neuro, etc) but I see less for bio and non-software engineering. Math and physics seems to have some really notable blogs, but beyond what gets posted to HN and linked further on those blogs, I can't comment.
There was this sociologist who had written a paper for us all to read ahead of time. I started to read the damn thing, and my eyes were coming out: I couldn’t make head nor tail of it! I figured it was because I hadn’t read any of the books on the list. I had this uneasy feeling of “I’m not adequate,” until finally I said to myself “I’m gonna stop, and read one sentence slowly so I can figure out what the hell it means.”
So I stopped-at random-and read the next sentence very carefully. I can’t remember it precisely, but it was very close to this: “The individual member of the social community often receives his information via visual, symbolic channels.” I went back and forth over it, and translated. You know what it means? “People read.”
Then I went over the next sentence, and realised that I could translate that one also. Then it became a kind of empty business: “Sometimes people read; sometimes people listen to the radio,” and so on, but written in such a fancy way that I couldn’t understand it at first, and when I finally deciphered it, there was nothing to it.
-- Feynman
I disagree. After going through quite a few research papers in my time, I've found the best are the ones that are direct and to the point. Many papers I've spent many hours/days trying to unravel just to realize the concepts were straightforward, not very novel, and there wasn't much of real substance to the paper.Meanwhile, some of the most impactful papers I've read are direct and to the point. Kadmellia, Bitcoin, BitTorrent, DynamoDB, Firecracker, etc.
It seems like, when you have something of substance to say, you say it. When you don't you overcompensate by falling back on building an intricate puzzle of jargon and convoluted equations in an attempt to make what you're saying sound far more important than it really is.
As LLMs get better, I look forward to the day where every journal has a standard LLM filter you're required to apply to your paper that unravels all of this nonsense and rewrites it a more straightforward way, if not to directly publish than just for the editors to verify there isn't a simpler way to convey your ideas. I suspect that if we had an EIL5 filter for most journal articles, we'd discover that a majority of the words that get published have very little substance at all.
In cryptography, certainly a paper with formal definitions and proofs can be much more valuable than a corresponding blog post. It's a field where formalism is desired, if not necessary. Otherwise you can't check other people's "proofs", or even know what model you're working in.
I think, since people haven't come up with better formalisms, sometimes it's quite obtuse, which gets mistaken as "academic writing", when really it's a best effort to formalize.
Certainly some papers are better written than others. But sometimes a blog post cannot replace a paper, unless it also goes into the depth and detail that formalism requires. (Then it becomes a 30 page blog post, where most people don't read past the intro.)
You can have both and weave them together into a digestible narrative. I see Physics textbooks sometimes written this way.
If you just wanted a digestible intro then you would usually buy a textbook.
I think the argument that every research paper ought to be a mashup of a textbook + the actual research to be a bit silly from a “people should specialize at what they’re good at” standpoint.
Put in another context, I also don’t want every recipe to reintroduce what it means to “fry” or “braise” or “marinate”. We have Google for that.
And this blog post probably could be condensed into 1/4 of its size or less with a less conversational/bloggy tone.
And there are words that are added to add empty padding, keep up academic pretenses, and appear smart.
The post could have been condensed, but it would lose the former, not the latter.
In short, jargon matters. People here can talk about functional, procedural, and object-oriented programming because each of the three words has more than just the dictionary meaning - to those of use in the field. In the same way we can talk about linear algebra and know it doesn't mean "algebra on lines".
Yes, it's possible to write scientifically without jargon and wordiness, but it's a lot of effort and takes much more space to say "a group who follow a social structure within a society (culture, norms, values, status). They may work together to organise social life within a particular place, or they may be bound by a sense of belonging sustained across time and space"[1]
1 https://othersociologist.com/2013/11/20/sociology-of-communi...
Besides, in the example Faynman gives the simple sentence is actually shorter. Maybe that shorter sentence loses some information that the jargon carried, but Occam's razor suggests the writer was just trying to sound smarter.
A lot of research involves lumping and splitting: what underlying properties do these seemingly-different share (or vice versa). For example, reading text is just one possible instantiation of a “visual symbolic channel.” Traffic lights, road signs, gauges and dials, logos, and clocks also carry information the same way. If you want to discuss “reading and reading-like activities”, you may want some kind of umbrella term.
Plus, you may want to contrast them with other ways of sharing information: non-symbolic systems that literally depict the item in question (photos on a picture menu, for example) or using a different sense altogether, like church bells for telling time.
Expressions representing numbers may be combined with an expression representing a primitive procedure (such as + or *) to form a compound expression that represents the application of the procedure to those numbers.
And an English professor haughtily responding, "you know what that means? 'Computers compute!' This SICP book is just a pile of jargon that could be dramatically simplified!"
His dismissal revealed nothing about the topic, but a whole lot about how so many in the "hard" sciences view others. Don't understand the text? It's the text's fault! For I am a real scientist, and if I don't understand it, it's not understandable!
He might have been a genius, but he should have stuck to subatomic particles and left exploring human behavior up to the people who'd done the prerequisite reading.
The criticism was """Haraway's work has been criticized for being "methodologically vague"[39] and using noticeably opaque language that is "sometimes concealing in an apparently deliberate way""""
So you're saying that "Her work is basically handwaving and bullshitting".
"Michel Foucault’s biopolitics is a faccid premonition of cyborg politics, a very open feld. By the late twentieth century, our time, a mythic time, we are all chimeras, theorized and fabricated hybrids of machine and organism—in short, cyborgs. The cyborg is our ontology; it gives us our politics. The cyborg is a condensed image of both imagination and material reality, the two joined centers structuring any possibility of historical transformation. In the traditions of “Western” science and politics—the tradition of racist, male-dominant capitalism; the tradition of progress; the tradition of the appropriation of nature as resource for the productions of culture; the tradition of reproduction of the self from the refections of the other—the relation between organism and machine has been a border war"
(donna was woke before woke was a thing)
Donna Haraway was born 6 years after “stay woke” in its sense as an admonition to maintain alertness to the racist context was coined. Leaving aside a debate over whether her work is a good match for “woke”, she very much cannot have been woke before woke was a thing. (Before its recent replacement of “politically correct” as the American Right’s preferred, meaning-stripped, label for everything it disagrees with, sure, but “woke” was a thing long before that.)
A game of being pedantic is always welcome:
She very well could have been "woke before woke was a thing", because "woke" as the parent means it in her case, refers to the modern usage (of like, 2 decades), not the original term of the 40s that might have preceeded her birth.
So take the parent's comment to mean:
"She was woke, in the modern, circa-2000s+ sense, before woke, in the modern circa-2000s+ sense was a thing, not in the 1950s namesake sense".
Similar to how somebody could have been a hipster (in the 2000s+ sense [1]) before a hipster was a thing (before 2000s), even if they have been born in the 70s. Sure, the term already existed before the 70s, but it referred to a different thing.
[1] https://en.wikipedia.org/wiki/Hipster_(contemporary_subcultu...
The only newer sense is the American Right’s use of the term to replace “political correctness” as an empty epithet for everything and everyone it disagrees with.
So, one side could see woke in theory as a noble activist/social consciousness practice, which can not go wrong and helps liberate us all.
The other side might see woke in practice as intolerable virtue signalling and self-aggrandizing whose actions often border on farcical.
Found the problem.
Can say the same thing about code. Some people just honestly don't want to give away how simple the core logic is seemingly, and will lead you through myriad twists and turns to finally see the point.
The original bitcoin paper is a great example. I was able to follow the paper almost fully at my first read itself—despite my not having a formal background in maths.
...and as you said, many of the insubstantial papers hide behind jargon and unnecessarily complex equations, just to camouflage their lack of substance. It's frustrating to spend time deciphering a paper, only to realize that you've essentially wasted that time.
Here is an added complication: succinct technical communication can be efficient when communicating to peers who work on the exactly same domain, similar problems as you, and want digest your main ideas quickly.
On the other hand, for any particular paper, the size of the audience to whom it is directly relevant and addressed to can be small. The size of the audience who got to reading it anyway may be vast. (Maybe I am reading your paper because someone cited a method paper that in lieu of a proof or explanation writes just two words and citation to your paper. Maybe I am a freshly minted new student reading it for my first seminar. Maybe I am from a neighboring field and trying to understand what is happening in yours. Maybe I tried to find what people have already done with particular idea I just had and search engine gave your paper. And so on.)
During my (admittedly lackluster) academic career I recall spending much more time trying to read and understand papers that were not addressed to me than papers that were and where I enjoyed the succinct style that avoids details and present the results. (Maybe it is just an idiosyncratic trust issue on my part, because I am often skeptical of stated results and their interpretation, finding the methods more interesting). But that is not all.
I also noticed that genuine misunderstandings coming from "brief" communication of technical "details" were quite common; two different researches would state they "applied method X to avoid Y/seek Z[citation]" in exactly so many and almost exactly same words, where X,Y and Z were complicated technical terms, yet the authors would have quite different opinion what the meaning of those words were and what would be the intended reading and how and why X should be implemented.
In conclusion, I think many a scientific field would benefit from a style where authors were expected to clearly explain what they did and why (as clearly as possible).
1. The raft paper is titled "In Search of an Understandable Consensus Algorithm"
2. The abstract of this tutorial on Understanding Paxos https://www.ux.uis.no/~meling/papers/2013-paxostutorial-opod...
3. Lamport's own "Paxos made simple" https://lamport.azurewebsites.net/pubs/paxos-simple.pdf
A known fact is that it's impossible to actually implement it correctly, and the "approachable" paper seems to be a significant factor in this.
They're also, more often than not, tedious, badly explained, error prone, oft-skipped, and hardly ever read carefully, even during peer review for the paper that contains them. That's how mistakes stay unnoticed for decades in influential papers with tons of citations.
In essense, a paper's tone and languge is often more formality, academic tradition, ritual, and padding for publication purposes, than serving a real purpose.
I'm skeptical that the only way for them to be precise and technical is to make them impenetrable. I think there is a culture of academic writing (many different cultures, really) that has adopted a voice and writing style which became a parody of itself over time.
Here's a trivial example: You frequently see papers use the passive voice, something a middle school English teacher would mark with a red pen. 500 participants were asked, vs. we asked 500 participants. In what sense is the former more precise and technical? It's not. It does not convey any additional meaning. People use it to sound objective and distant, even when they really aren't.
Realistically, academic writers usually don't even think about it as much as that. They're just copying the tone of other papers, because there is a culture and it enforces certain behaviors on its members irrespective of the value.
Nobody likes doing it, I think. We just do it because we’re scared our papers won’t be accepted otherwise.
That said the assertion that most scientific articles are written in passive voice is outdated för quite some time. Most journal style guides advise to use active voice, e.g. https://www.nature.com/nature-portfolio/for-authors/write
When scientific papers have a clear list of authors and delineated section headings, this point is moot. And in such papers, again, repetitive strings of sentences that begin with the same "we..." emphasizes the producers of the work over the work itself.
People complain all the time about news being biased for being told from a reporter’s point of view, but complain all the same when events are reported in an encyclopedic manner as researchers do when they remove themselves from the events and the outcomes of their studies.
"Make it sound like we do cool stuff; but don't make it so precise that they can re-implement what we do. Let them come to us so we can co-author papers."
"Spectrum sharing in an “apple-like” or a fixed set sense is not a coexistence. ". What does that mean? Coexist? Who knows, the author thought they were being precise, but they understood the statement they made with a head full of context that gave it precise meaning. As readers, we can only scratch our own heads as to what that context could possibly be.
> What should be encouraged is for academics to blog about their research as well.
Why so binary? A blog would be hard to find, why not have both in the paper?
My view is similar to that of code vs docs: code should be as small, and as precise as possible, whereas docs are best when they’re explaining to humans how things fit together, high level. Also easier to maintain.
Hyper technical natural language mixed in with math is almost the worst of both worlds: low density of the actual formulas, with an incomprehensible wall of text surrounding it. And clearly this is an issue also for phd domain experts.
Not saying academic writing could be super simple but I also see no reason that the status quo is optimized more for comprehension than say social posturing.
I want to see: an explainer of the science/ideas/experiments/hipothesis
And instructions on how to reproduce the experiments/results
Some YouTubers are going in this direction
For the rest of it I don’t care. As long as researchers understand what’s going on, that’s what matters.
Inconsistent math notatation in papers along with vague terms in descriptions makes me so mad.
When they include good videos, they really stand out
There are a thousand papers out there making minor tweaks to the transformer architecture. 99% of them are also worthless and forgotten.
That's precisely what he shared this for, though. So someone willing to train a model with this tweak tries it.
Ideas like sparse attention, tree attention, residual attention, etc, all sound good on paper, but when researchers try to reproduce them they either find no results or results that don't scale. Even AliBi is turning out to be less powerful than scaled-down positional embeddings. It's almost a bitter lesson on its own: you can't beat the original transformer.
Optimizations that do stick around tend to be the ones that preserve the original algorithm but help with caching or memory accesses.
With large ML models, there probably is no intuition like this. We just don't know "if I do the common sense thing X, it surely will produce better results for a given benchmark" ... well we have no idea until it is tried out.
Explainers and their folksy, imprecise tone are good for things we already know are true. I’m skeptical on things which are unproven.
> I lost a series of pitched battles against Pytorch and biblatex, so I figured I’d just write a blog post instead.
So I think your accusation of his burying the lede on the lack of experiment is unwarranted.
Blog posts are written by those who arrive first.
In a weird way my mental model is: blog posts are the recon team discovering a new idea. They might have errors. They might be incomplete. Maybe they’re outright wrong. Stakes are lower as it took less effort to get there and less loss if a position is abandoned.
Then papers are authored, often much later, and they’re the regulars coming in to fortify a newly captured idea. They provide (or at least are supposed to) rigor to the idea. A fortification of a position that we decide is worth holding.
Yeah, this analogy is probably sloppy. But in my brain there’s an eternal conflict against ignorance as we keep advancing into the unknown.
I can't imagine judging scientific papers based on whether the author might be looking down on me, or thinks he knows better than me.
> if we were "allowed" to cite research that reads like this
Maybe you're looking down on yourself? You can cite anything you want to cite.
But a good writer can write great articles in whatever format they wish.
What do you call it when somebody takes the time to write about "a big discovery" they've made, but don't take the time to check if somebody else already did it? It's not like it's in some forgotten paper nobody has seen. It's in Pytorch itself.
Also this: "I’m 99.44% sure that it will resolve the outlier feedback loop that’s making quantization the subject of cascades of research."
Papers should be structured like fractals - that is, they should be "self-similar". The main text of the paper after the introduction should go into all the necessary details demonstrating the origins of the idea and proving that it has value. Then the introduction section should summarize all this, and take a less rigorous tone. The abstract should be a summary of the introduction. And then the title should summarize the abstract. If you really have a lot of technical work to do, maybe you can write a super long appendix and have the main body summarize that.
I myself probably spend as much time reading paper introductions as I do reading paper bodies, which means that probably 90% of the papers I read, I only read the introduction. I do this because I enjoy it more - I like new ideas, and the intros are a great way to get a lot of them. This blog post reads like a great paper introduction to me. It's easy to trick yourself into believing something is easy though, so an academic paper would have to back this up with an experiment.
The explicit 1 formulation is used in binary softmax, and the implicit (not seen 1) is used in multinomial softmax. I suspect this is the old "notation B looks silly in terms of notation A's standards."
The latter seems like something training could figure out by itself (zero doesn’t seem like a hard place to land with the weights producing V, although a bunch of zero weights would be needed), but the former is a bit awkward, as QK^T is quadratic in the weights.
In any case, this seems intuitively quite reasonable. But I do wonder whether the 1 in the denominator (equivalent to an exp(0) vote) is the best choice if the goal is to quantize well. 0 is in the middle of the numerical range, and perhaps the implicit null vote should be weighted lower than the middle of the range.
I didn’t see anything relevant on alternatives to softmax, since TFA is specifically questioning softmax in a multihead attention context.
Ultimately, neural networks are arbitrary function approximators. It doesn’t necessarily have to be “right” internally to fit the data. But if this new softmax allows transformers to learn more, that’s great.
You'd have to train the model with the quiet softmax before inferencing with it would work.
1. Have attention spans been declining? (slimemoldtimemold.com)
338 points by janandonly 4 hours ago | flag | hide | 254 comments
2. Attention Is Off By One (evanmiller.org)
400 points by elbasti 4 hours ago | flag | hide | 129 comments
Note that the #1 post is probably there because the title earlier had the provacative "Yes, 65%" appended to it. So even more numerical.Otherwise, i fear some of the 'magic' of transformer networks is that this amplification effect allows it to encode/memorize some results verbatim. And we often are seeing a heavily tuned internet regurgitator. So similar to the rise of RNNs with attention, which supposedly allowed them to focus on some things and ignore others but really often was just overfitting stuff, yielded more interesting results with the overfitting than without.
it's messy though, bear with me for the full explanation:
- your initial post says "<bot>" token, which looked like a mix of "chatbot" and ChatML, used by OpenAI
- there is a bo_S_ token, which acts as you described
- I averaged my attention over your post and the initial reply, which answers as if you were using "<bot>" in the misunderstood way
- when I go back and read your post, I realize the chatbot interpretation doesn't quite make sense, since you're referring to much more technical aspects than general "how do I AI", i.e. you understand <X> as a way to denote special tokens, not necessarily an XML tag
So the <|beginning of text|> token, with no context before it, learns to predict the first-token-in-a-document distribution. That's not quite the same as predicting nothing at all.
There are interesting things to said whether or not this turns out to be the case.
>This vector seems to get taller every model year, for example the recent LLaMA 2 model from Meta uses an embedding vector of length 3,204, which works out to 6KB+ in half-precision floating-point, just to represent one word in the vocabulary, which typically contains 30,000 - 50,000 entries.
>Now if you’re a memory-miserly C programmer like me, you might wonder, why in the world are these AI goobers using 6KB to represent something that ought to take, like 2 bytes tops? If their vocabulary is less than 2^16=65,384, we only need 16 bits to represent an entry, yeah?
>Well, here is what the Transformer is actually doing: it transforms (eh?) that input vector to an output vector of the same size, and that final 6KB output vector needs to encode absolutely everything needed to predict the token after the current one. The job of each layer of the Transformer is quite literally adding information to the original, single-word vector. This is where the residual (née skip) connections come in: all of the attention machinery is just adding supplementary material to that original two bytes’ worth of information, analyzing the larger context to indicate, for instance, that the word pupil is referring to a student, and not to the hole in your eye.
Firstly, he is confusing representation with encoding--he's right that 2 bytes is enough to encode any token. That is in fact approximately how it's done: a code book is indexed into (with a longint in pytorch, at least last I worked with it ~6 months ago). The purpose of the embedding is to allow the model to learn a representation of the token, a la word2vec. (Though this representation is purely based on the characters comprising the token and does not distinguish between "student" and "eye" in the case of "pupil" as in his example.)
Secondly, his description of each layer's function as adding information to the original vector misses the mark IMO--it is more like the original input is convolved with the weights of the transformer into the output. I am probably missing the mark a bit here as well.
Lastly, his statement that the embedding vector of the final token output needs all the info for the next token is plainly incorrect. The final decoder layer, when predicting the next token, uses all the information from the previous layer's hidden layer, which is the size of the hidden units times the number of tokens so far.
>This vector seems to get taller every model year, for example the recent LLaMA 2 model from Meta uses an embedding vector of length 3,204, which works out to 6KB+ in half-precision floating-point, just to represent one word in the vocabulary, which typically contains 30,000 - 50,000 entries.
>Now if you’re a memory-miserly C programmer like me, you might wonder, why in the world are these AI goobers using 6KB to represent something that ought to take, like 2 bytes tops? If their vocabulary is less than 2^16=65,384, we only need 16 bits to represent an entry, yeah?
The reason we have 3204 2B allocations is that each of the 2B contains info (latent space dimension). If you go to just the 2B representation, it is effectively one hot encoding which completely defeats the purpose of word embedding
I think the author is more correct than you are. It is not necessarily the case that we need 3,204 dimensions to represent the information contained in the tokens; in fact, the token embeddings live in a low-dimensional subspace; see footnote 6 here:
https://transformer-circuits.pub/2021/framework/index.html
> We performed PCA analysis of token embeddings and unembeddings. For models with large d_model, the spectrum quickly decayed, with the embeddings/unembeddings being concentrated in a relatively small fraction of the overall dimensions. To get a sense for whether they occupied the same or different subspaces, we concatenated the normalized embedding and unembedding matrices and applied PCA. This joint PCA process showed a combination of both "mixed" dimensions and dimensions used only by one; the existence of dimensions which are used by only one might be seen as a kind of upper bound on the extent to which they use the same subspace.
So some of the embedding dimensions are used to encode the input tokens and some are used to pick the output tokens (some are used for both), and everything else is only used in intermediate computations. This suggests that you might be able to improve on the standard transformer architecture by increasing (or increasing and then decreasing) the dimension, rather than using the same embedding dimensionality at each layer.
I think his description is basically correct given how the residual streams work. The output of each sublayer is basically added onto the input. See https://transformer-circuits.pub/2021/framework/index.html
> Lastly, his statement that the embedding vector of the final token output needs all the info for the next token is plainly incorrect. The final decoder layer, when predicting the next token, uses all the information from the previous layer's hidden layer, which is the size of the hidden units times the number of tokens so far.
I think the author is correct. Information is only moved between tokens in the attention layers, not in the MLP layers or in the final linear layer before the softmax. You can see how it’s implemented in nanoGPT: https://github.com/karpathy/nanoGPT/blob/f08abb45bd2285627d1...
At training time, probabilities for the next token are computed for each position, so if we feed in a sequence of n tokens, we basically get n training examples, one for each position, but at inference time, we only compute the next token since we’ve already output the preceding ones.
Softmax(x_i) = exp(x_i) / sum(exp(x_i)),
we should use instead what the author calls the Softmax_1 function, Softmax_1(x_i) = exp(x_i) / (1 + sum(exp(x_i))),
which would make it possible for each transformer head's attention probabilities to be zero, i.e., attend to nothing, by computing x_i's with values well below zero.Giving each transformer head the ability to ignore all tokens surely can't hurt, but it remains to be seen if it will actually improve transformer performance.
But in terms of a written piece of technical content, this is brilliantly written. Easy to follow and stay engaged. Well done.
Can't the MLP that processes the concatenated outputs the attention heads handle this? I don't understand why it should be critical that a head be allowed to put something close to zero in its segment of the concatenated vector if it's immediately going to get projected by an MLP anyway.
> This is what’s been happening in LLMs – for reasons that are only partially understood, Transformer models contain these outlier weights and are emitting Black Swan mega-activations that are much, much, much larger, like orders of magnitude larger, than their peers ...
meaning that once quantized you can either have a finer quantization since the range of possible values is smaller or you can pick a coarser strategy that saves bits for each weight.
-Learn a near-zero representation for some otherwise low-importance token, like delimiters or whitespace.
-When a head wants to "pass", emit an outlier activation to attend to that token nearly-exclusively.
But I'm surprised the model can't just use its existing tools (the post-concat projection layer and the following MLP block) to achieve the same thing. And if the answer is that it could do that, but tends to learn to use the outlier activation trick instead, will giving it a new tool that still allows the use of outlier activations be sufficient?
This writing brought a happy tear to my eye.
I'm confused what his goal is though:
I could imagine some theoretical reason to add a 1 there, but he starts by saying this can lead to smaller, more compactable models. Is he talking about the size the compressed weights? or pruning to a smaller model? or resistant more quantization?
Parts of the essay seemed to throw me off track, because I'm not sure if they are relevant at all to the proposal (eg the of the initial embedding and how many bits it would take the store the vocab size, etc).
1. { -1.4, 0.8, 2.7, 7.3 } : With a range of 8.7, you have a resolution of 0.034. This set quantizes to { 0, 64, 120, 255 }.
2. { -1400, 800, 2700, 7300 } : Resolution 34.1, quantizing to the same as the above { 0, 64, 120, 255 }.
3. { -0.008, -0.001, 0.009, 0.019 } : resolution 0.000106. This set quantizes to { 0, 66, 161, 255 }.
3. { -1.4, 0.8, 2.7, 7329 } : Resolution 28.7. This set quantizes to {0, 0, 0, 255 }. Oops -- we can no longer tell most of our weights apart.
You can see how this quantization works really well when all the numbers are close together, regardless of their absolute scale. Major outliers completely mess up the entire system. You can make more and more complicated quantization algorithms, but those will always come with tradeoffs. The best option would be to tame your weights so that they are again close together.
One of the downsides of having a field be so empirically driven is that theoretical arguments suggesting improvements just don't have much value until someone actually tests it and shows that it works (and not just tests it, but tests it at scale).
What is the minimum cost to get information that would say "this is better" or at least "this has a good chance of being better, spending a million training a bigger model is worth it"?
Could you rent a bank for 8 A100's for a day and try it out on a smaller model and prove something. Not cheap, but doesn't need VC money either. Probably about $400 on LambdaLabs to test with/without "quiet attention"
The OP can be as sure as he wants, but it is not worth sounding all "I told you so" without a single benchmark. Should be more "what about this - did you miss it ML people...?" in tone.
I suggest ways to measure it here: https://news.ycombinator.com/item?id=36855881 but the TL;DR is to choose a metric and compare the reduction in performance for quantized versions of the LM compared to the same LM without the modified Softmax.
I think that also likely holds for the quants, the difference could very well be within the error bars.
Anyway, it's been posted to r/locallama so I'm sure someone will try it within the hour and report back soon :P
However, I do think that if perplexity showed a lower drop-off using this modified softmax under quantization that would be an exciting finding and enough to indicate further experiments would definitely be worth doing.
But you are right - if it doesn't show an improvement it doesn't necessarily rule out that it could be helping.
Edit: In the Qualcomm AI paper mentioned in this post, they experiment on BERT uncased (109B param) and OPT 125M and are able to show the effects using perplexity.
I hadn't read the paper when I suggested the same approach, so I guess that is good validation it is worth trying.
Edit2: Actually they also test on ViT 22M, which would be even quicker to try I think.
:D
“ Your life is the sum of a remainder of an unbalanced equation inherent to the programming of the matrix. You are the eventuality of an anomaly, which despite my sincerest efforts I have been unable to eliminate from what is otherwise a harmony of mathematical precision. While it remains a burden assiduously avoided, it is not unexpected, and thus not beyond a measure of control. Which has led you, inexorably, here.”
Hence ergo - you are a NaN. A divide-by-zero.
identify the algorithmic obstacle(s) that stand in your way.
summarize your results in the style of a blog post that will persuade a significant chunk of the AI developer community to commit resources to an effort to eliminate said obstacles.
make it interesting enough to entice FrameworkFred to begin to read the article, snarky enough that he finishes it, and use concepts and notational conventions that ensure while he reads it his inner dialogue will roughly approximate the mood evoked by Homer Simpson saying "it's nu-cul-ar"."""
https://timdettmers.com/2022/08/17/llm-int8-and-emergent-fea...
The fact that these only emerge in larger models is likely one reason the author hasn't actually tried it.
This change would have all the probability going into the null option, in that case, basically.
Rather, isn’t the autoregressively predicted single next token a combination (based on attention) of all 6KB word tokens in the attention window.
So the size of memory where all information for next token prediction needs to be ”crammed into” is more like window_size*6KB, right?
He just says he doesn't want to spend any more time on it, which is unlikely to convince or motivate anybody else that he has discovered something important.
I don't know what to say past that, but it's worth reflecting on.
Why is this true? Because their existence implies some sort of preferred basis that aligns with the dims of the neural network, which is surprising?
It's not obvious why their existence is so contrary to what we knew.
Its easy to see how for negative numbers the softmax operator could simply refrain from making a decision
e.g. ``` sum(softmax1[-100, -100, -100]) ~= 1e-43 ```
But is there any basis to assume commas and whitespaces will be negatively correlated with other tokens?
And then the sneering tone of this article, sounds unprofessional and disrespectful in my opinion.
I also am pretty sure he’s wrong or at least he has to change layernorm to make this work. Attention simply does a weighted average of the Value Vectors, his change breaks that and I think will push the output closer to 0 as you stack the layers (especially considering Layer Norm). He really should do some small experiments to validate his idea first!
I thought everyone does that, because you don't need to work long with these models to get NaNs, and when you check why you see it's because of the exp functions. Then you fix it. Apparently people don't.
It's not like the neural models care if you approximate functions. They couldn't care less actually.
Clearly softmax is not too bad, if it is used extensively in all the most powerful models.
You might attract more, ahem, attention if it was immediately apparent from the name only what this attention head does that the current one does not. There's also that small matter of distinguishing the internal vs output softmax functions.
Buuut, it's missing the forest for the trees. The goal of the last step of attention (ref., Fig. 2, left in https://arxiv.org/abs/1706.03762) is not to add/say anything (as the author is saying) but to compute the relationship between the tokens (QK^T) and V -- in layman terms, simplifying, which tokens are related to each other. The softmax is there because it gives a representation that is nicer to work with, it gives probabilities, instead of unscaled matrix multiplication.
TLDR; author isn't wrong but he isn't right, practically speaking, either.
As to what the commenter above meant I can only guess, but it should be noted that Bilbo's audience reacts with puzzlement, unable to parse his words.