764 karma · joined September 19, 2012
So why are so many people still employed as e.g. software engineers? People aren’t prompting the models correctly? They’re only asking 10 times instead of 20? They’re holding it wrong?
> Yes, but that's not because of Windows itself
Come on. There’s a reason Windows users all want to install crappy security products: they’ve been routinely having their files encrypted and held for ransom for the last decade.
https://en.wikipedia.org/wiki/Parkinson's_Law > The growth was presented mathematically with the formula x = (2k^m + P)/n, in which k was the number of officials wanting subordinates, m was the hours they spent writing minutes to each other.
(check the original paper for details since this obviously doesn’t explain what n is)
I think the author is more correct than you are. It is not necessarily the case that we need 3,204 dimensions to represent the information contained in the tokens; in fact, the token embeddings live in a low-dimensional subspace; see footnote 6 here:
https://transformer-circuits.pub/2021/framework/index.html
> We performed PCA analysis of token embeddings and unembeddings. For models with large d_model, the spectrum quickly decayed, with the embeddings/unembeddings being concentrated in a relatively small fraction of the overall dimensions. To get a sense for whether they occupied the same or different subspaces, we concatenated the normalized embedding and unembedding matrices and applied PCA. This joint PCA process showed a combination of both "mixed" dimensions and dimensions used only by one; the existence of dimensions which are used by only one might be seen as a kind of upper bound on the extent to which they use the same subspace.
So some of the embedding dimensions are used to encode the input tokens and some are used to pick the output tokens (some are used for both), and everything else is only used in intermediate computations. This suggests that you might be able to improve on the standard transformer architecture by increasing (or increasing and then decreasing) the dimension, rather than using the same embedding dimensionality at each layer.
I think his description is basically correct given how the residual streams work. The output of each sublayer is basically added onto the input. See https://transformer-circuits.pub/2021/framework/index.html
> Lastly, his statement that the embedding vector of the final token output needs all the info for the next token is plainly incorrect. The final decoder layer, when predicting the next token, uses all the information from the previous layer's hidden layer, which is the size of the hidden units times the number of tokens so far.
I think the author is correct. Information is only moved between tokens in the attention layers, not in the MLP layers or in the final linear layer before the softmax. You can see how it’s implemented in nanoGPT: https://github.com/karpathy/nanoGPT/blob/f08abb45bd2285627d1...
At training time, probabilities for the next token are computed for each position, so if we feed in a sequence of n tokens, we basically get n training examples, one for each position, but at inference time, we only compute the next token since we’ve already output the preceding ones.
Most disasters result from multiple things going wrong together, hence the importance of addressing them individually.
https://www.nytimes.com/2019/03/26/arts/television/jussie-sm...
What do you mean by continuous? Obviously if you take a number line and remove either the rational or irrational numbers, you will end up with infinitely many holes.
The thing that makes floating point numbers unique is that, for any given representation, there are actually only finitely many. There’s a largest possible value and a smallest possible value and each number will have gaps on either side of it. I think you meant that the rationals and irrationals are dense (for any two distinct numbers, you can find another number between them), which is also false for floating-point numbers.
Maybe your language isn’t so “brilliant” if some of the smartest programmers in the world still can’t figure it out.
Well, we have 2020s America as a comparison. Whether or not you think we’ve recognized the problem, it certainly hasn’t been solved to the extent that e.g. we actually flew people to the moon and back.
The issue with the current situation is that in principle I could be doing the work in New York (it doesn’t matter where I do it), but the office is closed and I don’t live in New York. So I have strong reasons for not doing my work in New York, but in principle I could be doing it there.
I guess there’s a further clause that you have to spend at least one day out of the year in New York to be taxed there, which unfortunately I have done (despite no longer traveling there for any work-related purpose).