From what I've read over the last few days, the "attention" mechanism used in chatGPT and similar LLMs does indeed dynamically change weights of a portion of the model.
From what I've read over the last few days, the "attention" mechanism used in chatGPT and similar LLMs does indeed dynamically change weights of a portion of the model.
Now, you may argue that because it's a multiplication with a linear or affine kernel, you might as well use commutative property of scalar multiplication and multiply the factor with the weights first, and then multiply with the input to the kernel.
But this only holds for very few kernels.
when training a model, the forward pass would happen i.e the generation and then depending on how close to truth it was, the configuration settings (aka the weights/neurons) would be adjusted to incorporate whatever little insight was gained from the text.
Weights are matrices. The values of the matrices aren't changing.
1: https://towardsdatascience.com/an-intuitive-explanation-of-s...
I agree that in some sense the attention weights are more like meta-weights that are applied to the context of the conversation to decide how to actually weight the various words. So it's totally correct to say that previous words in the conversation affect how future words will be weighted, and I think it's reasonable to call that 'learning': for example, you can tell ChatGPT new words and it will be able to use them in context. Again though, people usually take 'learning' to mean making updates to the trained parameters of the model itself, which obviously isn't happening here.