Maybe it comes down to semantics but when I read things like [1] I come away with the idea that the weights are altered. But it could also just be my misunderstanding.
1: https://towardsdatascience.com/an-intuitive-explanation-of-s...
1: https://towardsdatascience.com/an-intuitive-explanation-of-s...
I agree that in some sense the attention weights are more like meta-weights that are applied to the context of the conversation to decide how to actually weight the various words. So it's totally correct to say that previous words in the conversation affect how future words will be weighted, and I think it's reasonable to call that 'learning': for example, you can tell ChatGPT new words and it will be able to use them in context. Again though, people usually take 'learning' to mean making updates to the trained parameters of the model itself, which obviously isn't happening here.