Self-attention transforms a prompt into a low-rank weight-update
arxiv.org
arxiv.org
I wouldn't call it gradient descent. Residual connections of the form x_{i+1} = x_i + f(x_i) essentially form an "update rule" with one iteration per layer. Newton's method, gradient descent, fixed point iteration, conjugate gradient methods, ODE integration, etc can all be expressed as an "update rule" that takes a previous value and adds a modifier to produce a new value. It would be more accurate to say that each residual layer is a universal approximator of any imaginable update rule including gradient descent.