That looks like a classic Actor/Critic setup, yet it's not mentioned even once in the paper. Am I missing some large difference here?
Anyone who works with deep architectures and momentum-based optimizers knows that the first few updates alone provide large improvements in loss. In this paper the breakthrough is that computing these first few updates at test time enables one to describe the algorithm as "without training" and therefore attract hype.
But they aren't updating the model weights. They're iteratively updating the prompt. It's automating the process that humans use with generative models.
Agreed that it's conceptually equivalent though.