I was surprised he said this too. Even without the autoregressive part of GPT models you have a deep transformer with attention, so even a single forward pass can modify its own intermediate outputs.
The interesting thing though is that attention effectively allows a model to meta-learn based on the current context, so in many ways, it may be thought of as analogous to a brain without long term memory.
What do you mean? Attention is just matrix multiplication and softmax, it's all feed forward.
I wasn't disputing that it's feed-forward. I just meant that stacked transformer layers can be thought of as an iterative refinement of the intermediate activations. Not the same as an autoregressive process that receives previous outputs as inputs, but far more expressive than a single transformer layer.