Unless you consider the entire instance as a singular instance and don't use any hidden states, then I guess it could be considered feed-forward.
I don't know. Feedforward doesn't seem like a useful term tbh. Some people mean feedforward as information only goes one direction, but that depends on your arrow. Autoregressive seems more useful here.
Definition of Feedforward (from wiki):
``` A feedforward neural network (FNN) is an artificial neural network wherein connections between the nodes do not form a cycle.[1] As such, it is different from its descendant: recurrent neural networks. ```
Hofstadter expected any intelligent neural network would need to be recurrent, ie looping back on itself (in the vein of his book “I am a strange loop”).
GPT is not recurrent. It takes in some text input, does a fixed amount of computation in 1 pass through the network, then outputs the next word. He is surprised it doesn’t need to loop for an arbitrary amount of time to “think about” what to say.
Being put into an auto-regressive system (where the N-th word it generates gets appended to the prompt that gets sent back into the network to generate the N+1th word) doesn’t make the neural network itself not Feedforward.
But I'm also surprised that Hofstadter keys in on this so heavily. The fact that he wrote an entire pop-sci book on recursion would, in my mind, make him (1) less surprised that AR and R aren't so dissimilar and (2) more sensitive to the sorts of issues that make R more difficult to get working in practice.
(In my mind, differentiating between auto-regressive and recursive in this case is kind of the same as differentiating between imperative loops and recursion -- there are extremely important differences in practice but being surprised that a program was written using while loops where you imagined left-folds would be absolutely required seems a bit... odd.)
Recurrent neural networks have the recursion as part of the training regime. GPT only has auto-regressive "recursion" as part of the inference runtime regime.
I think Hofstadter is surprised that you can appear so intelligent without any recursion in the learning/training regime, with the added implication that you can appear so intelligent with a fixed amount of computation per word.
Consider a simpler case: a small neural network that takes 2 numbers and adds them together, producing 1 number as output.
This network is very obviously feedforward, and probably very tiny with few layers.
Say I have a list of numbers [1, 2, 5] that I want to sum. If I send 1 and 2 through the network, get 3 as a result, then I send 3 and 5 through the network, and get a final answer of 8, my network has not suddenly become non-feedforward just because I fed the output back into it.
The key distinguishing factor between feedforward and non-feedforward is if the network itself loops back around and, at training time, it learns how to make use of this ability to pass data to itself to maintain some hidden context between passes.
There is no such learned hidden context in my addition example, and none in GPT.
---
There actually is a tiny caveat here: the RL fine-tuning process OpenAI has done on its models ("RLHF" & friends) actually does allow for a very, very small amount of information leakage between passes because you are rewarding whole responses, so the model can learn little patterns of what tokens in the beginning of the response led to certain tokens at the end of the response and reinforce those patterns.
The model could learn to encode small bits of "hidden" information in particular token choices that the human raters wouldn't notice. In this case, there is a (small but non-zero) amount of learned hidden context. But this is not what Hofstadter is talking about -- the non-RLHF'd base model is just as intelligent, just harder to use.
Edit: Nope. TIL feed-forward means no loops.