Could someone explain the last hidden state of the LLM ? What it shape is and how it is normally used - and why it hasn't been used yet to augment the next input? (which seems logical)
Could someone explain the last hidden state of the LLM ? What it shape is and how it is normally used - and why it hasn't been used yet to augment the next input? (which seems logical)
There's typically an "unembedding layer"/"classification head" that uses this hidden state to produce a softmax distribution over the LLM's vocabulary. In this case, we can think of this as "snapping" the hidden state into a single token and feeding that token into the next position of the autoregressive LLM.
In this sense, the last hidden state _does_ augment the next input. The authors simply propose directly feeding this hidden state into the next step rather than reducing it into a single token—thus, reasoning in continuous latent space rather than discrete token space.
Don't forget, non-linearity is fundamental to the whole process, otherwise you'd just have one large linear transformation. Maybe there's a similar role for discretization? :shrug:
tldr; the method requires training in a way that loses one of the major benefits of transformers, but maybe in some scenarios that loss is worth it.