Is that an actual transformer, though? Like with encoder and decoder layers? That’s the part I never truly understood. Or is it “just” an example of a neural network? Thanks!
In contrast to seq-2-seq use, for generative language models such as ChatGPT you only have access to preceding (not forward) context in order to decide what to generate next, so the encoder part of the architecture is not applicable and a decoder-only transformer is used.
I think a neural network can be considered a transformer if it contains a stack of attention blocks as its core mechanism.