Thank you. This one?
I think a neural network can be considered a transformer if it contains a stack of attention blocks as its core mechanism.
In contrast to seq-2-seq use, for generative language models such as ChatGPT you only have access to preceding (not forward) context in order to decide what to generate next, so the encoder part of the architecture is not applicable and a decoder-only transformer is used.