> New Causal Encoder–Decoder architecture: just 8B active parameters for input, 16B for output.
Oh interesting, I can assume what the benefits is for including the Encoder, but whats the downside? I’m thinking GPT (which is decoder only) ruled out Encoder for a reason?