What are transduction neural networks, and how are they different from existing attention-based transformer models?
Attention based models don't necessarily need to be sequence to sequence. They can be classifiers, decoder only, etc. Attention is just one tool in the ML architecture toolkit.