Shallow Feed-Forward Neural Networks as Alternative to Attention in Transformers | Hacker News Reader