Generative Modeling with Sparse Transformers
openai.com
openai.com
But it's quite surprising that this also works for text data, especially that the fixed pattern performs better than the strided one, despite there not being anything analogous to image boundaries in the data.
It'd also be interesting to see what happens for other decompositions, such as 3 layers of ∛N or a logarithmic stack of dilated convolutions.
Maybe the reason "strided attention" didn't work as well is that it would require the network to put this context summary in every column lest the rows below be unable to access it. That would waste features since the summary wouldn't vary much over time but would still be stored in full at each step.
If this is true, the approach they used for images might actually be inefficient in a similar way.
The innovation here is that the transformer is compressed, allowing the system to deal with longer sequences.
Convolutions have a different set of weights for each position offset (with a fixed window size), and reuse those weights across the entire input space.
Transformer-based networks like this work compute attention functions between the current position's encoding and every previous position, then use the outputs to compute a weighted sum of the encodings at those positions. Hence they can look at an arbitrarily large window and the number of parameters they have is independent of the size of that window.
However, I'm a bit disappointed with the code release. I was expecting the full source code and setup.
They have a lot of incentives not to. Keeping the code under wraps allows them to maintain an edge over other researchers and companies and the space, which helps them secure more funding, publish more papers etc.
It's also not really fair to them if some other research group uses code from OpenAI to achieve results and then doesn't share their own code or modifications.