Google ‘BigBird’ Achieves SOTA Performance on Long-Context NLP Tasks
syncedreview.com
syncedreview.com
Yannic Kilcher - Big Bird: Transformers for Longer Sequences (Paper Explained)
GPT-3 is only using a sequence length of 2048. In most of our paper, we use 4096, but we can go much larger 16k+. Of course, we don't nearly have as many parameters as GPT-3, so our generalization may not be as good.
BigBird is just an attention mechanism and could actually be complementary to GPT-3.
The original implementation only took a couple of months and was primarily motivated by internal Google applications. Natural Questions was the first external benchmark we tried to validate on, which took a few months to find the right setup. All the other datasets, took a few weeks but the effort was done in parallel given the large team.
There was quite a bit of frustration dealing with Tensorflow, TPUs, and the XLA compiler that maybe set us back a few months, too.
I am thinking maybe longer context window, faster training and less memory use, but what about performance, will it measure up?
We believe something like BigBird can be complementary to GPT-3. GPT-3 is still limited to 2048 tokens. We'd like to think that we could generate longer, more coherent stories by using more context.
The moving pieces here in BigBird are a vector associated with each token in the sequence and a few more global vectors that you can think of as latent variables. Those pieces are present at each layer. The vector in layer i+1 at position t in the sequence is a function of a bunch of the vectors in layer i.
If position t depends on all of the other positions, then you end up computing all combinations. A sequence of length N has N^2 combinations. Papers like this are using different patterns.
In this one, position t only depends on local positions in some small window around t, those global/latent-ish variables, and a small number of random positions. The number of combinations is (the size of the local window + the number of random pieces to attend to + the number of global vectors to attend to) times the number of positions, so it's some k * N rather than N^2. That lets you scale to longer sequences.
I thought it had fixed length window. Can you explain how it differs from vanilla transformer other than the size.
> We use the same model and architecture as GPT-2 [RWC+19], including the modified initialization, pre-normalization, and reversible tokenization described therein, with the exception that we use alternating dense and locally banded sparse attention patterns in the layers of the transformer, similar to the Sparse Transformer.
In the referenced paper (sparse transformer) they showed a bunch of different sparsity patterns, and I believe they're referring to either their banded block diagonal sparsity or a true banded diagonal pattern (local windows like bigbird and some other papers). Unfortunately, that paper also was light on details and the repo they open sourced alongside it is inscrutible.
Not flying was a joke from the early design period, I heard. At the end it took off pretty well.