GPT-3 is only using a sequence length of 2048. In most of our paper, we use 4096, but we can go much larger 16k+. Of course, we don't nearly have as many parameters as GPT-3, so our generalization may not be as good.
BigBird is just an attention mechanism and could actually be complementary to GPT-3.
The original implementation only took a couple of months and was primarily motivated by internal Google applications. Natural Questions was the first external benchmark we tried to validate on, which took a few months to find the right setup. All the other datasets, took a few weeks but the effort was done in parallel given the large team.
There was quite a bit of frustration dealing with Tensorflow, TPUs, and the XLA compiler that maybe set us back a few months, too.
I am thinking maybe longer context window, faster training and less memory use, but what about performance, will it measure up?
We believe something like BigBird can be complementary to GPT-3. GPT-3 is still limited to 2048 tokens. We'd like to think that we could generate longer, more coherent stories by using more context.