The Illustrated Retrieval Transformer
jalammar.github.io
jalammar.github.io
A recognition model should use a similar mechanism to store short term context in a memory buffer from previous frames and a large external database of long term key value pairs that retain relevant semantic information for given embeddings.
Doing so will make it possible to update and expand the models without having to retrain and enable much better zero/few shot learning.
We already have a hacky version of this in our production app for food recognition. For new users we use a standard CNN to predict the items present in the image, once a user logs a few meals we use nearest neighbor search to match new images against previously submitted entries, which works extremely well.
The cherry on top is that you could get not only information from the model but also the sources it used to make up its mind!
Summary: The latest batch of language models can be much smaller yet achieve GPT-3 like performance by being able to query a database or search the web for information. A key indication is that building larger and larger models is not the only way to improve performance.
Hope you find it useful. All feedback is welcome!
Parameter count is not very accurate measure of model performance. Mixture of Expert models like the Switch Transformer [2] can be 1 trillion parameters in size, but are not 5X the performance, for example.
They clock the retrieval at 10 ms, unclear if that includes the BERT inference, however. My assumption is that it does not.
Any plans to do one on MOE for LM?
Thank you!
What is the size of the database and how does it compare to the size of the gpt3 model? "( 2 trillion multi-lingual tokens )" But how much memory is that
How much compute does each method need?
Can you run it on a laptop?
Compute: Building the database requires lots of BERT pre-computation. But at inference time, RETRO is at least one BERT forward pass (a batch of all the chunks that it broke the input prompt into). Then the neighbors are computer via the Retro encoder (which seems small at 2 Transformer layers). Then the input prompt is processed similar to a GPT of 32 layers (with attention to the neighbors).
Run on a laptop? Perhaps on CPU. Try running T0 [1] which is larger at 11B parameters.