In this case, by analogy, Ferrari made the comparison to the super cheap SUV from 2003. That is, OpenAI compared GPT3 to BERT on the SuperGLUE benchmark, in the paper announcing GPT3 [1].
They did so to demonstrate GPT3's ability to learn a new task given only a few examples of the task ("few-shot learning"). The limited amount of task-specific training data was a signficant handicap that GPT3 was able to overcome, like a Ferrari towing a two ton trailer outperforming an old SUV towing nothing.
What this paper claims is that encoder type models can also achieve few-shot learning. The headline should be "AI training method achieves few-shot learning with 99.9% fewer parameters than GPT3." That's the innovation here, not outperforming GPT3 on a benchmark that GPT3 isn't particularly good at.
This paper however, is not replicating that flexibility. Their proposed model is much closer to expectation maximization than it is an instance of what GPT3 does. The novelty of their work, what makes it genuinely useful, is they provide a practical and fairly general way to leverage pre-trained language models to produce classifiers for specific tasks using a very small amount of labeled data. Requiring less effort compared to what would go into fine-tuning. This approach to distillation is an instance of https://en.wikipedia.org/wiki/Semi-supervised_learning.
Compared to GPT-3, this approach remains at a severe disadvantage when amount of effort and time required to gather data and train something useful is accounted for. On the other hand, if you can fit your problem into the format proposed by the paper, you will likely have more control on the final model's behavior using a small amount of labeled examples (for your specific task), at a significantly lower cost of computation at inference time. Focusing on parameters however, is not even wrong.
It is true that this model requires a large amount of unlabeled data, in addition to a small amount of labeled data, but gathering unlabeled data is often easy.
So I think it's potentially misleading to say "this approach remains at a severe disadvantage when amount of effort and time required to gather data and train something useful is accounted for".
I'd say this approach actually has a serious advantage compared to GPT3, which is locked inside the walls of OpenAI, can only be used with their permission, and is too big for most people to use (let alone train) anyway. The cost and effort to use this approach on large real world problems is probably less than using GPT3.
And it may also have an advantage in terms of the total amount of data and training required, when all of the data and training in the original pretrained models is included. That is, it may be a more efficient way to train new models from scratch for tasks for which few labeled examples are available.
1: https://ai.googleblog.com/2020/09/advancing-nlp-with-efficie...
> It is true that this model requires a large amount of unlabeled data, in addition to a small amount of labeled data, but gathering unlabeled data is often easy.
Gathering unlabeled data is easier than labeled but can still be challenging. You'll often require careful thought in assembling a distribution of examples. Being able to skip that step yields a significant savings even if not as much as that gained from going from labeled to largely unlabeled data.
> I'd say this approach actually has a serious advantage compared to GPT3, which is locked inside the walls of OpenAI, can only be used with their permission, and is too big for most people to use (let alone train) anyway.
I fully agree and said as much too.
> The cost and effort to use this approach on large real world problems is probably less than using GPT3.
I'd say it depends. Most of the effort with GPT3 will involve edge cases. Having a system in front to handle these might eat into labor savings but you could still end up net positive. It's difficult to say without real world data, you might be correct.
> That is, it may be a more efficient way to train new models from scratch for tasks for which few labeled examples are available.
You're right in general, I think. But it's still worth pointing out GPT3's advantage. It combines a lot of general capabilities, which together with its generative ability and flexibility to input means the level of expertise required to get something useful will be much lower compared to this semi-supervised learning approach. And there are some capabilities it's displayed, one example of many being discussing, querying, pattern matching on computer code that seem hard to replicate with this method.
What is this sentence supposed to convey? I'm an NLP practicioner/researcher and this isn't even true - as GPT isn't a "state machine" as the latent space is continuous and not finite.
Moreover, there is nothing that makes GPT-3 "not capable of learning." It has had very exciting results from language modeling a zero-shot task at inference time, but there's nothing (besides compute) precluding fine-tuning of it in principle.
I agree with the rest of your comment.
There are examples where it is able to recognize and continue patterns in strings which if manually generated, would have required a FSM. In fact, some of the more impressive examples would require a stack of some sort so I thought I was rather underselling its capabilities in that arena.
> as GPT isn't a "state machine" as the latent space is continuous and not finite.
Technically speaking, that is impossible since these models leverage floating point numbers and are limited in memory to whatever hidden and self-attention layers.
Practically speaking, in order to generate strings based on patterns as mentioned prior, there must be abstract states which correspond to states and state changes such that thinking in terms of at least state machines is useful.
Studying LMs in terms of automata is not strange, there have been papers which do this for specific trained RNNs (such as https://arxiv.org/abs/1711.09576). I contend GPT-3 is capable, to a certain extent, of generating these dynamically at inference time.
As far as I know this way of extracting what LMs are doing hasn't been done for Transformers but you can also frame Transformers in term of RNNs so there's no reason why such methods wouldn't readily apply to them too.
> Moreover, there is nothing that makes GPT-3 "not capable of learning."
I specifically addressed that: to count as learning, without diluting the utility of the term, it has to be capable of remembering. Without permanent changes to its weights, the use of the term learning stretches the word beyond utility.
Yes, my claim is that there is nothing that makes its weights incapable of being fine-tuned and thus changed.