In this case, by analogy, Ferrari made the comparison to the super cheap SUV from 2003. That is, OpenAI compared GPT3 to BERT on the SuperGLUE benchmark, in the paper announcing GPT3 [1].
They did so to demonstrate GPT3's ability to learn a new task given only a few examples of the task ("few-shot learning"). The limited amount of task-specific training data was a signficant handicap that GPT3 was able to overcome, like a Ferrari towing a two ton trailer outperforming an old SUV towing nothing.
What this paper claims is that encoder type models can also achieve few-shot learning. The headline should be "AI training method achieves few-shot learning with 99.9% fewer parameters than GPT3." That's the innovation here, not outperforming GPT3 on a benchmark that GPT3 isn't particularly good at.
This paper however, is not replicating that flexibility. Their proposed model is much closer to expectation maximization than it is an instance of what GPT3 does. The novelty of their work, what makes it genuinely useful, is they provide a practical and fairly general way to leverage pre-trained language models to produce classifiers for specific tasks using a very small amount of labeled data. Requiring less effort compared to what would go into fine-tuning. This approach to distillation is an instance of https://en.wikipedia.org/wiki/Semi-supervised_learning.
Compared to GPT-3, this approach remains at a severe disadvantage when amount of effort and time required to gather data and train something useful is accounted for. On the other hand, if you can fit your problem into the format proposed by the paper, you will likely have more control on the final model's behavior using a small amount of labeled examples (for your specific task), at a significantly lower cost of computation at inference time. Focusing on parameters however, is not even wrong.
It is true that this model requires a large amount of unlabeled data, in addition to a small amount of labeled data, but gathering unlabeled data is often easy.
So I think it's potentially misleading to say "this approach remains at a severe disadvantage when amount of effort and time required to gather data and train something useful is accounted for".
I'd say this approach actually has a serious advantage compared to GPT3, which is locked inside the walls of OpenAI, can only be used with their permission, and is too big for most people to use (let alone train) anyway. The cost and effort to use this approach on large real world problems is probably less than using GPT3.
And it may also have an advantage in terms of the total amount of data and training required, when all of the data and training in the original pretrained models is included. That is, it may be a more efficient way to train new models from scratch for tasks for which few labeled examples are available.
1: https://ai.googleblog.com/2020/09/advancing-nlp-with-efficie...
> It is true that this model requires a large amount of unlabeled data, in addition to a small amount of labeled data, but gathering unlabeled data is often easy.
Gathering unlabeled data is easier than labeled but can still be challenging. You'll often require careful thought in assembling a distribution of examples. Being able to skip that step yields a significant savings even if not as much as that gained from going from labeled to largely unlabeled data.
> I'd say this approach actually has a serious advantage compared to GPT3, which is locked inside the walls of OpenAI, can only be used with their permission, and is too big for most people to use (let alone train) anyway.
I fully agree and said as much too.
> The cost and effort to use this approach on large real world problems is probably less than using GPT3.
I'd say it depends. Most of the effort with GPT3 will involve edge cases. Having a system in front to handle these might eat into labor savings but you could still end up net positive. It's difficult to say without real world data, you might be correct.
> That is, it may be a more efficient way to train new models from scratch for tasks for which few labeled examples are available.
You're right in general, I think. But it's still worth pointing out GPT3's advantage. It combines a lot of general capabilities, which together with its generative ability and flexibility to input means the level of expertise required to get something useful will be much lower compared to this semi-supervised learning approach. And there are some capabilities it's displayed, one example of many being discussing, querying, pattern matching on computer code that seem hard to replicate with this method.
What is this sentence supposed to convey? I'm an NLP practicioner/researcher and this isn't even true - as GPT isn't a "state machine" as the latent space is continuous and not finite.
Moreover, there is nothing that makes GPT-3 "not capable of learning." It has had very exciting results from language modeling a zero-shot task at inference time, but there's nothing (besides compute) precluding fine-tuning of it in principle.
I agree with the rest of your comment.
There are examples where it is able to recognize and continue patterns in strings which if manually generated, would have required a FSM. In fact, some of the more impressive examples would require a stack of some sort so I thought I was rather underselling its capabilities in that arena.
> as GPT isn't a "state machine" as the latent space is continuous and not finite.
Technically speaking, that is impossible since these models leverage floating point numbers and are limited in memory to whatever hidden and self-attention layers.
Practically speaking, in order to generate strings based on patterns as mentioned prior, there must be abstract states which correspond to states and state changes such that thinking in terms of at least state machines is useful.
Studying LMs in terms of automata is not strange, there have been papers which do this for specific trained RNNs (such as https://arxiv.org/abs/1711.09576). I contend GPT-3 is capable, to a certain extent, of generating these dynamically at inference time.
As far as I know this way of extracting what LMs are doing hasn't been done for Transformers but you can also frame Transformers in term of RNNs so there's no reason why such methods wouldn't readily apply to them too.
> Moreover, there is nothing that makes GPT-3 "not capable of learning."
I specifically addressed that: to count as learning, without diluting the utility of the term, it has to be capable of remembering. Without permanent changes to its weights, the use of the term learning stretches the word beyond utility.
Yes, my claim is that there is nothing that makes its weights incapable of being fine-tuned and thus changed.
See: hadoop->spark->spark with infiniband->spark with nvme->timely dataflow on a laptop; csv parsers that don't support scientific notation floats, etc. I'm sure others can give moren interesting examples.
The closest benchmark I could find was the "FASTER State Management for Timely Dataflow" paper from ETH Zurich, but that wasn't run on a laptop.
Few-shot GPT3 outperforms a BERT-based baseline.
That's like saying you can look at the Eiffel tower and it's schematics so what's so hard about getting a spare million dollars and building it.
But training a model of this size requires you to use thousands of GPUs, or wait forever. That will sum up to millions of $$$ in rental and electricity costs.
As far as my thinking goes, they are open more than enough. Thank you for your input!
Plus, they are called OpenAI but producing a closed source product...
I don't find that to be very open at all.
Note that GPT-3's approach to *GLUE involved no training on the task, just a good choice of prompt, whereas PET and iPET also use fine-tuning. Also, because distillation takes large amounts of training data, they use ensembles in the true few-shot regime, so their parameter efficiency is significantly worse than they advertise.
https://www.reddit.com/r/slatestarcodex/comments/itrcac/smal...
---
Timo Schick:
Finally, I do not really agree with your last two paragraphs, especially "One is about semi-supervised learning, that says by exploiting task-specific architectures you can do fairly well with low amounts of labelled data.": If you leave out the final distillation step (which is not required for good performance), we use the exact same architecture for all tasks. In what sense is this more task-specific than GPT-3? I would not consider "exploiting task-specific architectures" to be a (fundamental) part of the paper.
My reply:
So what I mean here is that masked training and bidirectional transformer models like BERT have always been designed as a way to get good scores in analysis tasks like Q&A, even if they are pretrained on general text, whereas unidirectional generative transformer models are now basically only relevant for generative tasks. You can say, well, both architectures can do both tasks, so is it really task specific?, but ultimately, yes, we've selected ALBERT because it's better for Q&A tasks, and we've selected unidirectional transformers in other things because they're better for generative tasks.
So I guess the problem I have is with the merits of your thesis, “Can we achieve similar few-shot performance to GPT-3 without requiring billions of parameters?” OpenAI didn't present few-shot learning as if it were an optimal method; their headline achievement was not “here's the best way to...” but “I bet you never expected that this could...”. And so while it's definitely true that a BERT-derived model will outperform a GPT-derived model even at lower parameter counts on these sort of tasks, nothing new or interesting is being said by it. Everyone already knows that a bidirectional GPT-3 would be better at Q&A, and so that's what a smaller bidirectional model should be competing against. GPT-3 is only interesting in this context because it's not the optimal model (or training routine).
So while it's also true that if your aim is SOTA in few-shot learning then you should definitely use a bidirectional transformer with all the new tricks, if your goal is to understand PET in a context that includes GPT-3, doing so merely makes it harder to see what's going on.
Do you know of any other models that should be used for such a comparison, or are there already any relevant results on SuperGLUE that should be mentioned?
PET (well, a version called iPET from the same author) is at #9 on the SuperGLUE leaderboard [1], and none of the models above it mention being evaluated by few-shot learning.