To the OP: please consider changing the headline to "Exploiting Cloze Questions for Few Shot Text Classification and Natural Language Inference" and link to the original paper at https://arxiv.org/abs/2001.07676 instead of this PR puff piece.
To the OP: please consider changing the headline to "Exploiting Cloze Questions for Few Shot Text Classification and Natural Language Inference" and link to the original paper at https://arxiv.org/abs/2001.07676 instead of this PR puff piece.
Which is fine?
What's the point of getting people outside of the field reading a paper by essentially lying about the content? They come expecting A, and if they read it they understand it's actually about B. Loss of time for everyone except the author who wants to make a little buzz.
That's trivial and has already been done. GPT-3 didn't even get SOTA on SuperGlue.
These are the sort of misunderstandings that could have been avoided if the title was better.
In general, paper with "new variation of cloze pre-training task for this specific task" is a new section of the literature that is rapidly becoming sort of mundane and uninteresting because there are so many papers doing small variations of the same basic idea.
Of course neither GPT-3 nor the PET paper claim SOTA on SuperGLUE. They used a few-shot learning setup with 32 examples per task The normal SuperGLUE setup has hundreds or thousands of examples per task [1].
> In general, paper with "new variation of cloze pre-training task for this specific task" is a new section of the literature that is rapidly becoming sort of mundane and uninteresting because there are so many papers doing small variations of the same basic idea.
Could you please link to some of the work you are referring to?
[1] Table 1 in https://w4ngatang.github.io/static/papers/superglue.pdf