> So if there is a method that achieves the same goal with a model that is simple enough to be used by normal developers and researchers without OpenAI-scale infrastructure, that does seem buzz-worthy.
That's trivial and has already been done. GPT-3 didn't even get SOTA on SuperGlue.
These are the sort of misunderstandings that could have been avoided if the title was better.
In general, paper with "new variation of cloze pre-training task for this specific task" is a new section of the literature that is rapidly becoming sort of mundane and uninteresting because there are so many papers doing small variations of the same basic idea.