Improving Language Understanding with Unsupervised Learning
blog.openai.com
blog.openai.com
The basic approach is the same as our ULMFiT (http://nlp.fast.ai/classification/2018/05/15/introducting-ul...) model - pre-train a language model (a model that predicts the next word in a sequence) on a large corpus, and then modify the language model slightly for whatever task you wish to do (e.g. text classification). Finally, fine-tune that model using your target corpus (e.g. texts labeled with classes).
This new paper has two significant leaps over ULMFiT:
- Replace the RNN with a transformer model
- Apply to many more types of problem.
Note that although the original language model takes them a long time to train (a month on 8 GPUs), there's almost no reason for anyone else to create their own model from scratch, except if you need to use this approach on a language that doesn't have a pre-trained model yet. The transfer learning fine-tuning doesn't take anywhere close to as long as the language model pre-training, and you can just use the existing pre-trained weights.
The previous HN discussion on ULMFit may also be of interest: https://news.ycombinator.com/item?id=17076222
Note however that the fine tuning stage adapts to the target corpus - it just doesn't require starting from random weights (so it's orders of magnitude faster).
Instead you either have to roll-your-own models in-house (which defeats the whole point of using a ready made cloud solution) or deal with whatever accuracy you happen to get from those APIs.
IMHO this is an area where you can make some serious competitive headway in commoditised AI/ML. Do all the heavy lifting of pretraining and give your customers an API to "fine-tune" with. Who is currently doing this?
Good guy Jeremy. Works hard during day, open sources it at night
Train this transformer model on a good amount of text (or grab a pretrained model), and then, with minimal fuss and very little tweaking, you can repurpose it to obtain state-of-the-art (or near state-of-the-art) results in a wide range of tasks, from document classification to textual entailment to semantic similarity.
This stands in contrast to prior approaches that involve much more tweaking and/or careful discriminative finetuning of the pretrained model, such as as Jeremy Howard and Sebastian Ruder's also-impressive ULMFit.[a]
The main downside to this new approach is that pretraining takes a long time.
Anyone working on ML/DL/AI with text should take a look at this, right now.
UPDATE: See Jeremy Howard's comment here: https://news.ycombinator.com/item?id=17288320
Another excellent example I came across recently (it also happens to be about unsupervised pretraining and transfer learning) https://github.com/bfelbo/DeepMoji
It's an absolute joy working with such papers and I suspect one of the best ways to get people to actually pay attention to your work in an era of Arxiv Sanity Preserver.
Wow! I enjoy playing with neural networks but this kind of thing reminds me that I'm not really doing deep learning...
I have no idea how researchers could have the patience and confidence to wait that long for a result. In my own (small-data) work, I get frustrated if it doesn't converge in half an hour.. I constantly end up Ctrl-C'ing and tweaking things if it doesn't behave as expected, or appear to be continuing to improve.
I enjoyed several talks at NACL 2016 that referenced the ROCStories data, but didn't really have the (personal, not work) compute power to do much. OpenAI's nice contribution fixes that.
Also interesting that one of the fundamental problems the authors note is "The limits and bias of learning about the world through text", which is essentially a Godelian incompleteness problem. One could say the reverse also applies to embodied/visual data, and a good argument for studying established literature in the abstract.