BERT: Pre-Training of Deep Bidirectional Transformers for Language Understanding
arxiv.org
arxiv.org
> The code and pre-trained model will be available at https://goo.gl/language/bert. Will be released before the end of October 2018.
* pre-trained - train on lots of language modelling data (e.g. billions of words of wikipedia) and then train on the task you really care about but starting from the parameters learnt from the language modelling task.