TensorFlow Code for Google Research's BERT: Pre-Training Method for NLP Tasks
github.com
github.com
Context to understand the importance of this release: http://ruder.io/nlp-imagenet/ Though not named in the post, BERT is part of the same family of models as these.
Would BERT help me by first enabling me to transform the input subject lines into vectors in a high dimensional vector space which could then be the inputs into a relatively shallow network that does the classification?
For the experiments in paper they actually fine-tuned BERT on the downstream task, but I reckon you'd get acceptable performance by just keeping it fixed and using its outputs as features for a shallow classifier.