This was a long project that required a lot of work. The quality of sample data hugely impacts the results, and so I built a (mostly automated) pipeline to clean, annotate, and improve the data. I have no doubt that the companies that build ML training sets will hold an edge in the future.
I'll be happy to answer any questions if there are any. I learned a lot from this. :)