Automatic categorization of text is a core tool now
abe-winter.github.io
abe-winter.github.io
The world has been sleeping on leveraging NLP, given the incredible amount corporations spend paying people to read and summarize text (answering questions like "How many people are having a bad search experience?", etc). The recent leaps in NLP make for many opportunities, and many products are coming to market that allow even non-technical folks to train and utilize text classifiers. If you're thinking about getting into NLP, now is definitely the time to do it.
Source: founded taggit.io
Duplex gave me shivers. Enough that I ended up wondering what would happen if I logically extrapolated from that: http://chir.ag/201812180030
The thing that worries me is not strong AI or evil-AI but rather the selective use of AI by humans who choose to create/remove barriers that were unthinkable just a few years ago. AI doesn't need to actually be intelligent, just pass off as close enough to an average stranger.
The scary part is that a large subset of people here will read that as a instruction manual and not a warning.
With a small corpus these would run into similar difficulties but would likely perform much better than RNN/CNNs trained on the same corpus. Regardless of whether you use NN or something else your approch is the right one: start with model trained on a larger but similar corpus and then use the smaller one to modify the transition probabilities.
Leaving out the "Deep" stuff will shed popular clicks on your post, probably a lot, but would get the job done with less cumulative effort. Important if for some reason you cannot use third party models. The stochastic grammars I mentioned are a lot less fiddly to train than * to vec
This approach to generating philosophy is the exact opposite to the philosophical tradition - there's no effort to determine if any of this stuff is "true" in any sense, there's no traditional formal reasoning model behind it; just if it sounds good. I expect to see it weaponised on political billboards and youtube videos this year.
In particular, “the limits of my language are always your limits” is a horrendously mangled Fascist version of Sapir-Whorf.
Two years ago the deep learning frameworks were difficult to use. At one point I even hacked together a CNN using background subtraction function in OpenCV at one point to simplify my code. I remember coding up a neural network in OpenCL and it took a couple months to verify it worked correctly (several years ago).
Now, I want to write a blog series on text classification (should be out in a couple days) and the actual coding takes 20 minutes, runs way faster and is 20 lines of code.
Neural networks are approachable, research is moving faster than most can keep up, and the accuracy of models is
As mentioned in the article, there’s also pre-trained models, but IMO that’s less important. It's the ease of access that's really the killer.
[1] https://medium.com/capital-one-tech/why-you-dont-necessarily...
(The specific problem is having a lot of tuples of (lat, lon, placename) and I want to build something that can canonicalize all of the placenames. It's different enough from common ML problems that I'm not sure where to start.)
For example, consider this guide for Tokyo's ports: https://www.kaiho.mlit.go.jp/03kanku/h22houkaisei/sozai/guid...
This one happens to be well documented, but most don't seem to be, and in any case many of the port labels in the movement data bear a somewhat hazy relationship to official labels.
I could do it all manually, of course: look at the data, figure out where ships stop, collect all the labels they apply, build a regex pachinko machine. But I'm wondering if I can bootstrap my way to a good text location classifier.
If single or minimal-word phrases are enough (e.g. 'kws' 'kei' 'hei' 'keih' etc), something like word2vec. Except instead of training to predict distance between words, you're generating an embedding between words and ports.
If more complex modeling is required (e.g. 'tokyo tuesday kei then chiba'), something like LSTM [1].
The lat-long data you'd want to squish into a closest-port training set via geospacial math. Haven't done anything in the space, but I'd imagine chewing through lat-long legs and computing closest-approach distances to various ports (that is, to canonical port lat-long). Optimized for computation and space. Probably compressed down into a "visit / not-visit" binary feature. Might take awhile, but the math seems straightforward for this bit.
[1] http://colah.github.io/posts/2015-08-Understanding-LSTMs/
There are a lot of factors contributing to widespread adoption of NLP -- certainly the availability of great tools like PyTorch/Keras, SpaCy, and Gensim, and more broadly growing technical competence and awareness among teams that traditionally handle text.
But the single biggest factor is that people are generating massively more purpose-driven text in interactive channels. A generation that grew up texting their friends is becoming core consumers and workers in every industry. People drop text in team chat, in project management tickets and pull requests, when contacting support, and more. Plus, we break our text into meaningful snippets and add our own metadata to make them more readable (and usable) for other people. Language and tools are co-evolving very rapidly in ways that make them more accessible.
As a human-and-organization-augmentation nerd, I'm incredibly excited. We are seeing so much more about how groups of people self-organize than we ever used to, and at the same time that our tools are starting to become capable of understanding us in the ways that we communicate most conveniently.
And I hate to pick at specific points that are more cursory to the article, but I can't help myself:
> Duplex feels like an extension of the gig economy; this is about college grads not wanting to waste their time negotiating with grunts.
I understand the sentiment here, but I don't think that's what's happening here. If there were a better way to do this sort of thing without "hacking" the human interaction I think we'd do it. I think Duplex has come about because many companies still insist on wasting everyone's time by forcing people to make time-consuming phone calls to perform very simple tasks. I don't think is about helping college grads not have to interact with "grunts".
I think that Jeremy Howard's results with FastAI do strongly suggest that deep learning helps. Specifically, it can be very quick to train a language model with your own data... so quick that he doesn't even bother starting with pre-trained weights.