How it works — imagine you’re having these sentences:
“Acorn is a tree” and “acorn is an app”
You essentially keep record of all word to word relations internal to a sentence:
- acorn: is, a, an, app, tree Etc.
Now you repeat this for a few gigabytes of text. You’ll end up with a huge map of “word connections”.
You now take the top X words that other words connect to (I.e. 16384). Then you create a vector of 16384 connections, where each word is encoded as 1,0,1,0,1,0,0,0, … (1 is the most connected to word, 0 the second, etc. 1 indicates “is connected” and 0 indicates “no such connection).
You’ll end up with a vector that has a lot of zeroes — you can now sparsify it (I.e. store only the positions of the ones).
You essentially have fingerprints now — what you can do now is to generate fingerprints of entire sentences, paragraphs and texts. Remove the fingerprints of the most common words like “is”, “in”, “a”, “the” etc. and you’ll have a “semantic fingerprint”. Now if you take a lot of example texts and generate fingerprints off it, you can end up with a very small amount of “indices” like maybe 10 numbers that are enough to very reliably identify texts of a specific topic.
Sorry, couldn’t be too specific as I’m on the go - if you’re interested drop me a mail.
We’re using this to categorize literally tens of gigabytes per second with 92% precision into more than 72 categories.