Organizing My Emails with a Neural Net
andreykurenkov.com
andreykurenkov.com
It is a general purpose naïve bayesian email classifier that you can integrate with almost any email system.
They took some of the concepts in the article mentioned here and expanded on them a bit.
For example, they have the idea of "pseudowords"[1] so that you're working with more than just the words in the email. Like html:td for example...it expands to the number html table cells in an email, which might help with choosing a bucket.
I rewrote POPFile in Perl and made it work with generic POP3 servers (hence the name) and later with IMAP and more. It is still actively maintained by a group of people I've never met(!) with the last release in December: http://getpopfile.org/
Well, thinking more about it leads me to tf-idf and naive bayes (of course), at which point you pretty much already have a classifier. So it seems feature selection is learning in itself and defines the maximum accuracy you'll be able to reach ? This is border philosophical but I'd love to read more about these matters. Pointers welcome !
The common practice with a small-ish dataset is to use e.g. the top 10k or 20k most frequent words, but filter out the top 50-100 so most frequent words, as those indeed do not carry much information. A commonly used weighting scheme is TF-IDF (https://en.wikipedia.org/wiki/Tf%E2%80%93idf), which comes included in Keras.
Anyway, this is a cool ML starter project. Keras makes it really easy to do this sort of fast experimentation with a range of different neural networks models.
In short, there are "good enough" rules that require much less processing.
Yes, but I guess you need to train it first. And if you have been bad at categorizing in the first place, you will start with bad training data.
How on earth am I to tell from that visualization how often it mislabels financial emails as personal? Eyedropper the colors and hope the values in the shades of blue follow a linear scale?
A quick curosry glance at the central diagonal tells me that Finance, Personal/Programming, Professional/EPFL, and Group work are the categories of e-mail which are most likely to be categorized incorrectly by the software. Looking at the columns tells me that the Academic and Personal are going to have the most misfiled messages in them.
Yes, but how likely is each mislabelling? There is no scale to indicate how the colors map onto probabilities.
As well as providing a scale, it would be helpful to make the heatmap an annotated heatmap, in which each square is labelled with the corresponding value (perhaps with values below some threshold omitted to reduce clutter).
Example: https://web.stanford.edu/~mwaskom/software/seaborn/_images/s...
(from the Seaborn docs: https://web.stanford.edu/~mwaskom/software/seaborn/generated...)
Edit: Consider another question - If an email has the true label 'Financial', is it more likely to be mislabeled or correctly labelled? I can guess, but without more knowledge of the scale I can't be certain.
As it is, the specifics of how it mislabels his e-mail is not all that interesting to me, as I don't have access to his e-mail and so the specific numbers are pretty much irrelevant.
I don't think you can create custom categories but Google's inbox.google.com does this
Usually you'll end up with a frequency threshold, but it's usually best to trim at the very low end --- like, words occurring 5 or fewer times. Further over-fitting can be controlled with regularization and parameter averaging.
This is pretty straightforward to implement in keras, you just need to supply pre-trained word-vectors weights to your embedding layer.
Embedding(vocab_size, 300, weights = [word2vec_weights])
Where 'word2vec_weights' is a numpy matrix with shape (vocab_size, 300).