Show HN: Subreddit classifier using scikit-learn and a high-performance Go proxy
ioloop.io
ioloop.io
https://www.findlectures.com/?p=1&class1=Technology&category...
For the first iteration I wrote heuristics based on a list of language-specific subreddits. The technique this blog post describes is the next logical step, so I'm thrilled about this write-up.
-- can't find an RSS feed for your blog -- ioloop.io has a blank homepage -- ioloop.io/blog gives a 404
So where's the best place to keep checking?
Thanks!
What I really like about this technique is that I can play with my scikit-learn while being offline, which seems to go hand in hand with the holiday travels ahead.
The focus of this post is the GoLang proxy used for the caching. It's actually used in a CI / CD environment, but I'm finding it incredibly useful for a whole variety of tasks.
Regarding the precision metrics, you can find all the information required here http://scikit-learn.org/stable/tutorial/text_analytics/worki... to compute the precision.
In the scikit-learn article the classifier scores over 90% precision, so I'd expect it would be possible to do the same.
I'll be posting more about this over the xmas period, so I'll write a part II, where I compute the precision metrics.
Another thing worth noting is that it's not caching the HN comments, while it is using the reddit comments. Despite using tfidf, this still completely skews the results towards reddit as opposed to HN. So that's something else any interested reader can look into.
Thank you for pointing this out.
https://bigishdata.com/2016/12/05/classifying-amazon-reviews...
But results and how well the classifier performs really just depends on the quality and amount of training data you have. So would be interesting to see how this does if you can get a bunch more data from each of the subreddits and have some more test examples!
I'm just getting started with experimenting in this topic, and it's great to see something that isn't about classifying flower petals that is approachable!
I know it's not the focus of the post, but was there any particular reason why you went with the MultiNomialNB classifier? I've been getting pretty good results recently with LinearSVC which seems to be a lot faster and in my case a bit more accurate too.
An interesting metric for a future post might be how your proxy compares with scrapy + httpcache middleware.