Algorithmic tagging of Hacker News or any other site
blog.algorithmia.com
blog.algorithmia.com
There are other words in both examples that I personally would not use as tags, but I can't really say they would be universally not-useful. I think a vast improvement could be made just by having a dictionary blacklist filled with things like these - from this tiny sampling contractions seem to be a big loser.
You can also seed LDA with a whitelist of words which we didn't do either - again all in the name of a quick and dirty solution to show.
Glad you liked it!
Didnt know about tf-idf thanks for the tip.
and this doesn't really strike me as much of a victory for the idea that it's just the implementation of an algorithm being the sticking point in practice.
The demo we show here is a version of how we used our platform to generate tags for all entries in our API by combining algorithms that existed already in Algorithmia. (modified for performance over quality due to the volume that HN would bring).
Cheers.
USA
NSA
DOJ
TCP
POS
etc.
I have been doing some research towards automatic tagging lately, and I found several Python project coming close to this goal : https://pypi.python.org/pypi/topia.termextract/ , https://github.com/aneesha/RAKE , https://github.com/ednapiranha/auto-tagify
but none of them is satisfying, whereas Algorithmic Tagging of HN looks pretty good.
I have been trying to implement a similar feature for http://reSRC.io, to automagically tag articles for easy retrieval through the tag search engine.
For example, try this on restaurant reviews like http://www.yelp.com/biz/el-gaucho-seattle. I get these tags:
steak reviews seattle food service gaucho restaurant review
Not useful, right?
The current state of the art would use much more sophisticated NLP for generating POS tags and use sentiment analysis. For example, check out MSR Splat at http://research.microsoft.com/en-us/projects/msrsplat/defaul....
This could be really useful in ecommerce for creating search keywords for category pages. The noise in the results matters not, so long as it gets 'T-Shirt' and someone searches for 'T-shirt' then all is well and good.
Are you looking to plug what you have into something such as the Magento e-commerce platform? The right clients could pay proper money for this functionality. It is something I would quite like to speak to you about.
after a while the system converges to a very useful structure and new members can see correctly tagged articles and the system learns their interests by itself
do you know anything like this already existing?
"Erlang and code style"
process erlang undefined
file write data
true codehttps://docs.google.com/forms/d/1UeSD11hrjwhsVbbPiv63VZBrEcz...
PS. Screenshot included + it's already in alpha in a company with 100 users.
Call it lobste.rs 2.0
Here is the trained topic model (Nov. 30, 2012) with only 40 topics (for file-size mainly) https://dl.dropboxusercontent.com/u/14035465/hn40_lemmatized...
You can load it with Python:
from gensim.models import ldamodel
lda = ldamodel.LdaModel.load("hn40_lemmatized.ldamodel")
lda.alpha = [lda.alpha for _ in range(40)] # because there was a change since 2012
lda.show_topics()
Now if you can figure out what is this file:
https://dl.dropboxusercontent.com/u/14035465/pg40.params I'll pay you a beer next time you're in Paris or I'm in the Valley. ;-)tags tagging hours link doppenhe reply ago lda
looks pretty promising!
Some of these points are related to encouraging users to tag content, but auto-tagging also seems problematic.
To me something more along the lines of entity extraction is more useful because it is a well defined problem, and can be used to improve a lot of other applications.
To understand the utility of tagging, look at some article, read it and then put 3 words that best describes the topics. I bet most others would find human generated tags very useful. Machine generated tags are usually no where close to what humans would generate.
Tags while cluttering the ui do help you find similar content. Still not as good as a good reccomendation system but a decent stop gap measure in some instances.
But doesn't the auto-tagging feature make to much noise for a business use-case? For example, it tags a article of Amazon and includes Google in the tags. White-listing words wouldn't fix this (Google is a whitelisted word if Amazon is).
I don't know about LDA though. Perhaps a proper tag administration would fix this, but then you'd have to remove tags on the go.
Are they using some alternative API that was blessed by HN?
How does iHackerNews show all the comments and everything?
Where are these RSS feeds? I doubt that's how "the pros" so it.
I would think lots of apps are still scraping pages.
> stream rotate type/page font structparents endobj obj endstream