Ask HN: Review ContextSense - deduces context and sentiment for any URL/text
wingify.com
wingify.com
Slightly Negative (0.33)
Extract From the same link (text only as shown below) Highly Positive 0.81!
>The usual German term for the extermination of the Jews during the Nazi period was the euphemistic phrase Endlösung der Judenfrage (the "Final Solution of the Jewish Question"). In both English and German, "Final Solution" is widely used as an alternative to "Holocaust".[16] For a time after World War II, German historians also used the term Völkermord ("genocide"), or in full, der Völkermord an den Juden ("the genocide of the Jewish people"), while the prevalent term in Germany today is either Holocaust or increasingly Shoah.
NLP is H A R D, I know and defining a range value for sentiment even harder.
ContextSense was made to demonstrate our contextual targeting capabilities -- and sentiment aspect was added to avoid traditional blunders of displaying ads on pages/news on catastrophies. Best to try ContextSense with text heavy URLs. Also try it with news items (both positive and negative).
I will be happy to provide API access in case any of you is interested in trying it out.
Few human practices have provoked such deep and widespread outrage as the practice of one human being enslaving another. So why has slavery survived for thousands of years? How did it become so important to civilization? Explore the ways that slavery has been woven into the fabric of societies in America and around the world.
Sentiment: Highly Positive (0.77)
Seems off in cases where the tone of the text is neutral but the subject matter is negative.
An issue we haven't brought up is tags. Current tags for HN: "ago, points, comments, com, hours, hour, hacker, discuss, minutes, tornado." I see some good ones, but it's quite noisy (and it seems to have grabbed the tornado story as a main topic--scraping the title?). This feature seems most useful to me, and most easy to improve.
http://en.wikipedia.org/wiki/C*-algebra
The results are better than I would've guessed, but not really so great (this article is pretty degenerate I think, versus more "human-interest" type stories).
One thing I'm curious out: the list of topics is here:
People (7.88) Algebra (4.47) Research Groups (3.41) Software (2.42) Math (2.03) Functional Analysis (1.53) Journals (1.01) Science (0.71) Publications (0.64) Logic and Foundations (0.6)
...does looking at the content give you any sense for why "People" comes in first?
That being said, this looks really cool and full of potential, now I just need a reason to be auto-generating that kind of meta data.
Anger: 25 Smugness: 50 Vulgarity: 15 Boredom: 5
etc.
I'm not even sure that's a good example or if it's clear, I just think positive/negative is an oversimplification of the sentiment of most sites.
I typically have cookies disable, but I turned them on and it made no difference.