Show HN: TextBlob, Natural language processing made simple in Python
textblob.readthedocs.org
textblob.readthedocs.org
Yay #2: an actually interesting programming-related article on HN. These get rarer every day, losing their place to gossips about what Snowden remarked following some or another NSA official's remarks about Snowden's even earlier remarks.
On the one hand, I agree with your yay #2(and your yay#1, of course. TextBlob looks great). I think you're right. I have many venues to discuss NSA issues and few venues to discuss startup/programming stuff. I like having a venue that is typically devoted to such stuff.
On the other hand, I'm not sure that saying "this isn't about Snowden" on tech-related articles that don't involve Snowden is the solution. Why bring him into the conversation when we're talking about Python?
Did you look at https://lobste.rs/ ?
I also feel programming stuff is depleting on HN, however I can relate to all the talk about Snowden on a forum like HN though.
I'd love an invite if you happen to have one. Contact info in my profile.
Just as words of encouragement for more programming-related content. I'm venting off steam, really. In the past few weeks I've been spending more time on /r/programming than on HN, something I could not imagine a year ago.
Honestly it's frustrating to me that a meta-discussion like this is necessary here. I think you're right, I just wish I didn't have to say it.
TextBlob is probably just using the en module, I would suggest everyone take a look at the other modules in particular the web module should you be doing any light data scraping. It has nice wrappers around BeautifulSoup and Scrapy among others, jumping into BeautifulSoup and Scrapy can be daunting for beginners.
One issue though is that it seems to choke with certain characters.
For instance the character £ it seems to complain with this error message:
>>> TextBlob("£") Traceback (most recent call last): File "<stdin>", line 1, in <module> File "/home/eterm/nlp/local/lib/python2.7/site-packages/text/blob.py", line 340, in __repr__ return unicode("{cls}('{text}')".format(cls=class_name, text=self.raw)) UnicodeDecodeError: 'ascii' codec can't decode byte 0xc2 in position 10: ordinal not in range(128)
(My source data is my own HN comments, it's funny doing sentiment analysis on them, seeing how objective or subjective it thinks my posts are as well as generally if I'm cheery or miserable.
(My end game is to produce an HN reader which only shows positive comments and news to reduce the amount of reading I do. ;))
It gets it mostly right, except the occasional hiccough, one of which is this following passage, which stood out as my most subjective post (1.0 on subjectivity!): "" Factorisation is unique, the addition of 3 primes is not.<p>e.g. 29 can be written 5 + 11 + 13 or 3 + 3 + 23<p>So even if it were a difficult operation to reverse addition of 3 numbers, it would be made easier by collisions. ""
I'm left stumped as to why nltk thinks this is not only subjective but a 1.0 completely subjective post!
I have no idea how sentiment analysis works though.
Sorry, but I think this thing is very much overrated by the HN crowd. There are many such libraries and this one adds exactly nothing. I also don't see how this is easier to use than, lets say, Pattern.
Try and add new functionality. One new functionality could be to use an ontology to calculate the distance between two words. Then you can do other cool things with that and place it in your module.
I'm curious to see exactly how it works and so I'll certainly check out the source when I have a bit more time. Thanks for posting this.
Edit: and it also uses pattern.
For example, an analyst might use sentiment analysis to see whether Facebook posts about a product are "positive" or "negative" in tone.
As another example, I hacked together this online sentiment analyzer using TextBlob: https://textfeel.herokuapp.com/
See also: NLP (Wikipedia): https://en.wikipedia.org/wiki/Natural_language_processing NLTK (a python library for NLP): http://nltk.org/ Twitter opinion mining using pattern: http://www.clips.ua.ac.be/pages/pattern-examples-elections
from the features list it doesn't seem to.
What you're referring to is text generated using a Markov Chain algorithm. This will generate text that seems at first glance to be human generated. On closer inspection you'll find that it only follows common linguistic patterns, the actual content is gibberish.
"Today, the Dow hit a high of 16,200, marking the first time it has crossed the 16,000 barrier. blah blah blah, etc"
Basically, use data points to create a market overview where readers wouldn't know that it was computer generated. That's one idea.
(posting before I commentstalk you to confirm, this is just as much to test my memory as anything else)
edit: maybe not...damn, I swear I've seen this exact phrase in several NLP threads
[1] https://textblob.readthedocs.org/en/latest/advanced_usage.ht...
Any thoughts or relevant benchmarks you would like to share about its speed?