Stanford Scientists Put Free Text-Analysis Tool on the Web
engineering.stanford.edu
engineering.stanford.edu
The paper is here: http://hci.stanford.edu/publications/paper.php?id=279
Great idea though!
Will there be an API available? Or will I have to get creative with Form POSTs :).
[0]: https://github.com/peterldowns/bookshrink [1]: http://www.bookshrink.com/
Don't be so hard on yourself. I review papers for CS conferences and just the fact that you used TF-IDF weighting puts you well above the average.
[0]: https://github.com/bumptech/tl-dr.js/blob/master/tldr.js
""" SPLIT INPUT INTO SENTENCES"""
god_awful_regex = r'''(?<!\d)(?<![A-Z]\.)(?<!\.[a-z]\.)(?<!\.\.\.)(?<!etc\.)(?<![Mm]r\.)(?<![Pp]rof\.)(?<![Dd]r\.)(?<![Mm]rs\.)(?<![Mm]s\.)(?<![Mm]z\.)(?<![Mm]me\.)(?:(?<=[.!?])|(?<=[.!?]['"]))[\s]+?(?=[\S])'''
Be advised that one of the nice things about python reg-exes is that they allow in-line comments (and naming of groups), if compiled with the verbose-flag:
""" SPLIT INPUT INTO SENTENCES"""
verbose_regex = r'''(?<!\d) # I can't actually tell
(?<![A-Z]\.)(?<!\.[a-z]\.) # what you're doing here...
(?<!\.\.\.)(?<!etc\.) # Is this one big group, or is
(?<![Mm]r\.)(?<![Pp]rof\.) # it several groups, with
(?<![Dd]r\.)(?<![Mm]rs\.) # different prefixes?
(?<![Mm]s\.)(?<![Mm]z\.)(?<![Mm]me\.) # Clearly, it's got something to do
(?:(?<=[.!?])|(?<=[.!?]['"]))[\s]+?(?=[\S])''' # with
# not matching the dot at the end of Dr.
# or as part of an ellipsis as the end of a
# sentence? But my point was that if such a
# regex is built-up and tested with comments
# it can be quite readable
god_awful_regex = re.compile(verbose_regex,
re.VERBOSE)
# continue here...
http://www.diveintopython.net/regular_expressions/verbose.ht...It pretty good for what it does.
Perhaps it does better with different hash-tags (there's a lot of bitter irony associated with #nsa -- possibly more than average for other tags) ?
They used it as a marketing pitch for AWS/EC2
Username response was here: http://www.etcml.com/jobs/8188 The only thing I would add is an overall score for how positive/negative/neutral a text is.