Learning Python while processing raw text: The NLTK book
nltk.org
nltk.org
Plain text files and not tied to a library.
It's machine learning as a service: simple API calls to train, cross-validate & classify your data. We'll also be at PyCon this week in Santa Clara so come on by.
Currently in private beta, but we're ramping things up quickly.
There are very good open source NLP libraries like Stanford NLP, OpenNLP, and NLTK. I think the opportunity for business might be in building custom language models for customers based whatever domains they deal with (e.g., medicine, housing, real estate, etc.)
I just had the experience, that the auto-slider slid away, while reading your explanation on the third step. I manually had to click, to get the text back.
Maybe you should A/B-Test, if a manual slider is better here, as it is independent of your users reading-speed. Just as a suggestion. ;-)
Edit: I'm obviously not saying NLP isn't useful, just that web scraping is more immediately useful. With NLP, besides learning the concepts, you have to find a source of raw text that's been unprocessed and yet contains something of real world value. With web text, you just have to collect what someone already thought was valuable to publish and find insights through the aggregation. It seems to me that the latter scenario is easier to grasp, with NLP being useful for going beyond what others have gathered and published.
But there is value in both, depending on your objectives. I find web scraping trivial, and mining the text hard, hence my interest in NLP and machine learning.
[edit] This coming from a text guy, who recently started down the path of python and is hooked ;-)
http://www.clips.ua.ac.be/pages/pattern