Natural language parsing for the web
nlp.naturalparsing.com
nlp.naturalparsing.com
I've been working on this website for the past couple of weeks. Please email me at andrew [AT] naturalparsing [DOT] com if you have any questions or suggestions, or want to use the API.
FYI, right now the API is not using all the capabilities of the Stanford Parser, just the word tagging part. More features will be implemented soon. Let me know if you have any specific requests.
Andrew
It's a POS tagger whose output is extremely close (read: identical) to the Stanford POS tagger.
It's also much faster. I was able to tag a corpus of 300K sentences with this in 15 minutes. With the Stanford POS tagger it took the entire weekend.
Sadly, this tool's license does not allow commercial use and it is not released under the GPL license.
But the top-poster is right: the linked website does part of speech tagging, not parsing.
Providing a wide-coverage parser for the web is still hard. The number of possible parses for long sentences is enormous. Even if a sentence is not ambiguous to us, grammars allow all kinds of ambiguities.
There is a web demo for a Dutch system that is developed in our research group[1], but heap size and time limits are used, to exit gracefully if parsing takes too much time/memory (and sentences with more than 20 words are ignored, since you'll really want to do offline parsing).
As to the time consumption and complexity, that's a known problem of unification grammars (or just any grammar that does a little bit more) - but see this paper by Matsuzaki et al for efficient techniques to speed this up: http://www-tsujii.is.s.u-tokyo.ac.jp/~matuzaki/paper/ijcai20...
There are also many other possible optimizations. E.g. with a left-corner parsers you can exclude left-corner spines that will probably not lead to a probable parse (Van Noord, Learning Efficient Parsing, 2009, and other work). And, of course, reduction of lexical ambiguity by restricting frames using a part of speech tagger.
Still, even in this best-1, optimized scenario, real-time parsing of long sentences is still hard. So, when parsing large corpora we usually apply time and space limits (which is easy to do in Sicstus Prolog, with good recovery).
Thanks for the link to the paper!
Still, I wouldn't call part-of-speech tagging parsing.
Here's some parsing demo: http://nlp.cs.berkeley.edu/Main.html#parsing
Hmm. Maybe sometime.
EDIT: The Stanford POS tagger is more complex and quite a bit slower than anything you'd do on your own. To quantify this, there's methods that are 10x as fast while sacrificing 0.05%-0.2% accuracy. (Or the easy ones that are 100x as fast, but are 1% less accurate - these would be fun to do in JS).
(model loading doesn't work yet for some reason, but you see what it's doing in principle).
This uses a smoothed trigram HMM, so it should, in principle, be a bit better than NLTK's HMM tagger but not as good as serious POS tagging packages (e.g. hunpos, or the Stanford POS tagger)
http://otlportal.stanford.edu/techfinder/technology/ID=24472
The tagger is licensed under the GNU General Public License (v2 or later).
The GPL, like other licenses meeting the 'open source definition', has no restrictions on use -- only on proprietary distribution under nonfree licenses.
http://dhigger.blogspot.com/2009/08/research-project.html
http://workproduct.wordpress.com/2009/01/27/evaluating-pos-t...
"Pride and Prejudice is a good book." becomes "Pride/NNP and/CC Prejudice/NNP is/VBZ a/DT good/JJ book/NN ./." I would have thought "Pride and Prejudice" would be lumped together.
go/VB fuck/NN yourself/PRP
go/VB lemon/NN yourself/PRP
doesn't sound good, whereas
go/VB shave/VB yourself/PRP
is fine. IMO, go should also be a VBP, not a VB.
go/VB fuck/VB yourself/PPL
It's very much in the amount and kind of training data and features used (assuming that the methodology is sound).