Stanford Natural Language Parser
nlp.stanford.edu
nlp.stanford.edu
I haven't seen too many things for NLP on .NET. I have run across a ton things on Python or C++/C. Would love to see more inn .NET.
Buffalo buffalo Buffalo buffalo buffalo buffalo Buffalo buffalo
Stanford fails miserably on this one, even at the tagging step:
Buffalo/NNP
buffalo/NNP
Buffalo/NNP
buffalo/JJ
buffalo/JJ
buffalo/NN
Buffalo/NNP
buffalo/NNP
... but then, most humans also find that sentence really hard to parse.Check it out:
https://en.wikipedia.org/wiki/Buffalo_buffalo_Buffalo_buffal...
New York oxen (that) New York oxen bully (themselves) bully New York oxen.
So I never thought there's something inherently difficulty in the syntax of the sentence- I thought it's just that unexpected third use of "buffalo" that makes it so strange.
Of course, now that I think of it, what you say makes sense also. For an NL parser, the variant use of "buffalo" shouldn't be an issue. You could even say that lacking the context of what a "buffalo" really is, a parser should even find it easier to parse the sentence than a human, who must pause to try and understand what it means.
Maybe it's not such a bad sentence to test parsers after all.
"History lamps traffic people house cry Michigan algebra" also has the same parse tree, but if I utter this, people will probably think I need medical attention.
I put the following phrase
> I nut on my girlfriends face.
and the parser correctly tagged 'nut' as a VBP (i.e. Verb, non-3rd Person Singular resent) i.e. that 'nut' in this context is a verb, and not a noun.
Here's what the tagging returned (all words with tags)
I/PRP nut/VBP on/IN my/PRP$ girlfriends/JJ face/NN ./.
Tag Descriptions: http://stackoverflow.com/a/1833718/325521
I wonder how the other NLP tools parse that sentence. These are the results:
* Spacy [1]
Buffalo/NNP
buffalo/NN
Buffalo/NNP
buffalo/NN
buffalo/NN
buffalo/NN
Buffalo/NNP
buffalo/NN
* NLP4J [2] Buffalo/NNP
buffalo/NNP
Buffalo/NNP
buffalo/NNP - pos2=CD
buffalo/NNP - pos2=CD
buffalo/NNP
Buffalo/NNP
buffalo/NNP
Well, that sentence really makes NLP tools confused.[1]: https://spacy.io/
How difficult is it to use NLP to do keyword extraction for non-English languages? Any pointer to get started? (not NLP in general, but how to tailor it for non-English)
Two examples that easily come to mind: TF-IDF [1] and TextRank [2,3].
These are good to get started, but if you want to know more about the state of the art, searching on Google Scholar for "keyword extraction $language" is your best bet (and maybe "keyword extraction overview", also in Scholar).
[1] http://www.joyofdata.de/blog/tf-idf-statistic-keyword-extrac...
[2] http://digital.library.unt.edu/ark%3A/67531/metadc30962/m2/1...
You could use this Dockerfile to run a local copy if you're interested: https://gist.github.com/charlieegan3/910276eef0f8658b44b42af...
https://demos.explosion.ai/displacy/?text=The%20complex%20ho...
vs.
https://demos.explosion.ai/displacy/?text=The%20complex%20co...
Their results for even the most simple english sentences are horrible, often with multiple mistagged words.
As soon as you add relative clauses, they break down entirely.
I've used the Stanford parser for university work (at Masters level). It was part of the pipeline for a sentiment extractor. I didn't get the feeling it performed badly, quite the contrary- the parses it generated were quite useful as features in my extractor. It had this rare feeling of a tool that you can actually use to do something interesting and useful.
Then again, I do have my expectations set very lowly, for this sort of thing, especially after completing my Masters thesis (on grammar induction). Language learning is hard.
Any every slightly more complicated sentence was completely misparsed by CoreNLP and ParseyMcParseface. As I mentioned before, as soon as you introduce relative clauses it breaks down.
It's hard when your research was supposed to discuss how to better handle the common knowledge problem of NLIDBs, but you have to spend a lot of time just to get a kinda useful parse out in the first place.
Plus, there is a part-of-speech tagger that provides an intermediate layer of nlp analysis, which makes a backtracking involving a change of part-of-speech («houses» being either a verb of a noun) quite harder (yet not impossible).
POS taggers improve parsers performance dramatically but usually are an indication that the parser have limited backtracking capabilities.
You can try it with «the complex houses married soldiers», houses is erroneously marked as a noun.
the/DT
complex/JJ
houses/NNS
married/JJ
and/CC
single/JJ
soldiers/NNS
and/CC
their/PRP$
families/NNS
But, honest, that's a rotten way to try out NL parsers. Nobody claims Stanford's, or anyone's, parser can handle that sort of thing. It throws human language organs because it's specially engineered to do that. And if it can confuse humans, there's no program that it can't confuse, except for one where the right parse is hard-coded.It's an unfair test, is what I'm saying.
E.g. AlchemyAPI has a cool demo, where they can detect names of diseases, cities, etc. I suspect the parts of speech would improve the accuracy for tagging words that have multiple meanings.
I used the disease bit on AlchemyAPI for https://www.findlectures.com, to make facets for a health topics: https://www.findlectures.com/?p=1&class1=Science&category_l2...
Just referencing the old Chomskian bit about statistical correlates and how relevant it is to understanding the fundamental principles behind the expression of language (rather than merely its syntax).
I guess it was a bad joke gone over. ah well
Unfortunately with zero comments.
Nice example of the application of recent NLP thesis topics to parsing problems, though. Quite suitable as documentation and reference for, e.g. future work on SpaCy - which is much more usable in productive contexts.
However, speed wise, its only slow in my experience, if you try to integrate via a shell script - if you run it as a service and make http calls it's been amazingly fast in my usage.
I believe they include a "service" mode as part of the package now, but for a project a little while ago I forked & extended a project that provides a JSON/XML http service: https://github.com/Koalephant/StanfordCoreNLPHTTPServer
(ROOT
(S
(PP (IN In)
(S
(VP (VBG keeping)
(PP (IN with)
(NP (PRP$ his) (JJ insurgent) (NN campaign))))))
(, ,)
(NP (NNP President) (NNP Trump))
(VP
(VP (VBD dispensed)
(PP (IN with)
(NP (NNS appeals)))
(PP (TO to)
(NP (NN unity))))
(CC or)
(VP (VBZ attempts)
(S
(VP (TO to)
(VP (VB build)
(NP (NNS bridges))
(PP (TO to)
(NP
(NP (PRP$ his) (NNS opponents))
(PP (IN in)
(NP (PRP$ his) (JJ inaugural) (NN address))))))))))
(. .)))