The Berkeley Document Summarizer: Learning-Based, Single-Document Summarization
github.com
github.com
Do any Python NLP engineers integrate Java tools in their setup, e.g. via Jython? How does that work for you?
Our current production back-end doing NLP stuff is written in Java and most third party libraries it uses are also written in Java. It was initially written in Python, but at one point we realized that most libraries we use are in Java and Python is just moving data between them. The choice of these Java libraries wasn't driven by any love towards Java either. By the time they were just a fair bit more advanced both in feature set and performance than their Python counterparts. One example is Stanford Parser and CoreNLP toolkit--up until Parsey McParseface it was the most accurate parser, and CoreNLP toolkit had more features (that interested us) than Python's NLTK.
>>> from parsedatetime import Calendar
>>> cal = Calendar()
>>> cal.parse("The day after tomorrow")
(time.struct_time(tm_year=2016, tm_mon=10, tm_mday=3, tm_hour=9, tm_min=0, tm_sec=0, tm_wday=0, tm_yday=277, tm_isdst=-1), 1)
Yes I've used Java equivalents before and they're impressive, but I don't think Python is inferior in this regard. For me, the advantage of Python is mainly being able to easily run small sections of project code in an IPython shell while developing that make the development workflow so much faster and more pleasant, rather than the quantity of code.See http://www.eecs.berkeley.edu/~gdurrett/ for papers and BibTeX.
berkley's summarizer takes syntax trees of sentences and cuts off unimportant adjectives (or these unimportant brackets, or comma separated explanations etc.) directly from sentences (since operations are done on syntax trees the shortened sentence is gramatically correct, if parsed syntax tree is correct).
it also uses coreference resolution. (for example, if you were to take the last sentence without context you wouldn't know what/who "it" is, they build up a map of named entities [name - Berkeley Document Summarizer] and replace any pronouns (if necessary) with the names to which they refer.
summarizing quality should be much better than the reddit bot.
- Length of sentence without stopwords
- Distribution of words across the article
- Words shared with the title
- Position in the article
- Position in the paragraph
- Does the sentence address the subject in third person? (avoid)
- Does the sentence contain direct speech? (avoid)
Stuff like this. It's a bit of a Mechanical Turk, really.It's unclear but still seems Really promising.
No shit.
Pages 3,4 of the paper http://www.cs.utexas.edu/~gdurrett/papers/durrett-berg-klein...
There are some examples of how the sentence compression works, but no complete automatic summaries that I can see.