Automatic Summarization in Medium
medium.com
medium.com
The biggest part of Medium's value proposition, to me, is manual curation. While the included summarizations are technically accurate, they're stilted, like they're being spat out from an algorithm (because they are!): the exact opposite feeling that you want when you're being gently told from a site that they're editing and proofing and including all of these human flourishes.
Perhaps it would make the most sense to have an rss reader or similar with the summarizer built-in. Or a page with links to medium articles (plus others) with summaries.
Edit: Saw the comment below asking if we could see your thesis, which I would be very interested in especially if you don't plan on sharing all of your code.
Basically, tools to help the Medium staff scale the white glove nature of the site.
- There is no list/bill of demands on the table that
can be negotiated.
- This is great for some of the effects the protests will
bring to government leaders. (this one really didn't make much sense)
- In less than a week a movement was formed that led
hundreds of thousands to the streets, without any kind of
leadership, warning or prediction.
- There is an explanation for all that.
- Despite all the anarchy, there is logic behind all this
that we are experiencing.I'd love to see this plugged into tldr.io, so articles on HN could be automatically extracted + summarized -- and later improved by real humans, as needed.
Like a Circa News app, but for the web at large.
Why would you read the article if you've just read the cliff notes? Previews should entice, not give away too much.
Movie trailers look awesome because it's the best the film has to offer, making the full length feature look pale in comparison, which is why I try to avoid them at all cost.
For writing as short as articles on Medium, it might be useful to compare the algorithm to a naive version that pulls the first two sentences from each paragraph.
[1] See, e.g. http://www.amazon.com/Style-The-Basics-Clarity-Grace/dp/0321...
I did a bit of testing of the java version, and it was pretty competitive with commercially available summarizers at the time.
I hope this kind of technology sees the day but I'm very skeptical about it working on general-purpose content and not just structured content such as news, as it does today.
Abstraction combines huge portions of two young research fields -- NLP & NLG (Natural Language Processing & Generation). NLG is even harder than NLP, and less researched. Without good NLG algorithm for presenting summary, you can't have more humane summaries.
Extraction simply takes sentences (or some portions of them), ranks them and presents a few best results.
Two years ago, I was at presentation of PhD about text summarization. There I've figured out that you can make fair summarization algorithm in a few hours. Here is an prototype: https://bitbucket.org/ivan444/textsum/src/1d09b0f4f72a60903d... Dirty prototype code, it took me just about 10h of work to prepare dataset, think algorithm, write program and tune it (this works only for Croatian language, if you want other language, you'll need to get list of function words for that language -- http://en.wikipedia.org/wiki/Function_word ). There is also java version of text summarizer (somewhere in repository) and simple tool to get clean, article-only text from any page containing some longer texts (it isn't tuned well, I didn't spent more than 1h of work in it, so I don't expect it works well).
Algorithm is simple: (1) break text into sentences, (2) extract features, (3) calc features score and sum them, (4) present ranked sentences (and, later, choose a few best).
Used features: normalized number of words, type of sentence (declarative, interrogative, exclamatory), order score (give first sentence a boost, as usually first sentence is the most important one), ratio between number of function words and all words (function words are words without semantic content; there is a fwords.txt in a repository which contains ~700 Croatian function words), normalized sum of three minimum TF-IDF scores (document = sentence).
I don't know the state of the code (it is more than a year old code), but anyone is free to use that code for anything they like.
Nevertheless, I do agree there's still considerable challenges when trying to perform text-to-text generation which involves trying to combine NLP and NLG together to abstract, interpret, and then summarise unstructured free text.
http://www.lexically.net/downloads/version5/HTML/definition_...