Tldr.it - Summarizer for RSS feeds and other web pages (built in 48 hours)
tldr.it
tldr.it
if (!url.startsWith("http://) { url = "http:// + url; }
(Edit: for some reason the closing quotes are being stripped)
For example, I tested tldr.it on a recent HN submission, Programmers: How to Make the Systems Guys Love You (http://news.ycombinator.com/item?id=1800839), and it seems to select the first paragraph or so of the article (depending on whether you select the short, medium, or long tldr).
However, so many authors these days start out with a long, rambling preamble, instead of the venerable Who/What/When/Where/(Why|How), that the first few paragraphs don't necessarily provide a real tldr-type summary.
Which is the case with that particular article - the first few paragraphs don't help me decide whether to read it or not, and the entire article is quite long, with some good nuggets further down in it.
We were discussing this on MetaOptimize recently (http://metaoptimize.com/qa/questions/2815/how-are-search-eng...), but I'm curious to hear about alternate approaches.
It's 5 years old now, but the summaries it generates are competitive quality-wise with most things out there (eg, the MS Word summarizer). Unfortunately I don't have an online demo working atm (like I said - 5 years old)
My algorithm is something I made up, and from memory it works like this:
1) Remove HTML, stem, remove stopwords etc
2) Sort unique words by popularity in the text
3) Split the original text on sentence boundaries.
4) Include each sentence that first mentions the next most popular word, until the summary is the maximum length requested.
Like most things, it's surprising how well a simple algorithm like that works.
There are ports for C#, and Googling just then apparently someone has done a python port too.
http://www.sfgate.com/cgi-bin/article.cgi?f=/c/a/2010/10/18/...
suffered from complete content extraction failure.
The third one seemed to work pretty well, but there were lots of spurious newlines in the output which made it really hard to read.
Nice idea but needs another 48 hours of polish :-)
Maybe you could pipe requests through their API:
In your version you said you weren't happy with the HTML extractor. It's pretty hard to generalize that part, but one technique I found useful was having a flag that told the program to ignore all text until it found the first <p> tag.
In my testing, that removed ~90% of navigation text (although I note you are only looking in <p> tags. I had a flag for that too, but found it was unnecessary most of the time).
Also, I found regular expressions weren't terrible for sentence boundary detection. OTOH, there was nothing like NLTK for Java when I wrote it anyway.
I invested most of my cleverness in actually getting the content out of the page since that's really where the money is for an MVP for this; no content == no summary. :)
It seems that OTS uses a word frequency strategy, so the algorithm is similar or identical to the one I demoed. Interesting.
I have gone through it carefully, and it is clever.
OTS is definitely word freq based.
Would really like to see this. :-) If it has an API I'd love to tie something like this to http://tldrd.com to auto-generate summaries of various texts.
I'm watching this thread with interest because I've been doing some work on building a good summarization engine for a few months now, and it is getting pretty close to ready for prime-time -> summarity.com
What about accepting complete urls as parameter? So that you just need to put tldr.it/ in front of an [URL] to obtain a summary?
Did you show this to reddit yet?
One thing i was thinking about was how to monetize something like this, sadly the only option seems to be ads or amazon/etc related to the content...
And yes monetization is definitely on my mind. I have a few options. The first is as you mentioned Amazon stuff for the arbitrary URL's. Then there's also injecting ads into the summarized RSS feeds. Then there's also adding users and letting them have enhanced functionality for $x a month (i.e., update their RSS feeds on a regular basis, cutting free users back to once a day, make the bookmarklet a premium feature, and so on).
I really think this could be a good product; just not sure if I'll have the time to really develop it.
I agree with drtse4 on that point. Not as many people know about bookmarklets as we let on to believe, at least thats what I found out among my friends. And well frankly this time last year I didn't know what they were either.
Secondly, the domain is quite short; telling people to add tldr.it in front of any url to get a summary is simple and could have better word of mouth advantages then say teaching/telling them to use a bookmarklet or coming back to the homepage and pasting it in the text box.
I have a special case extractor for things like the NYTimes (where their markup makes it really easy to pull things out) and 1-2 other sites. I want to add special cases for more of the popular sites on the 'net. But seems like not a ton of sites really need it unless their markup sucks (I'm looking at you BBC and CNN).
in short, you upload a text and it extracts: facts, summary, keywords, index and stuff, with the most important info.
in most of the cases pretty cool, but sometimes was slow in processing.
We think this results in a much more relevant summary - but we'd love to hear about how our peers compare to us for your particular summarization needs.
Also, we're starting to open up our API for select developers if you're interested in integrating summarization, fact extraction etc into your own application. Tweet @topicmarks if you want to be involved.
Roland (CEO Topicmarks)
odd content...