Show HN: TL;DRizer - an algorithmic summarizer webapp/api in java (weekend hack)
tldrzr.herokuapp.com
tldrzr.herokuapp.com
From:
Etiam tincidunt dolor at est sagittis a rhoncus turpis egestas. Integer elementum erat nec nisi molestie eu tempus magna feugiat. Mauris eu ligula et ligula vulputate tempor. Etiam vel lectus et mi vulputate rutrum. Cras libero ipsum, rhoncus at accumsan id, adipiscing iaculis turpis. Cras vel metus nec enim consectetur aliquet vel nec nunc. Proin at mauris purus. Nullam nulla dui, interdum nec pharetra sit amet, vulputate a lectus. Nunc vulputate pellentesque purus at euismod. Nam in justo quis ante porttitor pellentesque. Quisque quis purus a magna scelerisque egestas quis id sapien. Ut non felis sit amet ipsum sodales placerat. Proin nibh massa, sollicitudin et posuere a, placerat convallis magna. Duis lacinia mauris sit amet ante pharetra sed bibendum lorem euismod.
To:
Mauris eu ligula et ligula vulputate tempor. Etiam vel lectus et mi vulputate rutrum. Cras libero ipsum, rhoncus at accumsan id, adipiscing iaculis turpis. Nullam nulla dui, interdum nec pharetra sit amet, vulputate a lectus. Duis lacinia mauris sit amet ante pharetra sed bibendum lorem euismod.
So much faster to read, I never had the time to read through all those design mockups!
Here's a preview of mine (http://www.textteaser.com/ui/article?link=http%3A%2F%2Fwww.p...). Go to its home page to read more news. It caters Philippine news and will soon enters alpha stage. I'm planning to open up the API or open source it. HN, which is better? The API is ready, registration is the only thing that it lacks.
You can try the API here: http://api.textteaser.com/api/?url=http://www.theverge.com/2...
Just replace the url parameter with the URL of what you want to summarize. Some URLs are not tested yet, and may produce errors. :)
I still have other papers, but those can be a good starting point.
In my thesis, NLP is done via statistical approach. It learns from it's previous summaries does have somewhat learning. I don't see any problems combining NLP with Machine Learning. Can you elaborate on this?
Re NLP, I meant to say NLP from sentiment analysis perspective (not sure if that's the right way to put it). So my question was whether you had to extract meaning of one or few related sentences or only process the text in a statistical manner (which you answered).
Back in March, Yahoo bought a startup called Summly for $30 million. Before Yahoo shut it down, Summly was a news aggregation app for smartphones. According to Summly's own Web site, the technology behind the app was "built" by an organization called "SRI International," not by the startup's employees. And indeed, inside Yahoo, Summly is called "Yahoo's Siri." A source close to Yahoo says that CEO Marissa Mayer believes summarization technology is "going to be huge for Yahoo" as it builds "personalized news feeds" into mobile versions of its "core experiences," including Yahoo Finance and Yahoo Sports. The job of implementing this technology at Yahoo will not be given to anyone from Summly, including its young CEO.
--
Edit: Adding this...
Generated 3 Sentence Summary of Gettysburg Address
Four score and seven years ago our fathers brought forth on this continent, a new nation, conceived in Liberty, and dedicated to the proposition that all men are created equal. We have come to dedicate a portion of that field, as a final resting place for those who here gave their lives that that nation might live. The brave men, living and dead, who struggled here, have consecrated it, far above our poor power to add or detract.
edit: some work is needed on the tokenization also. currently, I don't preserve non-period punctuation.
$30 million to have the entire tech sphere spread the word on how Yahoo are evolving and are going to make peronalised summarized on the go news a core feature seems like a bargain (there certainly is a market for something like that I would say -- or at the very least, if implemented well, it certainly would be nice to have feature -- possibly a habit changer).
They no longer seem to me at the very least, the old who are they again company that they once were.
It's a shame they won't be changing their name anytime soon though. That on it's own puts me off a little. Forgive me for the snark.
Why Marissa Mayer Bought A $30M Startup - Business Insider: The deal got a lot of attention because Summly's CEO is 17-year-old Nick D'Aloisio. Acquiring Summly seems to have been an almost incidental side effect of a deal Yahoo made with SRI for a piece of "summarization technology." Until Yahoo bought it, SRI International held equity in Summly. The job of implementing this technology at Yahoo will not be given to anyone from Summly, including its young CEO.
Notice that the version from TLDR Stuff actually has the Answer in it. "Incidental side effect of a deal Yahoo make with SRI"
It also tells you that they aren't interested in the young CEO.
This is possible because it is not a keyword density Algo, the core technology called Liquid Helium is a Language Heuristics Engine and it can put weight on which sentences are Causality, and which are Subject matter. This creates a version of the text that tells you Who What Why and if there is still space, How. You can't do that with just a KeyWord density or where in the article is this system.
Summly claimed to have that tech, SRI has some of it, but what they really have is a nice Concept Tree, and sentence parser.
A far cry from a system that knows which points are important, not just which points are most talked about, because as you see in this Business insider article the important part isn't "what is summly" or "who is nick" or "Who is mayer" it is "Why did Yahoo do this" and that is captured in the TLDRSTuff/Stremor version, and not in TLDRizer.
Here's an example of one of PG's essays run through the algorithm: [http://paulgraham.com/startupideas.html]
The most important thing to understand about paths out of the initial idea is the meta-fact that these are hard to see. Empirically, the way to have good startup ideas is to become the sort of person who has them. If you know a lot about programming and you start learning about some other field, you'll probably see problems that software could solve. Some of the most valuable new ideas take root first among people in their teens and early twenties. So if you're a young founder (under 23 say), are there things you and your friends would like to do that current technology won't let you. But there may still be money to be made from something like journalism. Similarly, since the most successful startups generally ride some wave bigger than themselves, it could be a good trick to look for waves and ask how one could benefit from them.
If you ran examples of PG's essays through this, people would see the immediate benefit.
Are you planning on open sourcing this?
Will try to add url content type detection in the next cut and summarizing non-feed url's next up
And The TLDR Plugin works with HTML and on all western languages.
The problem TODAY is not whether or not a computer can summarize, but rather to what extent we as humans are satisfied with the computer's summary.
In some cases a dumb summary is good enough (first 200 characters for example). Given this baseline, and a target (human summary), you have to admit it's really an incremental process.
Also, do keep in mind... this is 2 hrs worth of coding time late on a sunday night. I don't have a CS degree, just a utilitarian/curious programmer who sometimes is stupid enough not to realize how hard a problem I'm tackling. :) Someone better qualified can do a much better job. Sometimes "just good enough" is good enough! :)
Right now, abstraction or paraphrasing is hard to do by a computer. But I think and hopefully it will be possible in few years time. There are various open source and academic tools that can do some pretty good NLP. I'm looking into Apache OpenNLP, and WordNet. I'm hoping for 2 or 3 years time.
BTW, I have an app similar to your tldr.io. Check my HN comment (https://news.ycombinator.com/item?id=5523770) for more info about it. ;)
Generating news highlights from lots of sources might be cool as computer generated content. But rewriting an author's story in new words is not adding value it is just ripping them off.
Does it expect a specific format?
So what does it do exactly? It seems to extract some sentences more or less at random from the text...?
Awesome work!