TextTeaser – An automatic summarization algorithm
github.com
github.com
Really surprised with both the quality and succinctness of the result: http://www.textteaser.com/s/t1bNud
Well done.
(Also to the project owner - copying the link is borked in Firefox, I had to type it out manually)
But I did not get what you mean by "copying the link is borked in Firefox". What link are you talking about? :)
Who is going to be first to couple this library with an RSS feed reader and mailer so that I can get auto-generated summaries of recently written articles sent to my blackberry?
Personally, I find the "email me news" scheme obnoxious. I get enough emails as is it. Would prefer to see a portal that shows summaries of all the news and lets the user explore.
Stremor's TLDR: Pointless reeling off the numbers. The challenge each of us now faces is a brand new one. Our filters were once the media, our friends, and our families. We were aware of, and understood how our filters operated. EdgeRank isn't something that Facebook users understand. Just 20 tweets out of thousands. We need better filters.
Text Teaser: How do we create a balanced diet of content with so much junk being thrown at us? Now, large media organisations create mountains of content, then track our reading habits and online behaviour in order to build a profile of us. A set of favourite news groups; a list of RSS feeds; a well-curated bookmarks folder; these are all filters we once built ourselves. The less we understand our filters, the more we will come to accept that the world they present us with is true. The more control we have over our filters, the more we can understand what we're not seeing.
Text Teaser goes over 350 Characters which is the Established People can't sue you for stealing it limit... So also weigh that when deciding which you like better.
If it is, I'm not using it. :)
I am also curious about how the NLP/ML parts are implemented, as it's claimed by the README on github. Briefly scanned the code but didn't really spot it.
If you leave the "Title" field empty and click "Summarize", nothing happens -- which I thought was very confusing. You have to fill in something for the Title.
EDIT it would be nice, from a UX POV, to request the title if it's missing, rather than silently deleting the story... also, you might emphasis the importance of it (because it doesn't seem important at all). Perhaps just labeling it as "headline" or "subject" instead of the generic "title" would help.
If so, I'll have to provide more literal subheadings!
At a cursory glance the algorithm seems like a variation Luhn's abstract algorithm.
Summarizing only blogs posts seems a bit limiting to me. (Btw, I'm not trying to be negative, congrats on your success! texteaser looks great!)
I implemented a custom version (mainly changed the scoring scheme to include TF/IDF of words for initialized scoring) of TextRank and loved it.
The main thing I liked about it was how general it was. Words are nodes and sentences are vertices. Then you basically use pagerank to rank the sentences according the graph representation.
I'm a little bit familiar with TextRank because I stumbled upon it when I'm doing my research. I also read several algorithms but forgot what they are called.
News is the most broadly applicable use for this so leveraging that isn't a bad thing. There's always a trade off of broad applicability vs overfitting for a particular case to get better results.
Thanks for the insight! Again great work.
What is the most important sentence in this:
Drakaal is a poopy head. He often posts to hackernews and calls people an idiot. When this happens I get mad.
The Premise is "Drakaal is a poopy head" so that is the most important, but it doesn't inform the user. What he does is the most telling of the sentences, but with the "He" as the first word you can't actually make sense of the content with out the prior sentence. The last sentence is the least important for the understanding of the content.
It is important to know that sentence number 2 is the most informative, but that to figure out what it means requires Sentence 1.
When computing the results of a summary you have to weigh sentence dependencies, density of information, amount of emotion expressed and number of characters available to you.
And Keywords aren't enough, you need noun entities and the ability to tell the relationships of words so that you know "Cars, Trucks, and Automobiles" are all the same concept in many contexts.
Check it out here: https://github.com/lukechampine/ADHN
Might I recommend taking CSS styles into account? Large text is usually headlines, <strong> text is usually important, and darker greys generally suggest a side comment. Would be much easier if everybody used <aside> and <h1> but even in 2013 that's too high an expectation.
results:
- Hacker Newsnew | threads | comments | ask | jobs | submit hnriot (1618) | logout upvote TextTeaser – An automatic - summarization algorithm (github.com) - If you leave the "Title" field empty and click "Summarize", nothing happens -- which I thought was very confusing. - reply upvote downvote MojoJolo 1 hour ago | link I require the title because I need it for the algorithm. - not a criticism of textteaser (which was behind this excellent project https://news.ycombinator.com/item?id=6498625), - reply upvote downvote wikiburner 3 minutes ago | link Is this a well known text summarization tool?
Not sure it captures the essence of the source article's argument. The fourth bullet makes no sense at all. I can't see it being useful at this stage.
Can you provide us with a list of articles that it manages to summarize properly?
It's great to see the project on github though. I look forward to seeing it improved over time. Thanks for sharing.
This isn't even on Par with Summly which was pretty hacked together.
https://www.mashape.com/stremor
Creates MUCH better summaries ans comes with all the stuff to separate Content from the web template.
If you contact Stremor there is also a version that scores every sentence for importance on a scale of 0-100 and maintains HTML so that you can return summaries of any length and still have images and other styling maintained.
( http://www.tldrstuff.com has several ways you can play with the tech )
you post non-idiomatic(?) scala in a comment to explain what you are doing, i think? not a criticism of textteaser (which was behind this excellent project https://news.ycombinator.com/item?id=6498625), but seems to raise questions about the language...
One minor nitpick that can be of help when dealing with tuples: A partial function ( {case xxx => yyy} ) is a Function1 so you can use it with map and filter. This way you can deconstruct tuples into names and avoid using _1, _2, etc. { case (name, value) => blah }
https://github.com/MojoJolo/textteaser/blob/master/src/main/... could be made more readable by giving names to the tuple elements.
Thanks for publishing this code. It yields impressive results.
1. perhaps `.reduceLeft(_ + _)` could be replaced with the use of a `sum` function or method (assuming one exists in scala?)
2. if the `topKeywords` collection returned a default value with a `.score` of 0 when queried with a key it doesnt contain, the headOption getOrElse null match null would not be necessary.
e.g. in python it might look something like this:
Keyword = namedtuple('Keyword ', ['score', ...etc...])
top_keywords = defaultdict(lambda : Keyword(score=0, ...etc...))
def sbs(words):
if words:
return (1.0/len(words)) * sum(top_keywords[w].score for w in words)
else:
return 0.0
(apologies for making superficial comments about the code. the algorithm itself certainly seems interesting)https://news.ycombinator.com/item?id=6498625
https://news.ycombinator.com/item?id=6049873
In TC:
http://techcrunch.com/2013/10/06/textteaser-lets-developers-...
Anyway, thanks for open sourcing - really cool.
Think MongoHQ for MongoDB.
Edit: the main TextTeaser web site is down right now, which is why I went straight to the API to test.
He can set up a limited free private API for you to test. Let me know if you have questions about this process - chris@mashape.com
What is the structure of the sent.model file inside the corpusEN.bin zip archive?
It's a strikingly small file for something called corpus. Say I have a larger corpus, or a corpus in a different language, how would I go about building one of these sent.model files with more data?
Also I remember it being a bit pricier to use the API. What made you to go down on price? I'm tempted to hook this up to my app right now.
Right now, the TextTeaser website is coded using Python and Flask.
If you have a single NLP model that already works, you wouldn't gain anything from rewriting it using NLTK. It would probably just get slower, because you're adding abstractions that you've already shown you don't need.
I say this as a fan of and (once) contributor to NLTK.
I created a news reader for Philippine news (http://www.readborg.com/) using TextTeaser. The word Philippine appears most of the time and I decided to make it as a stop word. Forgot to remove it in the stop words.
-Marissa