Hello. I recently finished my thesis for my MS CS degree. My thesis is about automatic summarization. It undergoes research, defense, and I think its result is good enough for me. It uses statistical approach and machine learning. My main issue about it is not the summarization part, but the text extraction part. I can't seem to extract article in a web page well enough. I'm using boilerpipe (
https://code.google.com/p/boilerpipe/) for it. It can do most tricks, but it's not that good for me. May I ask how you extract the main article in the page?
Here's a preview of mine (http://www.textteaser.com/ui/article?link=http%3A%2F%2Fwww.p...). Go to its home page to read more news. It caters Philippine news and will soon enters alpha stage. I'm planning to open up the API or open source it. HN, which is better? The API is ready, registration is the only thing that it lacks.
You can try the API here:
http://api.textteaser.com/api/?url=http://www.theverge.com/2...
Just replace the url parameter with the URL of what you want to summarize. Some URLs are not tested yet, and may produce errors. :)