try now. the blog.medusis.com/rss link works now. Thanks for the feedback. since this grabs the page text (when no rss text is found) a lot of junk like copyright notices etc. shows up in summary. Will have to add some logic to scrub those. It also behaves horribly with code snippets.
Excellent, thanks, it does work now.
So what does it do exactly? It seems to extract some sentences more or less at random from the text...?
the algo is listed at the bottom of the home page. Will be opensourcing this code soon.
no, i use ROME to parse RSS feeds. So it should be able to handle whatever that can handle. Let me check
ah, got it. if no entry text is present, I was assuming each entry to have a description field. Fixed it. if both entry text and description are missing, it fetches the url text of the link and summarizes it. pushing to heroku