there are couple things here:
1. scrapping web page to get text content
2. use NLTK to proccess text and get summary and keywords
3. wrap it into REST API and serve as web service
You could probably get cleaner input for step 1 via the Mercury API [https://mercury.postlight.com/web-parser/] — it has a lot of affordances for different kinds of HTML formatting.
Thanks will try at some point, my biggest concern was that those kind of API's are almost all paid and rarely open sourced.