Show HN: Extract Markdown, HTML or text from content-heavy websites
content-parser.com
content-parser.com
It also uses turndown and readability.
It's pretty finicky (readability doesn't always identify the correct content or misses pieces of the content). If you want to charge for it, you'd have to fix some of those edge cases.
Also, I don't think the value is this product is turning web pages into markdown, there are many free web clippers and archive sites that do this already. I see this as more of an "extra" in a product, like how Evernote has a web clipper built in to their note taking product.
Also, it's cool to see other people care about a stripped down web reading experience too!
- Readability (https://github.com/mozilla/readability) to strip down the page's HTML to a bare minimum.
- Turndown.js (https://github.com/mixmark-io/turndown) to convert the plain HTML to a markdown format with the GFM plugins enabled.
- Puppeteer (https://github.com/puppeteer/puppeteer) to download the page.
It costs me only several cents to parse an entire page, and I think OP can make some money out of this if they get the pricing right.
Also, some unsolicited feedbacks on the API:
- An option to enable/disable javascript would be great, since not all pages actually need to have it enabled to be parsable.
- You can probably tweak the header of the headless browser to bypass the paywalls of some sites. Some are as simple as setting the useragent to a crawler bot (like `googlebot`).
- Maybe an option to fill in the front matter (https://jekyllrb.com/docs/front-matter/) with a metadata given in the payload?
I should also find a better html-to-markdown parser, thanks for the recommendation there! From the "example", yes you guessed "readability" perfectly. And for downloading the page just fetch() + jsdom.
Suggestions:
- [JS]: I use fetch+jsdom, so no JS parsed at all! I've found most content-heavy websites (a.k.a. articles, blog posts, etc) are server-side-rendered, haven't searched too many but so far no issue without JS. Might move to puppeteer at some point for either failed parses with jsdom or for a domain whitelist if I keep one at some point.
- [header]: Already mentioned
- [Front matter]: Right now I'm actually returning two custom headers, `title` and `url`, might add more in the future. I did consider front-matter, but I want to keep the body as "raw" as possible.
- Edit: what I'm considering next is an endpoint to download articles with basic HTML style, or as pdf/epub.
Are you dividing the monthly hosting costs for a server by total seconds spent actually running this tool? I'm thinking if you did this with an AWS lambda it'd be free (maybe bandwidth cost, but again, trivial) unless you had way, _way_ more use than a single person could reasonably generate. Also, free if you used any of the free hosting services and were just doing it for a small number of users.