- Readability (https://github.com/mozilla/readability) to strip down the page's HTML to a bare minimum.
- Turndown.js (https://github.com/mixmark-io/turndown) to convert the plain HTML to a markdown format with the GFM plugins enabled.
- Puppeteer (https://github.com/puppeteer/puppeteer) to download the page.
It costs me only several cents to parse an entire page, and I think OP can make some money out of this if they get the pricing right.
Also, some unsolicited feedbacks on the API:
- An option to enable/disable javascript would be great, since not all pages actually need to have it enabled to be parsable.
- You can probably tweak the header of the headless browser to bypass the paywalls of some sites. Some are as simple as setting the useragent to a crawler bot (like `googlebot`).
- Maybe an option to fill in the front matter (https://jekyllrb.com/docs/front-matter/) with a metadata given in the payload?