Ask HN: How does outline.com work?
I'm pretty amazed that I can get around a paywall just by using https://outline.com/www.[my url] . I'm sure there's nothing too crazy going on under the hood, but does anyone exactly how it works?
- Pretend they're a crawler such as Google and pull down the HTML, potentially executing javascript
- Once it's pulled down, clean it up using open source code such as readability https://github.com/mozilla/readability
- Store that result as a document in a nosql database
Once they have pulled the article down once they don't need to get it again.