Creating a serverless function to scrape web pages metadata
mmazzarolo.com
mmazzarolo.com
It almost makes me feel that I am breaking the law when scraping a site, yet web scraping is on of the most basic programming things.
Just imagine where Google would be if it was a new startup and an existing giant like Cloudflare or Cisco blocked all attempts of access.
Yeah, same for me.
Regarding the denylisting, I guess it depends on what is being scraped and how often the scraping happens? I'm maintaining a remote jobs aggregator website and I've never been blocked before (but I'm not scraping more than ~5 times per day the same web page). And with a caching strategy, I think that even a scrape-as-a-service API like the one I'm building in the article should be "kinda" safe (besides edge cases that brute force the cache constantly, like by adding random query-params)?
“Most” sounds like an exaggeration. Wouldn’t this also create problems for virtual desktop services like Amazon Workspaces?
> It almost makes me feel that I am breaking the law when scraping a site
You might be violating their copyright, it depends what you do with it. If you overdo it, you could also degrade their service for actual users.
Spoofing your user agent is a must if you need to do anything nowadays.
To your second point that same would be applicable to Google and Bing and any other search engine. Even if you follow robots.txt and consume an equal or lesser bandwidth it does not matter much if you aren't an established player.
Google is not new to such complaints (news sites). Everything is very relative.
However, the term is generally accepted for cloud services where you run code on someone else's server. It is a product you can use.
I would be greatly disappointed if there are developers who think serverless code runs on air.