What you put online is public. Treat the web the same as a public bulletin board. Everything you post will be readable by everyone. Trying to prevent this is pointless.
What's the difference between a machine scraper and a human scraper? Humans use machines to access websites. All your users are machines. And behind every machine is a human.
If you don't want someone to access all your data in a short time-frame, then don't give it all away in a short time-frame.
If I just download a site – for local use, or similar cases – it’s not an issue.
And if I provide a scraper where the user customly has to enter the URL, it’s also technically just a very advanced webproxy, and also okay.
Think about it from the point of view of the site being scraped. Suddenly your webserver is being slammed with traffic that's way beyond what any human would cause. It's costing you bandwidth and compute resources. You have every right to say no, scraping is not allowed (in your Terms of Service). It's then up to a court to decide if those Terms of Service are legally binding. Some courts have decided that yes, they are.
Since you're clearly not a content creator, and have decided since you didn't work on it it has no value, but putting something online does not make content public domain.
Google scraps webpages as a core competency.
I might just want to train a neural network on extracting data out of websites, and, for that usecase, wget -m every website I can find.
How I use the data you present to me, as long as I don’t republish it, is my decision. If you give me a license to read it, you also give me the license to copy it onto up to 7 different media at the same time, and to show it to up to 7 friends at the same time.
My site, my robots.txt - if the scrapers do not obey, then I'll block them.
There are ways to force authorization, if you don’t use them, your fault.
If I just use TOR or VPNs or similar to get around your block, also, again, your issue. Unless you require authorization, your service is, by law, public.
And if I want to store the site offline for later reading, or read it now in a browser, is nothing that you have to decide.
No, it's not "by law" public, or could you cite a few laws for that, both US and EU. Would appreciate that.
HTTP directly says you should use authorization to make stuff non-public.
And third-party content scraping tools are NOT violating copyright law, as they do not redistribute it freely.
Otherwise services like Opera Turbo – which scrapes a website, removes cruft, compresses data, and sends it to your phone – would also break copyright law, and there is a nice exception for this case, under US law, it’s covered under fair use.
(Unless the scraping service makes your content accessible publicly – and even then, fair use applies usually, as in the case of pure archival like archive.org or archive.is)
Are you a lawyer? I am not - anyway, I can't really agree with that falls under fair use.
If I charge money to scrape someones site and then hand over the material in whatever form my customer uses - I am bound run into problems sooner or later.