A great case is the UK Post Code system (which is much more specific than the US zip code system here - no more than 15 letter-boxes share a unique post code in UK).
A number of people have created Post Code to lat/long conversion tables. They're all facts but because they have been worked out they are copyrightable. To stop copying the providers introduce slight variances in the data to track duplication and usually jump on top of it.
I don't believe that would fly here in the USA
Afaik there is no law that mandates you must read/agree to the ToS before loading a website. In fact, what if the ToS states 'If you visit this site, you must pay us 1000 dollars', are you now liable? Despite the fact that just to read the ToS you had to visit said site.
My point is that only a court will decide what happens, and just because something is in the ToS doesn't mean it gives you carte blanch to do whatever you please.
There's no physical analogue to scraping a database-driven site. The question of exactly how legal it is probably won't be settled until that fact penetrates the legal system; until then it depends on what metaphor you sell to a given judge.
I'm really playing devil's advocate here, I do very large amounts of scraping online for various projects and do not bat an eye lash at what I'm doing. If it's online, it's there for the taking. If you don't want me to scrape it, hide it.
In fact I agree with you that it is broadly speaking incumbent upon someone who does not want to be scraped to have at least some protections technically enforced, however feebly, and you should not go out of your way to violate such protections however feeble they may be. But I come to that conclusion thinking about the monetary issues and bandwidth issues and ethical issues directly, not by making a bad analogy to people leaving doors open or locks on gates in the middle of the field or houses constructed out of glass.
My point is that someone works to create something that is freely available to peruse (website content = book at library), and anyone who comes in and copies and sells that content will make said author upset.
"Ethics" in cases like this are so subjective that it's almost useless debating them.
It's funny how public perception of this can be driven by folks who have never met you or done business with you. You would be hard pressed to find anyone in business, an investor, a partner, etc. that would ever say anything to the contrary. Thus the reason I'm able to raise money, partner, sell companies, etc. over and over again despite the fact that i get negative comments/blog posts from haters.
There is reality, and there are comment threads. :-)
a.) Google follows robot.txt. You can disallow Google to index your website. Most of the websites, OTOH, want Google to index websites.
b.) Google does not republish the content. All the traffic is directed to the content owner, ie the other websites.
Of course they are not, they are simply re-published snippets of the websites along and Google surrounds the results with ads.
While it's a symbiotic relationship that most websites want - sharing their content for placement in Google's webpages - it's not necessarily universal and ROBOTS.TXT is hardly a "contract" covering your data's usage.