Now try scraping Google and see what they do to you.
Now try scraping Google and see what they do to you.
If you provide value to Google they will make an API to allow accessing that data easier.
By scraping do you mean scraping their search results? They offer this, which is nice: https://developers.google.com/custom-search/
Many large sites don't allow scraping because of unnecessary server load (denial of service sometimes) so they'll offer an API where you can download content in a controlled (and monitorable) manner.
If one sends an email to an GMail/Outlook.com/Yahoo email address, one should be able to opt-out of their email crawler, advertisement analysis, artificial intelligence analysis, etc.
I don't think they do too much storing of email details, they know that it's a sensitive area and that an employee will eventually blow the whistle and it will hurt user trust, which is a big part of their business.
1) Alice sends an email to Bob (GMail user).
2) Bob receives the email, meanwhile Google scrapes the content and extracts its meaning to show Bob some ads and use the facts to improve Google's A.I.
Alice wants a way to mark her email content as "no index".
So that email service provider don't crawl through the content. Exactly like the robots.txt for domains or the "noindex" metatag in HTML head element!
It's about the email service provider, that should stop analyze the email text to extract its meaning. Gmail uses it to display ads to Bob, builds a shadow profile for Alice (like Facebook) and trains an artificial intelligence (see headline link).
If an unknown person tries to scrape, he/she will promptly get banned by those very same people (Google wouldn't like someone scraping their stuff either).
Different players different rules, I guess.
For example, if you went one by one through Stack Overflow and sucked out every question and answer, your scraper bot would get banned (unless you're doing one request per minute, in which case you'll never finish).
Or if you tried to scrape Twitter.