Case in point: OpenStreetMap (run by volunteer sysadmins) provides completely open, machine-readable data dumps. You can use these to set up your own geocoding, rendering or similar service. There is copious documentation, several post-processed dumps for particular purposes, etc. etc.
OSM also provides a friendly, human-scale, user-facing interface for map browsing and search. There are clearly documented limitations/TOUs for automated use of this interface.
Does that prevent people from setting up massive scrapers to scrape the human-facing interface, rather than using the machine-facing data dumps? No, it does not; and the volunteer sysadmins have to spend an inordinate amount of time protecting the service against these people.
DataHen's proudly admitted practices ("No need to worry about IP bans, we auto rotate IPs on any requests that are made."; "our massive pool of auto-rotating proxies, user agents, and 'secret-sauce' helps get around them") is directly antithetical to this sort of scenario. I find this incredibly irresponsible and unethical.
Almost all web sites that authenticate use "submit a form with username and password and respect cookies"; often sites that don't authenticate to use the web site require authentication for the API. Every API uses a different authentication scheme and requires custom programming: for web sites you have the URL of the form and the name of the username and password field and you are done.
If you feed most HTML pages through a DOM parser like BeautifulSoup you can extract the links and interpret them through regexes. You might be done right then and there. If you need more usually you can use CSS classes and id(s) and... done!
I wrote a metadata extractor for Wikimedia HTML that I had working provisionally on Flickr HTML immediately and had working at 100% in 20 minutes. No way I could have gotten the authentication working for an API in 20 minutes.
What this conversation is about (since we're talking about web scrapers here, not "data downloaders" or whatever you would call it) is when there is no other avenue to get the data, you should be able to access the same data via a machine as you could when you're a person.
Similarly to HTTP, the tool does not decide if usage is "nice" or "evil", only the user with their use case decides that.
I generally agree with this, but I do see the problem for certain spaces.
Concert tickets, for example. People write scrapers to get the best seats and sit on them so they can scalp them later at inflated prices. Or other first-come/first-serve situations, online auction "sniping", etc.
Because that's the only way to be sure.
It's unfortunate that it also allows havoc and burdens good services. But device-over-user is not a future I want to live in.
Determine a reasonable rate limit and apply it to the human-facing version, with maybe a link to the machine-readable version in the error message?