How did we start with consent using robots.txt and end up here?
A good neighbor doesn't use circumvention.
How did we start with consent using robots.txt and end up here?
A good neighbor doesn't use circumvention.
Google sends some traffic so I can afford to let them scrape. Bing crawls almost as much as Google but sends 5% as much traffic. Baidu crawls more than Google and never sends a single visitor.
I hate reinforcing the Google monopoly, but a crawler that doesn't send any traffic is expensive to serve.
The interfaces for many sites actively and with brutal effectiveness deny ready access to any content not currently featured on the homepage or stream feed. Search features are frequently nonexistent, crippled, or dysfunctional.
Last week I found myself stumbling across a now-archived radio programme on a website which afforded exceedingly poor access to the content. The show ran weekly for over a decade, with 531 episodes. Those are visible ... ten at a time ... through either the Archive or Search features.
Scraping the site gives me the full archive listing, all 11 years, in a single webpage, that loads in under a second. I can search that by date, title, or guest to find episodes of interest.
The utility of this (a few hours work on my part) is much higher than that of the site itself.
Often current web sites / apps are little more than wrappers around a JSON delivery-and-parsing engine. Dumping the raw JSON can be much more useful for a technical user. (Reddit, Diaspora, Mastodon, and Ello are several sites in which this is at least somewhat possible.)
Much of the suck is imposed by monetisation schemes. One project of mine, decrufting the Washington Post's website, resulted in article pages with two percent the weight of the originally-delivered payload. The de-cruftified version strips not only extraneous JS and CSS, but nags and teasers which are really nothing but distraction to me. Again, that's typical.
I'm aware that many scrapers are not benign. More than you might think are, and the fact that casual scraping is a problem for your delivery system reflects more poorly on it than them.
What you are doing stripping out the junk threatens those organizations at the core.
If mobile phone companies kicked back a fraction of the revenue they get to content creators they'd be better paid than they are now and Verizon would get the love that it has sought in vain. (e.g. who would say a bad word about the phone company?)
Gobal ad spend, which mostly accrues to the wealthiest 1 billion or so, is about $600 billion. Some complex maths tells us that's $600 per person in the industrialised countries (G-20 / OECD, close enough). Global content spend is somewhere around $100 -- 200/year per capita. That's roughly the annual online ad spend.
Bundled into network provisioning, that's about $30--40 per household per month, all-you-can-eat. Information as a public good.
(My preference is for higher rates in more affluent areas, ideally by income.)
Trying to figure out WCPGW.
Think of the old phone company slogan "reach out and touch someone." If I can accomplish that and spend less than I do on food or clothes or my car then I win.
The challenge, as I see it is that information is a public good (in the economic sense: nonrivalrous, nonexcludable, zero marginal cost, high fixed costs), and provision at scale requires either a complementary rents service (advertising, patronage, propaganda, fancy professional-services "shingle") or a tax. Busking or its public-broadcasting is another option, though that's highly lossy.
Any truthful publishing also requires a strong self-defence mechanism (protection against lawsuits, coercion, intimidation, protection rackets, etc.), a frequently underappreciated role played by publishers.
Charles Perrow's descriptions of the music industry (recorded and broadcast) circa 1945 -- 1985 is informative here (see his Complex Organizations https://www.worldcat.org/title/complex-organizations-a-criti...), notably the roles of publishers vs. front-line and studio musicians.
The web needs a more efficient system for distributing an index of its content. Having web developers design websites in a gazillion different permutations, when they are all bascially doing the same thing, and then having a handful of companies "scrape" them is neither an efficient nor sensibly-designed system.
The web (more generally, the internet) is a giant copy machine. Google, NSA and others have copies. Yet if we were to allow everyone to have copies by faciltating this through technical means (e.g., Wikipedia-style data dumps), many folks would panic. When Google first started indexing, it was not a business, and many folks did panic and there were many complaints. It's 2021; folks are still spooked by others being able to easily copy what they publish to the web. However it's OK to them if it's Google doing the copying. If there were healthy competition and many search engines to choose from, if one search engine did not have the lion's share of web traffic, it's doubtful Google would be given "exclusive" permission in robots.txt.
Case in point: OpenStreetMap (run by volunteer sysadmins) provides completely open, machine-readable data dumps. You can use these to set up your own geocoding, rendering or similar service. There is copious documentation, several post-processed dumps for particular purposes, etc. etc.
OSM also provides a friendly, human-scale, user-facing interface for map browsing and search. There are clearly documented limitations/TOUs for automated use of this interface.
Does that prevent people from setting up massive scrapers to scrape the human-facing interface, rather than using the machine-facing data dumps? No, it does not; and the volunteer sysadmins have to spend an inordinate amount of time protecting the service against these people.
DataHen's proudly admitted practices ("No need to worry about IP bans, we auto rotate IPs on any requests that are made."; "our massive pool of auto-rotating proxies, user agents, and 'secret-sauce' helps get around them") is directly antithetical to this sort of scenario. I find this incredibly irresponsible and unethical.
Because that's the only way to be sure.
It's unfortunate that it also allows havoc and burdens good services. But device-over-user is not a future I want to live in.
What this conversation is about (since we're talking about web scrapers here, not "data downloaders" or whatever you would call it) is when there is no other avenue to get the data, you should be able to access the same data via a machine as you could when you're a person.
Similarly to HTTP, the tool does not decide if usage is "nice" or "evil", only the user with their use case decides that.
I generally agree with this, but I do see the problem for certain spaces.
Concert tickets, for example. People write scrapers to get the best seats and sit on them so they can scalp them later at inflated prices. Or other first-come/first-serve situations, online auction "sniping", etc.
Determine a reasonable rate limit and apply it to the human-facing version, with maybe a link to the machine-readable version in the error message?
Almost all web sites that authenticate use "submit a form with username and password and respect cookies"; often sites that don't authenticate to use the web site require authentication for the API. Every API uses a different authentication scheme and requires custom programming: for web sites you have the URL of the form and the name of the username and password field and you are done.
If you feed most HTML pages through a DOM parser like BeautifulSoup you can extract the links and interpret them through regexes. You might be done right then and there. If you need more usually you can use CSS classes and id(s) and... done!
I wrote a metadata extractor for Wikimedia HTML that I had working provisionally on Flickr HTML immediately and had working at 100% in 20 minutes. No way I could have gotten the authentication working for an API in 20 minutes.
> Till was architected to follow best practices that DataHen has accumulated
I know there are plenty of sites out there out to prevent any scraping whatsoever, or to improperly prevent some situations most of us would agree is reasonable behavior, but this appears to be blatantly hostile to web admins out there.
My concern is that tools like these that use circumvention by default become the go-to when someone needs scraping and makes life hell for us running sites on hardware that's enough to service our small customer base but not an army of bots.
Since Feb 2021 we've seen a substantial increase in scraping of our customers' inventory to the point we now have more bot traffic than human traffic. Annoyingly, only 5-10 requests come from a single IP so we've had to resort to always challenging requests that come from certain ASNs (typically those owned by datacenters).
This type of project frustrates me because it knowlingly goes against a site's desired bot action (via robots.txt).
None of this gets around over-eager Cloudflare or Akamai rules set up years ago by some contractor that the businesses have no real ability to change.
If the business has no ability to even change the CloudFlare settings I don't expect them to be able to provide an API.