Web Scraping Is Vital to Democracy
themarkup.org
themarkup.org
This applies not only to personal life, but businesses as well. Instead of hiring 100 interns to collect some data manually, you hire one programmer to automate that data collection process.
I think that the ability to automate tasks on the internet is absolutely crucial to further development of our society and limiting it in any way will be detrimental to the world as a whole. The amount of information these days is so vast, that no human labor force could possibly analyze that information and use it to drive our progress.
Disclaimer: We run a web scraping platform (https://apify.com)
I say this as someone who runs an adblocker, installs them for family, and doesn’t derive income from hosting ads, so I’m not pro-ad; I just realize that while it’s not a crime for me to dump the whole bowl of waiting room candies into my backpack, that’s going to be frowned upon.
If I run a search engine which potentially links real human eyes back to you, then should I pay the "ad" toll as well?
I don't believe so.
I do believe there is an agreeable middle-ground, but Google walked away from that conversation years ago.
That being the case, any argument founded on "but advertisements" does not hold water.
While you may have a point, that point just doesn't matter within the context that we're talking about. Making companies money is not my responsibility.
Consuming the content means agreeing to the premise that the content is paid for using ads.
If you disagree with said premise you may happily browse another website.
Malware is already illegal to install and you are free to sue the website for damages.
I don't like ads either but trying to justify that their should be a choice of browsing a website without the ads it hosts is ludicrous.
It is in fact not this way, because the content arrives with or without the ads. In the EU EULAs (the dystopian construct that you'd expect to enforce that bit of lunacy) that purport to apply to content you have already accessed are not legally valid. Leaving the legal interpretation aside, me doing one thing doesn't mean consent for something else. Believing otherwise is both unethical and amoral, stances I don't hugely feel like interacting with.
But that's simply a technical implementation detail.
If the articles you read had first party ads or videos with embedded ads in them, that choice wouldn't exist like it already doesn't when you watch TV.
> Leaving the legal interpretation aside, me doing one thing doesn't mean consent for something else. Believing otherwise is both unethical and amoral, stances I don't hugely feel like interacting with.
Yes and you stealing content (consuming without paying for it) is somehow moral?
It's easy to find all kinds of free content, should it be music, movies or series but lets not kid ourselves in thinking that it is some great human right to have access to things we didn't pay for.
I could not do this for any of other several large sites advertising used trucks, like commercialtrucker.com, truckpaper.com, and machinio.com. I kept bumping into artificial javascript and captcha limitations that were not worth my time to try to work around.
The thing is that I will probably find what I want on a craigslist site that I can scrape, I'm getting a lot of great info from them, and I'm not going to bother with any of the ad-based sites, it takes to olong to run all the manual searches. They have effectively done a disservice to their (presumably paying) customers who want to sell their vehicles.
But you don't mention what you do with this data after. 1. You store it somewhere (not in your brain) 2. You extract value out of it (directly or indirectly)
That's why I understand why it is a problem for those who publish this data.
One thing I always try to say to everyone who argues about web scraping: web scraping is not a problem, problem is what you do with the information you scraped.
Disclaimer: We crawl the web for news (https://newscatcherapi.com/)
I still don't understand why it's a problem for those who publish the data.
If the "data" is facts (e.g. lists of values associated with objects, like the colors available on automobiles), this data is not protected by copyright, there is no remedy if it's republished, and if a business relies on limiting access to this data, that business needs a new business plan. Charge for access to cover the costs of obtaining and organizing the data, but understand that clients are legally allowed to make copies and use them however they see fit, even if that impacts the supplier's ability to charge for access.
If the "data" is prose, it's under copyright and republishing without a license has remedies under law. Maintaining copies of articles for the purpose of processing them to obtain other data (e.g. how many nouns? what adjectives are near those nouns? etc...) isn't protected.
If it's not prohibited for some reason, e.g. because the term of the prohibition, such as copyright, has expired, why not.
Sure, you could argue that news are facts, and that they should be free from copyright. But that would only work if your news is just bullet points of the factual stuff. But if any element of bias is present, it become a distinct property of the author, hence it should be copyrightable.
Of course, with today's tech, it's easier to remove the bias element from the factual data.
Totally different? Yes, but it shows that scale can affect whether something is OK to do or not, at least for some (I accept that people watching the streets can be useful, but also have my doubts about doing that at scale, whether by cameras or by hiring a million agents)
With both cameras and copy-pasting web content, there’s the issue what you do with the data. If, for example, I start scraping all the articles on a newspaper’s web site, publish them on a web site, adding my own ads, most people would think that shouldn’t be legal.
If you agree, we’re now haggling over the price (https://quoteinvestigator.com/2012/03/07/haggling/). That’s where things get difficult, but I think the entire spectrum from white (scraping for this goal is fine) via grey to black exists.
Publishing? Ads?
Another similar trick: "If someone will block ads and then murder the publicist and burn his house, most people would find that it should be illegal".
> If, for example, I start scraping all the articles on a newspaper’s web site, publish them on a web site, adding my own ads
Using a Q-Tip to murder someone is illegal. We don't have special laws against it—it's just that murder is illegal, we have criminal punishment for murder, and the ordinary machinery of the courts and the law is sufficient to handle any instances of murder by Q-Tip. Because it being done with a Q-Tip is the least important part.
If you're scraping and republishing someone else's content, it's not the scraping part that's the problem.
I agree that scraping should be allowed. We wouldn't have Google otherwise.
However, there's something to be said for the fact that some activities are okay at a small scale but become problematic at a large scale—particularly the scale which becomes possible when a task is automated.
For instance, I recently bought an expensive camera and have been having fun walking around the city and taking interesting photos. Many of the photos have people in them. I don't think there's any harm in this.
I store my photos in Aperture, which automatically performs facial recognition on everything in my library. Perhaps some day, if I take enough pictures, Aperture will notice that the same stranger is present in two completely different images. That might be kind of cool—I can see myself fancifully trying to imagine this person's life story. I don't think there would be any harm in that, either.
However, if I aggregated millions of photos from different sources, and used them to track people's movements across the city, that would clearly be a huge problem! Sure, it would be merely automating the work a human could do, but the scale just changes everything!
For reference this is what one site (findlaw) has given as what constitutes the crime of stalking:
The crime of stalking can be simply described as the unwanted pursuit of another person. Examples of this type of behavior includes following a person, appearing at a person's home or place of business, making harassing phone calls, leaving written messages or objects, or vandalizing a person's property.
Sure the information is free. But to find my property records, phone number, and "aggregate them" like Spokeo, fastpeoplefinder and similar sites, is akin to digital stalking IMO
For example, if someone's ex (who was explicitly warned that he was unwelcome) continues showing up at the doorsteps because they remembered your address from a long time ago, then it is stalking. Them knowing the address isn't.
In light of this, I don't see how it matters whether they remembered the address from past experiences or just found it through a website that aggregates publicly available info. As long as the data was obtained legally and without breaking any other harassment clauses, why would just the knowledge of something be a crime?
I see it just like firearms. Having a firearm (in a lot of US jurisdictions) is not a crime, as long as it was obtained legally. Doing harmful actions with it (such as threatening people or shooting someone who didn't pose a threat to your life) is a crime. Having a firearm feels like just having data, in this scenario. As long as you don't use that data for criminal actions, why would just the potential of you being able to do something criminal with that data is a crime?
Except for the last point, that sounds like ads :)
Or are you referencing computers with ads in them (websites, etc.)? By that definition, anything disproving a conspiracy theorist is vandalism to them. That’s not how that works.
I did this with Python regarding market prices for ETFs, Stocks, and Mutual Funds.
I wrote Python scripts, one of which was to web-scrape current market values, investment distribution, &c for ETFs, stocks, and Mutual Funds for whic I was invested in. I then would have to manually port that into a spreadsheet for "my own special graphs" and such.
I'm sure there are places online which would do this for me if I logged in, entered in all of my information and more, and provide that all for me ... but this story should sound familiar. And due to personal reasons, I've been lagging behind for far too long.
My Point: This information is publicly available and and it serves my purposes. There's no reason why I should not be able to do this. Yes, it's my fault for not using the information [and yes, the information dies with me], but the point is that I automated the gathering of that information for my own self -- and "everyone" else on the planet has that same information.
I never felt like I was illegal doing this. I'm happy to know if there is a "save" way to do this type of aggregate situation otherwise.
I'm not sure about that. If I had a phone book full of phone numbers (those heavy ones from 80s), would calling every number in that book to find the one I'm looking for be legal/ethical?
P.S: I agree that "Web Scraping Is Vital to Democracy".
Sure, why not?
Fight against web-scrapers just seems like a complete logical oxymoron: they want data to be public but also select who gets to see it. Our whole web infrastructure are based around clearly distinct public/private exchanges - there's no middleground this and yet people create these absurd hacks like captchas and fingerprints to fight the nature of the internet.
Finally everyone wants benefits of public data (search engine indexing etc) but don't really want to give anything back to the ecosystem. It's just pure greed and law, our society and government shouldn't aid it in any way, shape or form.
There's a POV that, on the internet, "anything goes", i. e. whatever you can do, you're allowed to do.
Then, there's a perspective that works a lot like the offline world, where any clear communication that a reasonable person would understand as denying them access becomes legally binding.
In the offline world, we derive great benefit from following the second model. Indeed, if only measures that successfully prevent people from entering your house without permissions were to count, you wouldn't need laws in the first place! You would, however, need a bunker. Which is quite a bit more expansive than a functioning legal system.
In that offline world, we have created all sorts of additional rules to balance rights for specific situations, and we rely on a canon of expectations that say, for example, that it's usually not ok to enter a private house, but you don't need explicit permission to enter a supermarket.
These are still developing for the online world, and your idea of the "ecosystem" hints at that. But you're just taking from those contradictory ideas above to arrive at the outcome you intuitively feel is "just": a bit of might-is-right when it comes to "whatever is online is fair game", followed by principled ideas of rights and obligations when websites try to defend themselves in that jungle of yours.
What's really needed is something that can, for example, distinguish between a journalist scraping Facebook to map out a terror network vs. some other entity scraping Facebook to sell your embarrassing photos to the highest bidder ten years down the road.
There is no middle ground _in the protocol_ and that's why I explicitly said hacks. Any system can be rehashed into anything else with unlimited extra layers on top of it - by your definition everything is everything.
> What's really needed is something that can, for example, distinguish between a journalist scraping Facebook to map out a terror network vs. some other entity scraping Facebook to sell your embarrassing photos to the highest bidder ten years down the road.
Sorry but that sounds uneforcable and rather absurd. We have the framework in place already - if you don't want something to be public don't put it out in the public _explicitly_. To add we already have legal framework in place for copyrighted and/or protected content like photos and against any sort of malicious attacks like ddos.
Our web is getting extremely centralized and most of these majors are natural monopolies: google search becomes stronger the more data it has - google can scrape the entire web freely yet it's competition can't; facebook becomes stronger the more data is has etc. etc. One way to restore balance is to ensure that public data remains public and the ecosystem can have healthy competition and growth otherwise we're moving to a very dystopian corporate owned world.
If you want to put restrictions on the usage of your data, make people sign a contract before accessing it.
Also a contract should not have "the public" as a party or a subset of, you should be able to identify the parties you have contracted with. Else you may end up with warrants targeting everyone or a subset of...
As long as people keep noticing how stupid Ayn Rand is before they come of voting age, we do have some protection against surprises: you can't just sign over your house or your first-born by clicking on a cookie banner. But I'm pretty sure Facebook could make you type "I won't scrape Facebook" into a box and it'd be (civil-law) binding.
I think you're proving too much here. Your argument applies to all published authors, and would strike a crippling blow to their copyright.
Secondly the Internet is best viewed as a public noticeboard purely because of the way the protocol works. There's just no getting around that. I think you'd agree that putting up a notice on a street corner and then getting offended when people read it would be viewed as rather odd, if not something else.
Wouldn't it be ironic if the non-selective people used this as leverage for discrimination.
They could. But if they filed a claim under the CFAA without ever sending a cease and desist letter to the alleged intruder, I think the claim would be dismissed.
Is simply changing TOS enough for a CFAA claim to have a reasonable chance of success. I could be wrong, but I believe for every CFAA claim we have seen so far based on "scraping", there was some notice to stop directed specifically at the respondent. If someone thinks just changing TOS (public notice) is enough for a CFAA claim to survive a Motion to Dismiss and there is no need to also contact the alleged offender asking them to stop, then let's hear about the precedent that supports that idea. I do not think there is any such precedent, but I could be wrong.
Other than that, you're mixing civil and criminal law rather liberally. I agree that it would be insane to create criminal liability for run-of-the-mill violations of ToS, and there is a decision from the MySpace era saying as much (and not even involving any changes to those ToS).
But once you have been specifically asked to not do something, by any means that would reasonably get that message across (so, not just C&D), it becomes... murky?
I could set "Terms of Use" for the data stored on protected computers I own. I could place limits on "acceptable use" of this data. But can I really argue that I gave adequate notice to all the tech companies and partners that try to access this data.
Key idea is this: access to internet and its services is essential (and at the web scale, so is automation), but bot abuse is real. It should be possible for any company to ban bots but then for a person with a account on the website to say: I am giving this tool my cloud power of attorney (best if signed through a digital govt id system, e.g. https://en.wikipedia.org/wiki/BankID) and I take responsibility for what it does; you may not block it or erect CAPTCHAs for it, it scrapes/takes actions on your website on my behalf for my personal needs. This would make running a bunch of scripts on your own NUC or Pi an inalienable right while still allowing companies to fight unfair competition and plain simple DoS attacks.
That's basically what API keys are, aren't they?
Sometimes to do some task at a 3rd party tool, it requires many clicks and page loads. My scraper wrapping the tool automates most of the task and only keeps the manual task, reducing the amount of required clicks (and time) required to get the task done.
I use a mix of headless browsers (for JS-heavy apps) and raw API calls. Sometimes I even use the browser to login, trigger a single sample APi call, extract all headers and content, close the browser, and re-run that api call directly using a HTTP lib. The request body obviously gets modified on the user’s requirements. We’re bypassing the slow login process as long as the session is valid. We’re also sharing sessions/logins/accounts this way without exposing credentials to users.
It can also be done to bypass the only-one-session-per-user systems. This is done with permission from 3rd parties. They’re fine with it, they just didn’t want to provide a proper API or let us bypass the only-one-session rule because it requires code changes they’re not willing to make only for us.
Sometimes the tool breaks when the html Or api changes but it usually only takes a few minutes to modify the code to fix it.
While coding it it was briefly on stock but I was too slow with my card input hahaha and so I neeed to now automate the buy part as soon as it's available, or make that run every 5min and mail/notify me if available...
Btw this was super easy, playwright as a puppeteer successor really rocks, another thing hard to hate from MS like TS or VS Code
PS: I brought down script execcution from 7 secocnds to 2 seocnds, by not loading any unneeded stuff css images external js fonts.
I felt like a good scraper citizen doing so, it was just two linnes to block by request type on the network level, and it worked like a charm to speed up the process.
The article then goes on to defend web scraping which is very different because (1) the person in question accessed data manually, (2) the person in question had access to confidential information, (3) the person in question used the data in a way he must have agreed not to as part of his employment. It's hard to see how someone could connect whatever precedent this case sets to web scraping.
The link is another case, Linkedin v hiQ, that is being held pending this case and presents the same question of what counts as "accessing without authorization or exceeding authorization" under the CFAA. The dispute there is whether hiQ could scrape public LinkedIn pages.
The crux of the issue is that if instructions on how to use data that someone has access to without "breaking and entering" don't count as revoking "authorization", this law doesn't cover this officer's actions (though other laws / job requirements may). If just breaking verbal or written terms of use is enough to criminalize it, that covers a whole bunch of things we'd think of as not federal crimes, like lying about your age to set up a facebook account.
Tons of our customers use us not just to archive a snapshot of the site but also to keep tabs on any visual and scope changes over time. All without any coding. 100% Automated.
and ERP means "enterprise resource planning"
(maybe everyone else knows these, but just in case)
https://en.wikipedia.org/wiki/Oracle_Applications#Oracle_E-B...
One spider may be equivalent of hundreds of real users. And it spoils caches too.
I'm working on simplifying web scraping for developers with https://webscraping.ai and see how important it is for almost any business or researcher.
https://news.ycombinator.com/item?id=25254499
Can't really fault The Markup for looking out for their own techniques/strategies
Now all we have is SPA 30MB main.min.js mess.
I really thought Obama Admin's data.gov initiative signaled a sea change.
Share your data. Show your work.
Science, journalism, others, face this same crisis. Legitimacy, credibility, authenticity, accountability, etc.
It's this obvious? Why are we still talking about it?
The few times I've written scrapers, one of the things I've done is put contact information in the user-agent so that if there is an issue, the site admins can reach out. So far that's worked out ok for me.
Websites and businesses do not factor bandwidth for the rogue scraper.
In all seriousness, your suggestion is woefully naive in the face of any serious entity.
You can take a meme (i.e., song/web site) and change/combine it to make a new meme.
What used to be a five minute "put tcpdump on the access point / router to work" job pre-HTTPS everywhere now is a many days worth job of messing around - out of reach for all but the really dedicated, and as a result it is very hard for an user to have actual visibility over where their data flows.
Doing the same on our mobile devices is much more tougher, because we've let feature phones become walled gardens.
Not if the app uses certificate pinning, ships its own version of a SSL library and uses code-signing and obfuscation to prevent you messing around with it.
As for the walled garden part: I agree with the general sentiment, but on the other hand I also see the lengths malware authors go to gather data from people. There really is no one-fits-all solution here, because anything that allows the user to intercept and monitor SSL communication can automatically be used by an attacker! :(
Are there desktop apps that behave this way? Atleast in my experience - I haven't come across anything like this on Linux.
And I seriously hope cert pinning gets adopted by more applications.
As for data being sent from your device to 3rd parties, again, if you don't trust the app's developer not do to that, you also won't trust them not to be sharing it from their end where you have no way to look at the traffic.