Web scraping for me, but not for thee
blog.ericgoldman.org
blog.ericgoldman.org
- LinkedIn sues HiQ, Ninth Circuit sides with HiQ
- LinkedIn pushes to Supreme Court, Supreme Court vacates citing Van Buren
- Ninth Circuit re-reviews and affirms their decision
- LinkedIn moves to get the injunction preventing them from blocking HiQ dissolved, which is granted
- A mixed judgement is finally issued in Nov 2022 ultimately resulting in a private settlement
Where exactly does this leave things at? I feel like everyone loves to cite this case but never goes into the finer details.
Reading a summary of the mixed judgement from Nov 2022, it looks like maybe the issue came from HiQ using people to log in and thus the ToS came into play...? If I'm reading correctly, it looks like the court did eventually side with LinkedIn in stating that HiQ violated the ToS.
https://www.natlawreview.com/article/court-finds-hiq-breache...
Edit: Formatting.
The question that matters is if this establishes any precedence.
So you can't be extradited to the states or go to jail there as it doesn't violate the CFAA (as the supreme court sent it back)?
I guess 500k and no lawyer fees sounds like it isn't punitive given it's, i assume, decently sizeable company?
I'm wondering if we're looking at another MPAA/RIAA situation where they threaten 6 figure sums at individuals.
I lived through that, but this time it's not some 32kbps mp3s of metallica, it's just the entire future of human thought and power itself.
You never really know how a common law justice system is going to act.
Calling this Order on the parties' motions for summary judgment "precedent" would be a mistake. Nor is the Consent Judgment and Permanent Injunction "precedent". The Ninth Circuit decision is precedent.
People in this thread are stating that hiQ was "defeated". Of course. However if "defeat" means a party settling, paying a large sum and agreeing to refrain from certain conduct in the future, then Google and Facebook have been "defeated" many times.
Having "web scraping" remain a "gray area" by limiting the number of final decisions and thereby the amount of precedent might be beneficial to so-called "tech" companies. Putting aside hiQ's predicament, if more of these cases went to trial instead of settling, then we might have some clarity.
We should be thanking whomever funded hiQ's litigation costs. Getting the Ninth Circuit decision was something every web user can be thankful for.
i.e is the common take that people have of "scraping is legal after HiQ vs LinkedIn" just completely wrong?
Edit: oh, I didn't realize you wrote quite a bit here: https://blog.ericgoldman.org/archives/2022/12/hello-youve-be...
-hiQ sues LinkedIn for injunctive relief in the ND Cal., win on its CFAA claim.
-LinkedIn appeals to 9th Circuit, which sides with hiQ on CFAA claim
-hiQ loses its antitrust claims at the motion to dismiss stage
(somewhere in here hiQ goes out of business, but rich benefactor keeps paying its legal bills)
-LinkedIn continues with breach of contract and other claims, wins at motion to dismiss
-LinkedIn appeals to the Supremes, who vacates and remands back to 9th Circuit after Van Buren
-9th Circuit sides with hiQ a 2nd time on the CFAA claim
-injunction is dissolved
-hiQ suffers near-total defeat at summary judgment
-hiQ waves the white flag, agrees to permanent injunction agreeing to nearly all of LinkedIn's demands and pays LinkedIn 500k
With "contracts" of adhesion proliferating, and how impossible it has become to exist in the modern world without acceding to them (something as simple as buying a new SSD involves agreeing to one), this problem is getting worse by the day.
The law is becoming increasingly irrelevant, and more and more we are ruled by one-sided "contracts" from giant companies that are in a position to push them on us.
The craziest example of this is how all these contracts are appearing in the physical world as well. There are stores that actually have a sign indicating that entering the store constitutes acceptance of contract terms (with a QR code that you presumably can scan with your phone to read the contract). I've also seen public parks with the same thing basically indicating that entry binds you to a legal agreement to not sue the park/follow posted rules/etc.
I don't know about the changes made due to this lawsuit, but i reckon they changed their cups to keep ~120F coffee at that temperature longer.
edit: aside: I just realized why i prefer Fahrenheit even though i understand Centigrade. Stuff over 100F is "hotter than i am" It's intuitive in a way that merely remembering 37 is body temp (my dad was born in '37 so that's why i remember) and "anything over 40 is hot" doesn't really roll off the tongue, even though it's roughly as accurate as my ">100F is hot"
And the problem with that has a lot do with corporations. For instance, if you are a pedestrian and get hit by a car and end up in the hospital, in a lot of places in the USA your health insurance will not cover you at all -- you have to sue the driver and get compensated from their auto insurance. The logical method would be for your insurance to cover you and then the health insurance would recover costs through appropriate parties.
It is the same with ridiculous lawsuits like the aunt who sued her sister because the nephew jumped on her and threw out her back. In order to recoup medical costs she had to sue her sister since the sister had homeowner's insurance.
You can't entirely blame the legal system when the corporations are using it to perpetuate the problem for their own gains at the expense of everyone.
Don't get me wrong, I think there is some good in it (mostly around risk calculations), but the entire industry feels scummy.
If a company doesn't accept the customer's contract or won't let you bypass their own, you walk away -- no sale. Other companies will get your business.
which begs the question - what if you need the company's services, and they have no competitors (within reach of you at least)?
The fact of the matter is, the company has a higher bargaining position than the customer.
/s
There are two ways of thinking about what a webpage is:
1) A web page is a billboard
2) A web page is a pamphlet
If a webpage is a billboard, then it is morally wrong for me to paint over those sections of the billboard that I do not like (i.e., using an ad-blocker). This viewpoint is held by those who own webpages (because they want control over it) and by those who cannot change what a webpage looks like (common users).
If a webpage is a pamphlet, then I'm free to cut it up and re-arrange it however I want. Naturally, those with knowledge to cut and re-arrange are more likely to take this view. This viewpoint is more technically correct, a webpage is just a few bits of information handed to me, and to the extent I control my own computer I can cut up those bits and view them however I want.
It's fair to say that Amazon.com contains Amazon's webpage, and that Amazon owns that web page. And yet, I've never once viewed Amazon.com without using an electronic device owned by myself or another non-Amazon entity. Amazon.com doesn't exist on a billboard, it requires the use of electronic devices owned by other people. What rights do the owners of those electronic devices have? Any? At what point do the pixels on my screen become your protected space?
You're not painting over the billboard, just blocking the light reflected off it for yourself. It does not affect others.
What if you only block your competator's ads? What if you replace them with your own? (..does Brave do this?)
What if you block all ads (yours, and your competitors), but in so doing, exploit the inertia of a consumer who is a.) already using your product, and b.) actively sheltered from exposure to the market?
ohh, i got it: just don't block ads for FAANG /s
Ad blockers having paying partnerships can be a problem, but something like Ublock Origin does not.
It also keeps a handy "cost to advertisers" metric you can view. I'm over $40,000 at this point.
Why is it morally wrong to paint over billboards?
But even if we accept that unfounded premise, the equivalent of ad blocking would be to hold something in front your eyes blocking the billboard as an ad blocker does not prevent other visitors from seeing the ads but paint does.
Naturally, they're going to say "web scraping uses resources, stop it!" but then keep web scraping in the background.
To be clear, it's bad behavior, it's just not hypocritical behavior, as it's completely in keeping with what amoral corporations locked in constant battle would be expected to do: maximize benefits to themselves while minimizing benefits to others.
It is a public policy issue also outpaces "competition" which is merely a subject change.
Trying to prevent certain classes of behaviours via legal means is more like trying to prevent certain types of play, by appealing to the referee, while still doing them yourself. Clearly, this often does happen in sports, but _is_ generally seen as hypocritical.
There's NOTHING natural about our economic systems. They're all COMPLETELY made up, let's treat them that way.
(and yes, here it is about 'lobbying the ref')
Genetics however is not only a useful model, it's hard science. You can experimentally find out whether some characteristic is e.g. Mendelian (I'd doubt the greed is, as normally defined).
It got me thinking that to cross the two domains, there is also a meta-concept of cultural viruses ("memes") to which Dawkins applied Darwinian model. Definitely not hard science, but they kind of counter your point that "there's NOTHING natural about our economic systems".
Honestly, what helped me a lot is: Economic systems are more like video games than "nature." Sure, video games can look like nature, but also are extremely malleable.
In football, the rules have been extensively tuned to promote a fair fight.
Perhaps we should do a bit more of that sort of thing in corporate law.
that's the expected cost of publishing something to the public internet. People are going to access it. No one has a right to complain when people access something that was put there for the public to see. Scrappers can be dicks about it too, they can get lazy and endlessly hammer away at some server or repeatedly pull down the same content because they messed up, but we don't need need litigation for that. If something raises to the level of DoS that's already covered under existing laws.
> it's completely in keeping with what amoral corporations locked in constant battle would be expected to do: maximize benefits to themselves while minimizing benefits to others.
Maybe we need to rethink giving some of these corporations the privilege of corporate personhood if they are just going to make things worse for everyone else while only enriching themselves. We don't need to allow parasites and pillagers to take whatever they want at our expense.
It's not always about individual bad actors. You can have lots of small players causing problems. I wonder how many python developers there are right now trying to make their own offline copy of stackoverflow.com.
Wikipedia has a great defence against this. They ask you not to scrape, and at the same time, provide torrents of the data (https://meta.wikimedia.org/wiki/Data_dump_torrents)
There was a hiccup around June but that seems resolved now: https://meta.stackexchange.com/questions/389922/june-2023-da...
What is your opinion on spam email?
Some people don't publicly publish their email address, they instead selectively give it out only to those they want to get email from, but their address gets leaked/sold and abused. Ultimately people who do publicly publish a contact address (email or even a physical mailing address) are basically on the hook for deciding what to do with whatever people send them.
The spam situation got out of hand pretty fast though. The only thing that kept email spam from reaching the level of a DoS attack were blacklists and server-side filtering, and even with those things (plus client-side filters) spam is still a huge problem today. Spam is just a much bigger problem than web scraping. Even the junkmail the mailman delivers to my door has an environmental cost that's much worse than the "harm" of a web scraper's http GET requests.
We have many alternative ways to contact each other online that aren't as vulnerable to spam, but for all of its shortcomings email continues to be widely used because at the end of the day people think giving strangers the ability to reach out to them uninvited is valuable. Anyone can set up a whitelist and trash everything that comes into their mailbox unless it's from an approved sender, but almost nobody does because they want to be more reachable than that.
- it seem a lot of web developer / product manager either do not know or do not care about robots.txt
- some web application are so badly optimized, that some of them are not able to handle more than 1 hit per second at a sustained rate, which admittedly have worked fine so far. But crawlers are persistent, causing the normal crawling activity to cause denial of service for normal users.
You also would not let one football team buy up 60% of the other football teams in the world and merge them into one MEGA-TEAM.
But apparently we still cannot summon the willpower to enforce simple and straightforward 100+-year-old antitrust laws (so instead we make new ones and just won't enforce those either).
Why not? If that mega-team wants to destroy their revenue stream (People want to watch competitive football, as opposed to curb-stomps), they are free to do that.
The whole point of competitive sports is for there to be, well, meaningful competition. All the money in that industry only exists if there's meaningful chances for either team to win. Winning too hard is a detriment to the winner.
That is incredibly different from capitalism, where the whole point of it is for one of the participants to win the competition, to the deteriment of the rest of us, which is why society comes up with elaborate rules for preventing that.
I don't understand why this demonstrates hypocrisy. There is a big difference between crawling the publicly accessible web (which legitimate search engines do all the time) and scraping an authenticated web application or API.
Of course this is not on the level of child labor, nor environmental pollution.
1. OpenAI (etc.) scrape the public web to train and build their models.
2. They use these models to sell subscriptions (make money). None of this goes to the creators of the data used to train the models.
3. They deny others to do what they have done themselves.
If you compare to, say, search engines scraping the public web:
1. Search engines scrape the public web to create their search indexes.
2. They use these indexes to provide search results and sell ads around them. Importantly, these search results direct people to the websites that they have scraped (much of the time), offering opportunity for them to make money.
What about "by continue using this site you agree to this TOS"?
What if you don't have rights for published content?
What if you make your content free and opensource, but don't want big greedy corporation to use your content to train AI?
What about author rights? If I publish my painting does that mean that any corp can sell t-shirt with this content, because "content available to anyone who can send a GET request or use the Internet Archive"?
Terms of service determine when a provider will refuse to service a user's requests. If the website responded to the request with the requested content, TOS is a moot point.
All of the other examples are covered under copyright law. Whether or not copyright has been violated depends on whether the training of AI models falls under fair use. That remains to be decided in the courts, but I think there is a plausible argument that an AI model counts as "transformative use" and wouldn't be a violation of copyright.
If you give away your content for free and expect ads to sustain you, that will start failing once others get the value out of your content without seeing the ads. Examples are ad blockers, answers embedded in Google results, Stack Overflow clones, and things like ChatGPT.
If ads weren’t your business model you wouldn’t be using revenue from it.
The other issue is scale, and I don’t know how to address it.
It’s easy for someone (say the government) to have a friendly policy and say “you can use dig in a park” thinking it’s useful to campers and such.
But when someone shows up with a professional strip mining crew, things are different.
If you run a site providing quality information for free, making money off book sales or professional services or such can be a good living. Even if answers end up in the Google answer box, more complicated stuff or analysis still requires a visit to read and people can start following you from there.
But if ChatGPT or whatever can “read” your stuff and give out 80% of the value without anyone even knowing it came from you, you’re screwed. Your business model no longer works. Any kind of “give away good information” business model fails. Same issue artists are now seeing.
And I don’t know how you fix that without some kind of ban. But unless every country everywhere enforces one… you have to work with the lowest common denominator and lock all your content up. No web search. No Google answers. No chat GPT. “Please don’t scrape me” in robots.txt won’t work.
Unlimited scraping makes some of privacy regulations moot. Such as right to erasure (ability to delete personal data from a platform).
In case you really need an example to elucidate, consider reproducing an image. A scraper can quite literally accomplish that, trivially; a great artist would still be limited in multiple facets of the recreation, such that even one with the best memory and hand would find themselves far short of pixel-perfect.
Let's suppose an embarrassing image of Person X is shared on Facebook and Person X uses their right to erasure with Facebook to delete their profile. Facebook has no control over the folks who may have downloaded or screenshot-ed that photo and turned it into subsequent memes. Likewise, if someone straight up scrapes and re-shares, that's not Facebook responsibility.
What I don't want to see happen is for:
1. Facebook to make it somehow impossible for anyone to ever copy or screenshot that or any photo, preventing anyone from ever doing anything with photos on Facebook without Facebook's explicit permission. This would seem to be quite the loss of user agency for very little society wide benefit (also, how would they do this?)
2. Facebook to somehow "control" that photo so closely that Facebook is able to remotely revoke folk's copies and screenshots of said photo in the spirit of "abiding by a persons right to erasure"; that'd be a huge overreach, but seems like the only other way to approach this (though "how" is also an open question).
Even asserting that "unlimited scraping makes some privacy regulations moot" seems like an implication that we can only have privacy laws by going towards situation #1, and that doesn't seem accurate given that folks can use existing privacy laws to remove content from any distributor (as long as they're compliant).
I don't think a paywall would fix this. One paid account is all a scraper needs. It couldn't really even be rate-limited if it's just "reading" articles as they become available. After the data is acquired it can be dispensed. If directly posting it violates copyright, then obscuring it behind AI will do the trick just fine.
And unlike now to sign up you have to agree to a very enforceable EULA.
So instead of going to court with “FunAI read my public website and is making money off it which I don’t think that should be fair use”, you have “FunAI violated a contract they signed and committed fraud by lying on signup”.
Seems to me that’s much easier.
There will always be people who get the content for free somehow. You don’t have to stop 100%. Even stopping 95% would be a lot better than the current 0%.
The original problem with copyright was that the website owner's content could be duplicated elsewhere, and thus, violate copyright (as well as suck away the web traffic and presumably lowering revenue).
The new AI issue is not that the content is duplicated elsewhere, but that the knowledge contained in the content is "learnt", and used to produce a different work (totally copyright free - in the truest sense, as it is original). An example would be a recipe website. The site owner could've painstakingly collected recipes from the literature, and cataloged, labelled it, etc, making searching and such easy. But the recipes themselves are not copyrightable, only the expression of the recipe.
So given this info, the AI scrapper now has a large labelled dataset for which to learn from, and to generate new recipes. These new recipes do not violate _any_ existing copyright, as they are entirely original in expression.
I say, as an AI advocate, that the old business model of recipe hosting is destroyed by this new AI, and legislating it to remain by legal means is just fighting against the tide. After all, the world doesn't have a unified jurisdiction, and the internet is world wide, so any would-be violators could just as easily move to a different country to operate.
I have two thoughts:
- EULAs aren’t written for companies to sign.
- I think EULAs are garbage anyway. They’re completely one sided and in most cases probably illegal or wouldn’t hold up in court if anyone actually had the resources to fight one.
Imo, the burden of ensuring someone has read and understands a EULA should be on the company who creates it and they should not be enforceable unless they can prove the person understood the EULA entirely before accessing the site. EULAs are not a business agreement. They’re some kind of corporate pseudo-law companies try to attach to the usage of a product. But what other product in the world has a big list of rules that come with it that way how you can use it (or be sued)?
So how does this all come back to this “company vs company scraping”? If you put it on the web, and you don’t have REAL copyright on the content (that is, you didn’t make it yourself), you have no right to protect it from “theft.”
PS yes, I know John Deere doesn’t let its customers work on its tractors but that’s some bullshit too.
In the case, Verio was calling Register's API for a purpose that Register disallowed. However, they only supplied the "contract" text that declared the restrictions after the call was made (as part of the API response I think).
The court did in fact agree that this was too late: If you can only find out the terms and conditions for an API call by making that API call, then this is a "shrink-wrap agreement" and the terms are void.
The restriction the court made on this was that it only applied to the first time they called the API: Verio has employees who can be expected to have common sense. So after having called the API for the first time, they had an opportunity to read the text and become aware of the restrictions. This meant that for all the subsequent API calls they made, Verio's staff was aware that they were doing something that Register explicitly forbid, but did so anyway - hence the court ruled that as "breach of contract".
The important point here was that the courts never abandoned the principle that an individual must be aware of a contract's conditions before it can enter a contract - the case was just shooting down a situation where a party was pretending to be ignorant of the terms when they really weren't.
They "released" a dataset scraped from public domain stuff under a license that restricts how people can use it
Kinda. Yes, Facebook says that content belongs to users (otherwise they'd have harder time explaining they are not liable when it's illegal), but users also agree to give Facebook “non-exclusive, transferable, sub-licensable, royalty-free, worldwide license to use any IP content that you post on or in connection with Facebook.”
For example, if a user deleted their* content, Facebook can still use it and show to their friends. That's why it's "kinda".
I don't think it is correct. If you asked Facebook to remove your data from platform, it will be a GDPR (and probably CCPA, etc...) violation for Facebook to not delete your data within 1 month.
That said, ever since I added some invisible unicode characters into my name on LinkedIn and can see the fraction of spam that I receive that addresses me with a bunch of `?`s where I put the invisible characters is rather astounding. Probably half the garbage email I get has clearly scraped my name from LinkedIn.
All that to say: if you're someone who feels affronted that you aren't allowed to scrape LinkedIn, and you try to garner my sympathy about the situation, please accept my invitation to get stuffed.
I shared info on LinkedIn for use by myself and other linkedin members, and for linkedin to use in their various products. You can't be pro-privacy and simultaneously believe that my putting info on LI has granted the entire world license to do with it what they will. Including running research programs or ML stuff or whatever hiq's business was, selling it to spammers, and using it in products by other companies that I've never heard of.
You gave information to LI, trusting that they wouldn't disclose it to people you wouldn't want to have it. Meanwhile, untrustworthy third parties were free to view that very same information. Blame LI for being casual with data about you. Blame yourself for trusting LI. Don't blame others for reading what they were free to read.
HiQ's claim is that LI shouldn't be allowed to prevent HiQ from reading data that I had no intention or interest in sharing with HiQ, simply because LI makes it available to other LI users. It's wildly anti-privacy to claim that because I shared data with LI, and LI put some piece of it on the website, that HiQ has a right to read it and use it for arbitrary things.
When we get facts separated from the model it will be enough to have just one article about, say, landing on the Mars for the model to be able to actively use this new knowledge. In this case 'AI providers' will need to cooperate with one news agency, buy right for only for a few classic books. Judging by how fast everything is moving this will happen within a couple of years if not a few months.
If I were to guess, it probably originated in an edgy, hyperbolic statement like “the new AI empires will be built on the free data from the old ones”, uttered by someone scary and important, perhaps Altman.
Through the rumor mill all the feudal lords in the data-kingdom, like Steve Huffman, Musk etc are freaking out and are fortifying their defenses against this new perceived existential threat, sacrificing whatever minuscule openness remained in the process.
To me, it’s obvious that a new copyright interpretation is needed for commercial AI applications, which would alleviate the gold rush panic. However, the regulatory bodies (ironically, eroded by the same corporations), don’t have any teeth left to do anything about it, so it’s a Wild West, and they all know it. The legal defense against scraping has nothing to do with principles, even less so these days, it’s all just about abusing the legal system for dominance games.
Personally, I think the threat is overblown. Yes, AI will force eg Google to improve their shit ranking of recipe sites, but AI can’t provide up-to-date stuff people care about, so Google, Facebook, Twitter etc will keep providing services just like today.
Twitter barely works on Tor. Reddit barely works if you're logged out. Reddit is pushing their app really hard and obviously going to drop old.reddit one day when their internal API finally breaks it off.
If you make revenue from advertisements, having a web page that can be interpreted by any machine besides your own code is wasting money on freeloaders. Be they 3rd-party clients, Tor lurkers, or scrapers for search engines and AI.
https://archive.org/details/2015_reddit_comments_corpus
(oh there is a 2TB torrent from 2005-06 to 2022-12)
Unless you think of feudalism or indentured servitude as Wild or Western.
It’s a place of a few royal palaces, with walls & moats, surrounded by the masses who submit the value of their creative & social efforts through little slots, for the right to receive in kind from others, but filtered & diluted with whatever the royalty wants you to see and hear.
The evolution of aggregation sites is not on a path that enhances or maintains individual freedoms or independence.
Amazon charging suppliers a fee for NOT using its services ought to be an Onion article. [0]
[0] https://www.theverge.com/2023/8/16/23834653/amazon-seller-fe...
Regulators are out of control, yes. But that does not change the fundamental nature of the internet. Especially as it pertains to something which is purely data, copyrighted material. Amazon is not stopping filipino and bulgarian kids from downloading mission impossible any time soon. Nor is it stopping people from bulgaria from paying for AI services in the Philippines which are trained on American data.
Any more than a Duke or Earl wants to kill off peasants or stop them from playing and singing with their families in the evening.
The lords just want to extract most of the value that they create, most of the value that they spend, and let them see what the lord wants them to see, an increasing percentage of time.
The takeaway for those of us in the trenches who feel we have as much a right as Google or Bing to crawl websites, is that we are not subject to arrest and imprisonment, just tort law. The ruling has apparently caused the major social networks to switch from criminalizing web scraping to putting most content behind a login and threatening to sue small companies out of existence.
[1] https://www.davidrevoy.com/article977/artificial-inteligence...
No. This is not true. Quit saying this. A small fraction of the world’s knowledge is on the Internet, much less scrapable.
There must be a good chunk of the world that doesn't have any laws forbidding it. This isn't under the jurisdiction of WIPO or anything like that, it's just a completely insane evolution of anglo common law
Courts feel very fine doing the bidding of private decision-makers on what should be questions of public interest, thank you very much. Reminds me of this joke:
“Wake up, sir, wake up, you have shat yourself!” — “I’m not sleeping.”
For that matter, who will build the next LLM if anyone like Elon Musk can slap terms of use change on an account because he wants to stop fair use. This is anti innovation and means challenging all these monopolies- Google in search and twitter in Social media will be all the harder.
This needs to get to SCOTUS fast for clarification as it's evil, wrong and will kill innovation.
If we can't be clear on Web Scraping, do we really have any hope for more clear rules on more complicated things like AI or privacy or copywrite?
Who was this "litigation funder".^1
Perhap the funder wanted to settle, even though hiQ could have prevailed at trial on the CFAA claims. Arguably, hiQ's goal was not to establish new CFAA precedent, it was to compete with LinkedIn. It wanted a Court to order LinkedIn to stop blocking its access to public data.
This took too long. hiQ ran out of money. (Or its funding source(s) ran out of patience.)
If hiQ had the funding to go the distance, then it's impossible to predict what would have happened.
It's sad that the top comment in this thread believes LinkedIn sued hiQ to stop it from scraping. That's not what happened. hiQ's "business" was going to fail because LinkedIn was effectively blocking its access, even when hiQ was using proxies and mechanical turks. hiQ filed for a declaratory judgment against LinkedIn: hiQ sued LinkedIn.
Here is the Order for anyone who cares to read it.
https://ia600100.us.archive.org/29/items/gov.uscourts.cand.3...
1. According to the Order, the funder is identified in some correspondence filed as an exhibit. It is not clear from the Order whether names in the correspondence were redacted, i.e., it's not clear whether it's public or private information. I'm assuming the former.
At time of writing, I believe that top comment is mine - and I'll be clear that I wasn't saying that I think that's why LinkedIn sued HiQ.
In fact I don't particularly care why they did it, I'm just interested in what precedents the case has set and whether the typical comments thrown around regarding the case have any actual merit.
There's a comment replying to mine that has a better breakdown of it all that I'd trust over my own anyway. ;P
What are the "typical comments thrown around". Perhaps could provide some references.
Umm, no; author needs to study the word "hypocrisy" more deeply than a cursory glance in the dictionary.
Doing something to others, while defending against the same thing, is not hypocrisy.
For instance, soccer player isn't a hypocrite for defending against the ball going into his net, while trying to put it into the other team's net.
A soldier on the war front isn't a hypocrite for shooting, while also taking cover and dodging bullets.
These subjects are not hypocrites because they are not acting in one way, while preaching that they, or others, ought to be acting in a different way.
Microsoft would be hypocrites if they published an official statement such that nobody who engages in web scraping has the right to defend their own site against web scraping, because that would not resemble their actual behavior and position which could be inferred from their behavior. (Is there such a statement somewhere?)
For hypocrisy to take place, you have to actually preach that you and others should behave in a certain way, and then not actually behave in exactly that way. If you only act, and don't preach, you cannot be a hypocrite.
Moreover, your team's net is not the same object as another team's net. If a soccer player loudly professes "it is morally wrong for anyone to kick the ball into our net", but then kicks the ball into the other team's net, that is not hypocrisy. His statement references only his own net; he didn't proclaim that it's wrong to kick any ball into any net whatsoever.
Hypocrisy is the professing of a moral position to which one does not conform.
If you don't preach that web scraping is wrong, and engage in it, that doesn't mean you can't defend your site against it.
Everyone shits; it is not ipso facto hypocrisy not to want it on your lawn.
Hypocrisy could be involved in the way you articulate your wish not to have it on your lawn.