The Scraping Problem and Ethics
blog.osvdb.org
blog.osvdb.org
One of the things that stands out is called out in this article. The people involved really want this information, so much that they are willing to expend time and effort to construct scraping bots and what have you . Why not just buy it? How is it that someone gets a request from their boss to get some information, but their boss expects them to get it for free? Can you imagine if they said, "We need pens, pencils, notebooks, staplers, the works for the office here. Oh and you can't spend any money getting that stuff, just get it here." Would they construct some elaborate raid on a nearby Office supplies store using a mercenary army of criminals? Why do that with information?
We did an experiment where we would 'grep the web' for you, basically run a regex over a multi-billion page crawl, give you the first 50 results for "free" and you could buy the complete set. I think we sold exactly one of those.
It is a weird thing, the OP captured it perfectly.
[1] Its a violation of the terms of service.
I have absolutely no explanation regarding McAfee though, considering the billions McAfee and its parent company, Intel, makes in revenue yearly.
Here you have a great case study, about an organization that tried to do a volunteer model, and it didn't work. Then they pivoted to a commercial model, but fundamentally they still believe in a free tier. But they have to cripple that free tier pretty thoroughly, and even still people abuse it.
I have a product I'm working on, that some people are apparently willing to spend lots of money on. Ideally I would have some kind of low tier, so that people without lots of money would be able to use it too. But I can't figure out a way to segment the product so that everybody pays what they can afford without bad apples abusing the low tier and ruining it for everybody. The result is that I may end up only selling it to customers with deep pockets, even though the product is much more broadly applicable.
I have worked at lots of companies where I had a monthly budget of 10k+ that I could spend on whatever I wanted, but if I wanted any sort of complex deal (can't just put on CC with a line item) -- had to bring in legal and other groups -- instantly killed any interest.
"Licensing is based on the data needed (e.g. all of it vs subset), how it is used (e.g. internal only, external, product integration), etc."
What a goddamn horror show. I simply want a product, I want to pay for it, and I want to use it. Turning on Dropbox for Business was a decision made in about 5 minutes... "You all like it, already using it, awesome! I will get team setup." -- 5 minute later I had given Dropbox $3800.
I really think they are getting in their own way for no benefit. They have created a very high barrier to EVEN HAVING A DISCUSSION about buying the product. So, if I don't know exactly how will use it -- I can't purchase it. Stupidity.
What is the difference in value between a CD with the latest release of Ubuntu burned on it, and the download? download and a bootable Flash drive?
There is a great experiment you can run which goes like this; At one end of an athletic field, place a chess board with a queen on it on one of the squares. At the other end of the field have a table where people can get a quest. Offer to pay a person $5 if they will walk to the end of the field, note where the queen is, and come back and tell the quest giver. At the mid point of the field set up an information seller. They offer to sell you the location of the queen for anywhere between 20 and 80% of the reward price.
This simple experiment lets you see all sort of mechanisms in play that control information value. On the one hand you can see the range of values people apply to their own time (acquisition cost), their willingness to retain value (do they then go back mid-queue at the sign up table and start offering to sell the information for some fraction of the price to anyone?) At what threshold to people start trying to break the rules (a notion that is similar to price inelasticity but has a component like the 'black market demand').
Interesting questions to be sure.
For example:
i) Are search engines web scrapers?
ii) Should search engines pay the scraped sites if they are charging to access their indexed data? probably some of the scraped sites has a specific license forbidding the search engine to sell their information in any way.
iii) Regarding Internet policies, is it fair/unfair that a site has a robots.txt configuration to avoid being indexed by a search engine other than Google? I would call this "search neutrality".
Any reputable search engine will respect robots.txt.
Take the classic example, a search for 'bilbo baggins'. What that is, is a request to identify documents on the web that have referred to Bilbo Baggins and return their locations.
1) It is absolutely true that you could sit down at your computer and look at each site, from aol to zillow, read all their pages, and note the ones that mention Bilbo. Then you could go back and order the list by the ones that had more of the information you were looking for to the ones with less useful information. Along the way you would find some sites that would not open up to you unless you had an account, those you could not visit.
2) A search engine can look at all the sites, it can note which sites mention Bilbo along with a bunch of other terms and can essentially "pre-compute" that list you were looking for. Along the way it will find sites that, through their robots.txt file, will say "We'd rather you not look here." and it will respect that, thus not indexing those sites.
In both cases figuring out which web pages have information about Bilbo on them is creating 'new' information out of existing data. You can do it on your own and it will cost you time, or you can do it with a few thousand machines and it will cost you money. Either way you get a list of possible sites.
That list forms a distribution, where there are a lot of sites that don't care one way or the other if you read them, sites that won't show up because they asked to be excluded, and sites that will show up because they paid to be included. Some sites really want you to find them, some sites really don't. A good search engine caters to both types.
ii) If a search engine was taking the page, copying it, and then showing that instead of showing the page (this is what got Google's news product in trouble) then its pretty clear that they should not do that. But in terms of location information? The sites themselves derive a huge economic benefit from being in the index that isn't reflected at all back to the search engine that sent traffic there [1], so on a pure economic basis the search engine is on the losing side of that transaction. However, the marginal cost of additional transactions is small (search engines are general purpose) so they make a small amount on large volumes.
To put your question in more specifics, where is the economic value in the list; bobs middle earth atlas, wikipedia entry on bilbo baggins, middle earth web ring, imdb pages on characters in the "Lord of the Rings" movie.
Is it that Bob has an atlas of Middle Earth? Or is it the list itself? Who made the list? Bob or the search engine? (or some human curator of a bookmarks page[2])
iii) It is completely up to the site to allow or disallow access to its content by search engines. Some sites do only allow themselves to be indexed by Google and they find they get less search engine directed traffic that way. Some sites don't allow anyone to index them and they get no traffic (sometimes they are surprised by this, sometimes they don't care, sometimes they are angry that the only way for people to find them is to be in a search engine index)
[1] Google broke that by creating AdSense for Content and created a pretty interesting conflict of interest for themselves.
[2] Good luck finding a book marks page these days :-)
Not sure I understand, is there a whole ecosystem of web businesses that feed off for free of your search engine, or you meant from several legit search engines or ... ?
Crawling others' websites then selling ads. <--- Blekko
Scraping others' websites then selling ads. <--- ????
Selling access to user-generated content. <--- OSVDB
Amusing to watch these folks argue about ethics.
Who owns the copyrights in this data? Surely not the one who is demanding that you pay a license fee. These "services" are middlemen, plain and simple.
This might be why McAfee was wondering about how much manual curation is done.
Maybe that is the only possible theory of how OSVDB could assert any rights in the data (and only in a select few jurisdictions).
So what drives the folks at McAfee to do this?
Maybe it is the same thinking that drives programnmers to not want to write code.
"Don't reinvent the wheel."
"Code reuse."
"Use a shared library or a scripting language with batteries included."
Why crawl the web when Google has already done it for you?
And so on.
Personally, I do not have trouble understanding why McAfee would do this.
What I have trouble understanding is why OSVDB would think they could crawl some public data and then charge a fee to access it.
It is the "sale" of "free" information that puzzles me.
By all means go ahead and try, you may well recoup your outlays for compiling the free data and even make a profit.
But should we really be surprised when someone does not want to pay?
The argument that if you really want this information, they should pay for it doesn't work. Basically, you are punishing people for automating manual labor. The fact that you can hire bunch of people and tell them to manually copy data from the website and achieve the same result means that it shouldn't be any different from automating the process itself.
Terms of service on a website is not the law. If there is something you don't want people to have access to, then don't publish it online at all.
Sounds like you didn't advertise it right. I've been looking for a "grep the web" service for a while.
We're going to use https://builtwith.com/ to achieve that exact thing. It will cost me $295 : https://builtwith.com/plans
For that $295 I will get a list of all domains using a rival technology... a list of sales leads.
I'm still short of what I want to have... a prioritised list of sales leads.
So I will write my own scraper to go through those results and scrape every one of those so that I can pull some info from the HTML page to tell me how large those customers might be. I'm not aiming for the largest (easy to find, costly to win), nor the smallest (time consuming and pointless to win), but the median.
I would definitely pay for a "grep the web" that allowed me to match pages by text signatures in the HTML, and then extract part of the DOM as values, and return the list of "url + extracted values" for the matching hits.
I'd consider that to be worth similar amounts to what BuiltWith are chaging, but I'd add more and would go up to $500 assuming that the results come with the extracted values as a CSV file of some kinda and the quality and completeness of the report is high.
People will pay for "grep the web", especially if you sell it to them as something they know they really want: "sales leads".
The internet isn't about "don't take my stuff", it's about spreading that stuff around. I'm confused by people who want to make their data public, but want to control exactly how people access it.
Is that a bad thing?
>The internet isn't about "don't take my stuff", it's about spreading that stuff around.
Try asking Google if they want to "share" their database of crawled data.
>I'm confused by people who want to make their data public, but want to control exactly how people access it.
Me too.
I'm not sure if this is a joke. So, I'll refrain from replying.
That is not what I said at all. I don't believe that more competition necessarily leads to better product/service. In fact I believe that in most cases it does not. In capitalist economies, companies try their hardest to avoid competing. Competition forces companies to reduce costs, and not necessarily increase quality. The quality of a product is not some number that people can read and go "oh yeah this product is better". Marketing people try hard to invent such pointless numbers (e.g. Megapixels in cameras.. horsepower in cars , etc etc). Also many CEOs don't have the first clue on how the product is actually made, much less increase its quality. They rely on these same 'marketing numbers' that their underlings serve them with. So, they too can go "oh yeah this number is increasing so our product is getting better".
It seems like a lot of people are brainwashed into believing this free market utopia where things just automatically get better because everyone is competing and the customer is this genius who can figure out which company is delivering a better product.
Its sort of like thinking "Well if I'm nice to everyone, everyone will be nice to me.". And then you realize that the real world is a dark place filled with assholes, where slavery is still rampant and many of the goods and services we consume are dependent on the exploitation of natural resources or other fellow humans.
Sorry if reading all that bummed you out. I'm really a quite a cheerful person :P
I talked about this a while back: https://news.ycombinator.com/item?id=6572937
Someone decided to ban all bots from accessing their site except for Google and Bing. So much for "if you're so worried about Google just use another search engine".
You mean like TV? or Radio? or print....
McAfee were just being the usual assholes, but the first guy mentioned in the blog post could have been converted into a paying customer if the pricing scheme were clearly outlined on the web. Since "Contact us" account tiers are usually reserved for the very high end, he probably assumed that it would cost him an arm and a leg.
It shouldn't be too difficult to come up with a handful of moderately priced account tiers, each targeting a different type of customer, as well as an open-ended tier at the end ("Contact us") for those with special needs and deep pockets.
I understand the sales psychology in making sure you enter into a proper discussion with people to make sure their needs are right, and extracting the maximum consumer surplus from them. But there is a non-trivial pricing point where this makes sense, and by not having any anonymous-sign up, you're cutting off your sales curve below this point.
I think it's very easy to convince yourself it's easier just to sign up customers over the phone, but doing so without at least testing takeup of simpler tiers is an incomplete picture.
In most cases I don't believe these services are forcing customers to talk on the phone because they think they will convert them. Chances are that sign up isn't as easy as a 1-2-click, and involved set up and understanding of the customers needs is required before services can be rendered.
If your signup process can be automated and you only charge $25/mo, and you still tell people to give you a call, then you will lose business because 1) the friction of a phone call is worth more than the difference between your offer and a competing offer; 2) it's impossible to tell whether your offer is worth the friction in the first place, because your pricing is unknown; and 3) people just assume that you'll charge $2500/mo because that's the usual price point where people say "contact us". If someone else comes along and offers an inferior product for $35/mo, they'll get all the business because monkey psychology.
What I am saying along with that is that - by doing this - you're ignoring a lot of the market at (or even just below) your price point. That might be OK - as long as it is a conscious decision to do so, including acknowledging that your market is completely above that point.
In the case being discussed here, it would seem that there is an interest below this price point.
Fuck that non-sense of pricing based on what ever they think they can scam me out of
However, the lack of up-front pricing data isn't being used to justify "I won't do business with you." It's justifying "I'm going to take all your stuff anyway."
A conversation I've had a few times:
"We need it to do $THING_IT_WON'T_DO."
"In that case, it probably isn't a great fit for your needs."
"You don't understand. I won't buy it if it doesn't do that."
"I think I do understand. That's fine. You might consider trying $COMPETITOR, although you should know their minimum spend is $1,000 a month."
"That's outrageous. You have a $29 plan."
"Yes. So you should go with the competitor if that requirement is worth $971 a month to you."
"No, I want to spend $29, but I absolutely need that."
"I understand where you're coming from, but we do not offer that feature, and if we did, we would charge prices close to what our competitor does for it."
"You're not working with me here."
"I'm trying to find a resolution which works for you, but including that feature at $29 doesn't make business sense for me, so I won't do it."
"Put me on the phone with your boss."
"I'm afraid that isn't possible, as I sort of run things around here."
"What sort of businessman turns customers away."
"You're not a customer. If you were, you would be purchasing a product I sell for the amount I sell it for. That isn't happening. That's fine. Have a nice day."
I'm saying rather this argument won't find much favour here unless everybody is a hypocrite.
The fact is, if it's harder to buy something, would-be buyers choose another route.
Imagine something of trivial value that was very arduous to obtain. Say, the scores to last weekend's football game are only available via sending the NFL 2 cents taped onto a postcard then getting a user account and password back in the mail. Yes, you could apply for an account and who cares about the 2 cents? But most people wouldn't and who could blame them? There'd certainly be a market for pirated 'score data'.
But, if you just put ads on your site (the NFL does), voila! You make the same amount of money and people are happy to use your service. Netflix, Hulu, iTunes Store and Spotify all get this. If you make it a pain to do something, you can't feign surprise when everyone goes around you.
Is it legal to circumvent something just because it's arduous? No. Is it ethical? Not completely, but downloading pirated football scores wouldn't keep me from running for office.
You should be rate-limiting how many requests free API users can make, like; Twitter, Facebook and every other Internet provider does via their API. Make it harder for people to obtain the information (to the best of your abilities) and paying will become more of an option because they won't go to the trouble of scraping it if it's impossible and will take a lot of time to do so.
Think of your offering as a car. You currently have no car alarm or immobiliser, if you install a immobiliser and car alarm, you will make it very hard for a thief to steal your car.
At one point the effort to circumvent would cost more in man-hours than just buying the product.
When I used it (which was almost a decade ago) never ran into problems, plug a list of 10,000 proxy and scrap away.
Not condoning that, which is a bit hypocrite of me, at the time I was mostly doing what I was told and I thought I was clever. Now that I'm in a position to have a positive impact, I do buy data and pay appropriate licence fees on all software/data purchase, which still baffles some of my programmers who constantly ask "why not crack it?", "you know I found a .zip on Google with the data, why buy it?", and so forth.
I don't know what in programmer culture makes it so hard for us to pay for something, some people put some effort behind that software / data collection, and it's only fair to pay them.
Maybe I want to use your data casually once, and I don't want to sign up and give you all my contact details and subscribe to your annual plan with all the other optional extras.
Tough shit, you say? I'll just steal it then, and not because I can't afford it, but because you're making it hard to pay.
Scraping is not necessarily a no-victim situation. Even today after this stuff has gotten cheaper, you're costing them bandwidth fees, and likely increasing their server storage and CPU fees if it's on a metered hosting service, which is quite likely nowadays. If you degrade their site's functionality, you may chase away paying customers.
We need not hypothesize crazy third-order effects; you are taking money out of their pockets by the act of scraping itself, independent of the question of the value of the content.
"What about Google? etc." - robots.txt-honoring scrapers that don't hammer the sites at least have a plausible claim to permission. Scrapers are quite likely to be ignoring the robots.txt.
While technically correct, you are conflating the issues, because in none of the cases (that I've seen mentioned so far in this thread) the problem is with bandwidth/storage/CPU costs of retrieval to any significant extent.
Instead, it appears that almost all of the costs are incurred before retrieval: curating, sorting, etc.
I'm not arguing that it's okay, but it's just as much not stealing / thievery as downloading movies or music isn't.
I still hesitate when it comes to thousand dollar licenses when it's for my personal use , though.
A better method would be to set "default" pricing (something high but not ridiculous, that could easily be negotiated downwards if they contact you) and make access beyond a few requests a click-through (or better: have them respond to an email before progressing further) where they agree to that pricing if they are using the information commercially.
The problem I see is with giving out bad data while leading people to believe they have obtained useful information (that they then embarrass themselves by using/re-distributing).
I did misread the grand-parent post though: his was suggesting the bogus data was the message, and I read it as handing out fake data for the scraper along with the message to be seen should a human be looking.
That sounds pretty ludicrous, do you have anything to back it up - caselaw, settlement report? It would be analagous to serving a fake image to combat hotlinking; or a fake page to combat framing.
But leading someone to believe they have correct data when what they have is potentially embarrassing when used could be something they'd take objection to. Even if not there are two other points of risk: your reputation if something goes wrong and you accidentally give bad data to your paying clients and your reputation if someone, paying or otherwise, shows off the bad data as "the sort crap these people try to sell".
I don't have any specific references, but it is something I would be careful of as there have certainly been similarly ludicrous (IMO) cases on unrelated matters in the past (yes the right side would win, assuming they can afford to).
I may be being too cynical here, then again maybe not...
I did misread the grand-parent post though and this isn't what he was talking about. He was suggesting the bogus data was the message, and I read it as handing out fake data for the scraper along with the message to be seen should a human be looking.
Or is that about as ethical as spiking trees to prevent illegal logging?
It could be like spiking trees I guess, but that depends on the potential for harm. I just looked it up and was surprised to find that only one injury has ever been reported due to tree spiking. I guess it's a better talking point than actual tactic.
And like lukejduncan said, this is definitely done in practice.
As someone who sees both sides… I don't know what to say. I'm running test landing pages on Heroku with New Relic that pings the sites every minute to ensure the dynes keep spinning and my users don't experience downtime. While I'm careful to stay within fair use, this is at best obnoxious, because if everyone did this, Heroku would certainly need to redefine what's free. From my POV though, I am a bootstrapped entrepreneur and supporting 5 landing pages. I simply don't have resources to pay for a dyno and test everything I have in my head, especially not combined with the many other resources I'd need to start paying for as well.
Or consider the kid in Florida who used Parse' free account for hundreds of thousands of users. [1] (The article was on HN a few weeks ago, this was not it's central point, just something I took away relevant to this comment.)
Part of the cause I think is that we live in a world where we're so used to having things be free, it becomes an entitlement. Another is that all these examples of start-up hacks and hustle stories, we kind of laud, don't we? Everyone talks about how Airbnb scraped Craigslist and got a huge boon that way, but few in critical tones. Should we? Or is that how competition and new products get created (i.e., if the scraping hadn't happened, perhaps Airbnb and the whole sharing economy would be less successful today).
These are philosophical questions, and I don't really have a solution, but they are things to think about.
[1] http://pando.com/2014/04/30/how-a-florida-kids-stupid-app-sa...
How is this relevant? Well I would have much preferred paying for scraping, than trying to learn some new api. It increased the transaction cost. If you further have to negotiate with a partner company, that sounds like even more transaction cost, from having to send emails back and forth and the mental effort of negotiating.
Freelancer sites has a lot of offerings for web scraping but this niche has its own issues.
I dontbhave time to look but couple of things striked me as odd. Is there a reason you dont lock down your request limits? Also why dont you secure access? Allow free accounts to be made and issue apis keys for them, at least that way you can much more easily rate limit access to the apis and heavily rate limit front end web requests down to reasonable numbers.
user-agent: *
disallow: /
user-agent: Googlebot
allow: /While I don't necessarily agree with the concept of only allowing google to index your site, comparing a search engine which feeds you business to a company reselling your data with no attribution is not really fair in my opinion.
Of course, forging your user agent, disobeying robots.txt or scraping after you were told not to, is wrong ethically.
This doesn't even rise to the level of what Weev did.
I'm not sure whether there is a legal precedent though, in some cases you could call it a denial of service if the requests are not rate limited, and in other cases it might be considered an inappropriate access (see Weev, though he eventually won appeal).
Even calling this "bandwidth theft" is quite the hyperbole—if the server can't handle the bandwidth, then rate limit the requests.
I think if you're serving out pages to the public, you don't really get to tell me what kind of browser I'm allowed to download it with. As long as I'm speaking HTTP, it seems fair.
Sadly, that law has been slowly creeping against this mentality... Lately I feel like I'm some old internet hippy with these views. On a site called "Hacker News", no less.
I guess it's because everyday more and more people on here are finding themselves on the other side of the fence, i.e. finding that some of their users are ripping off their content/site.
Sell a product or service.
https://web.archive.org/web/20130714002216/http://www.osvdb....
I guess they were hoping no one would notice the subtle but substantial change to their service
I would avoid the word expansive. Too similar to expensive. Especially when emailing people who are not native English speakers.
This is straight from the open security foundation website:
"We believe that security information and services should be easily accessible for all who have the need for such information and services"
Aaron was scraping private resources to share publicly.
The system is changing, Aaron helped.
Since then, JSTOR has released these documents freely themselves[1]. They have also stated that this was their intention all along.
[1] http://about.jstor.org/news/jstor%E2%80%93free-access-early-...
I feel bad for OSVDB from a sysadmin perspective, but if Aaron's case was so polarizing for essentially the same thing, why isn't everyone jumping on the Hate Train here?
This is naive. Just because research is backed by public money doesn't mean the publications are automatically free to the public. If your argument was valid, you could use it to demand access to the emails of every FBI employee. Just because something is funded by the public doesn't automatically make every component of it open to the public.
The article is basically saying "I want to charge people for information I have made public online, they won't pay, so they are obviously thieves by refusing to do it manually by hand like they are supposed to."
Gimme a fucking break here. If you don't want the information to disseminate, DO NOT PUBLISH IT ONLINE.