Scrape like the big boys
incolumitas.com
incolumitas.com
Rate Limiting by login,
Limiting data to know workflows ...
But our most fruitful effort was when we removed limits and started giving "bad" data. By bad I mean alter the price up or down by a small percentage. This hit them in the pocket but again, wasn't a golden bullet. If the customer made a transaction on the altered figure we we informed them and took it at the correct price.
It's a cool problem to tackle but it is just an arms race.
If information is your competitive advantage maybe you shouldn't have it on a publicly accessible website, and should instead stick it behind an API with pay tiers and a very clear license regarding what you may do with it as an end user.
Note, a simple sign up being required to view a website makes it not publicly available information any longer and you can cover usage, again, in a license.
Then you have a whole bunch of legal avenues you can use to protect your work. Assuming you can afford it that is.
Seriously? Do I need to explain why a song doesn’t enter the public domain when it is played on the radio?
At some point people are gonna have to accept this.
This sentence added nothing substantive to your comment, and made it rude; could you please be a bit more polite in the future?
Yours is a somewhat orthogonal point and one I don’t entirely disagree with.
On the opposite end of the spectrum might be a photographer's website containing a gallery of their sample work. The fact that the gallery is openly published doesn't represent a relinquishing of copyright over those images.
You appear to have claimed that rights over material is broadly relinquished if it's published in public:
> If information is your competitive advantage maybe you shouldn't have it on a publicly accessible website
And your distinction was further clarified when you argued that placing barriers to access fundamentally changes the equation:
> Note, a simple sign up being required to view a website makes it not publicly available information any longer and you can cover usage, again, in a license.
Perhaps you meant to speak only of material which is not subject to copyright? In which case I think your argument does track.
If you build a database of touristic places and display in your website, the information is not protected by copyright.
In Europe they have laws covering _sui generis database rights_, but they are from another era and unenforceable nowadays.
By my understanding any website with a copyright disclaimer warrants their data as exclusively their own and are granting permission for other web users to generate it, ie people are not entitled to share their web data with anyone. So if they are, and we agree that it’s good that they do, and continue to create information for others to know, how do we avoid the implicit harm in extracting data without nothing being given in return but possibly harming the internet’s experience for everyone accessing the same information?
If you can't make that trade then you've weighed the value provided by an organization like google to be more valuable than the copyright of these content creators and I want other players who may want to be able to challenge google to have the same protections and access google does to have a chance at providing the same value.
Some scraping services make their money by offering scraping services to companies for specific information and you could argue they provide value to other businesses that way, but not to the broader "rest of humanity".
So I'm not sure it's as simple as just "aggregator" good "scraping service" bad as value provided takes on many different forms, and that's what makes this difficult.
I guess it may come down to your take on what you think of middlemen, because they are all effectively middlemen in the data economy.
Edit: I was rereading your comment, in respect directly to the value added to the content, then yes maybe it is more clear that aggregators are in principle different because they do add that value where scraping services that sell the data do not offer any enrichment to the content creator. I personally think protecting content aggregators that republish the data to create visibility or other value for the content creator to the extent that they're not worried about being sued for that is probably a worthwhile thing to happen because of the net benefit to our ability to find information/content.
Like displaying a table with semantic elements, then divs, then using an iframe with css grid and floating values over the top.
This almost seems like a problem for AI to solve.
Still, maybe AI comes into it. Maybe poisoning the data is the right way to do it conditioned on ML-juiced anomaly detection.
Plus, it's one you're going to lose. I was once asked at an All-Hands why we don't defend ourselves against bots even more vigorously.
My answer was: "Because I don't know how to build a publically available website that I could not scrape myself if I really wanted to."
Is that legal? It would be a big blow to trust if I was the customer, but that's without knowing what you were selling and in what market.
I feel like I'm getting a glimpse into the dark underbelly of the web.
We use it for data-entry on a government website. A human would average around 10 minutes of clicking and typing, where the bot takes maybe 10 seconds. Last year we did 12000 entries. Good bot.
I believe the future will make us more free by using more bot / AI technology since who wants to spend their whole day in front of a computer and research information if machines can do the job just fine?
In the past we've had the most success defeating bots by just finding stupid tricks to use against them. Identify the traffic, identify anything that is correlated with the botnet traffic, and throw a monkey wrench at it. They're only using one User Agent? Return fake results. 90% of the botnet traffic is coming from one network source (country/region/etc)? Cause "random" network delays and timeouts. They still won't quit? During attacks, redirect to captchas for specific pages. During active attacks this is enough to take them out for days to weeks while they figure it out and work around it.
With only 10 dongles and 10 dataplans, you can have a lot of IP addresses that are extremely hard to block. It's an one time investment, paying proxy providers is a fixed cost.
We tried to get some, but all of the ones we could get were various levels of broken or unsupported.
>>Because I could not fully trust the other customers with whom I shared the proxy bandwidth. What if I share proxy servers with criminals that do more malicious stuff than the somewhat innocent SERP scraping?
For the site I work for, about 20-30% of our monthly hosting costs go towards servicing bot/scraping traffic. We’ve generally priced this into the cost of doing business, as we’ve prioritised making our site as freely accessible as possible.
But after this week, where some amateur did real damage to us with a ham-fisted attempt to scrape too much too quickly, we’re forced to degrade the experience for ALL users by introducing captchas and other techniques we’d really rather not.
I had a particularly bad time not so long ago, when a customer's site - a shop - was brought to its knees because someone, probably a competitor, hired some scraper-company of some sort to scrape every product and price.
The scraper would systematically go through every single product page.
And by scraper, I mean - 100's of them. All at the same time, using the old trick of 1 scraper requesting 3 or 4 product pages at a time then pausing for a while.
They used umpteen different IP address blocks from all over the globe - but mainly using OVH vps IP address blocks from France.
Now, maybe if they'd just thrown, say, 5 or 10 of the scraper "units" at the site, no one would have noticed in amongst Googlebot (which they wanted to use anyway because they are using Google Shopping to try to bring in more sales).
But no. This shower of arseholes threw 100's of scraper "tasks" at the site. They got greedy.
Now, the site was robust enough to handle this load - barely - which was massive, however, having to do that /and/ also handle normal day-to-day traffic? Nah. The bastards got greedy and like you I spent a few days unfucking the damage they were causing.
Seriously, I hate scrapers. I hate the people who make scrapers. I hate their lack of ethics. Fuck those guys.
Not everybody in this space is out to destroy your site. Some of us actively try to put as little load on your site as possible. My scraper puts less load on sites than I do when I browse them normally, I've measured it. Really sucks when we get lumped together with the other abusers and blocked.
Yeah.
First scraper I ever built was for my school portal. Absolutely atrocious user interface. It got to the point that I seriously hated that site so I built a script to log into it and download my information. I just wanted to see my grades without suffering.
At their request, we built a method to flag accounts for data poisoning. Once flagged, those accounts would start getting plausible-ish looking garbage data.
It was pretty effective. One competitor went offline for a few days about a week after that started, and had a more limited offering when they came back up.
I'm now making an entirely new shop for them - I shall bear this in mind. Thanks for that!
My favorite is Varnish,[0] which I have used with great success for _many_ web sites throughout the years. Even a web site that 10+ millions of requests per day ran from a single web server for a long time a decade-ish ago.
Wait till you find out what half of Google's business is based on (spoiler - scraping).
I really don't think scraping itself is an issue 90% of the time. It's the behavior of the out of control scrapers that are the problem. A well behaved scraper should barely be noticeable, if at all.
I might argue that what google actually uses their scraped data for is their search engine - which is private. They simply allow us access to specially crafted queries, which they can and do manipulate (for many reasons, some good some bad).
The only thing I'd say meets that definition would be like Common Crawl.
if a scraper is effectively DDoSing you, call it what it is -- a denial of service attack.
i've found from experience that most scraping attempts originate against host-sites that are generally user-hostile; no APIs to use, JS tricks to bother user browsing, or groups that profit from first-mover advantage and thus try to obscure data.
So, if your sites are commonly the victim of scrapers that are harvesting publicly available data i've found that it's more useful to ask myself what alternatives I could provide those that feel the need to scrape.
As for a 'lack of ethics' on how publicly available data is wrangled -- well, i'll just say that I feel that it remains the responsibility of the administrator rather than being something to push the blame onto clients for. There are plenty of technical avenues to pursue before appealing to morals and ethics for help.
At the same time, even as someone who runs a web crawler, I have zero qualms about blocking misbehaving bots.
Stuff like redirect resolution is very easy to overlook. You may think you're fetching 1 URL per second, but if you are using the wrong tool and you're on a server that has you bouncing around like in a pinball machine and takes you through a dozen redirects for every request, the reality may be closer to 10 requests per second.
On top of that, sometimes the same server has multiple domains. Sometimes the same IP-address serves a large number of servers (maybe it's a CDN).
Usually it's error pages that really drive the large redirect chains. They often have a vibe of like some forgotten stopgap put in place to help with some migration to a version of the site that is no longer in existence.
Of course you don't know it's an error page until you reach the end of the redirect chain.
Not saying that doesn't suck - it does, it's why many ideas don't work in practice as an online service.
We chose to switch to the JS challenge screen as it requires no human interaction. We now block 75% (estimated to the best of our knowledge) of bot traffic but some customers are livid over the challenge screen.
The irony for us as a provider is that it's one of our customers (party A) paying a third party to scrape data from another one of our customers (party B) which in turn affects the performance of party A's site. We've started blocking these third parties and directing them to paid APIs that we offer.
Also, not all types of company will provide API endpoints. It all depends on the type of site - for example, an online shop might not wish to provide easily accessible data on offered products and prices, to their competitors who may wish to undercut them. Why would an online shop do that?
I absolutely would pay for an API that provides that data. I'd be willing to pay 10x more than the cost of maintaining and running the scrapers.
But the sites being scraped have no interest in that.
Because right now, I sure wish that the bots - which comprise probably 2/3 of my traffic - are causing me huge headaches and I wish that the people doing it would tell me what the heck they want.
The scraping company WILL use the API/CSV file... they will probably also still charge their customer for scraping, so it's a win-win :D
You can think of it this way, the prices and product data are publicly visible already on the website, there are no real secrets, none of it is password protected.
You can be principled and insist on blocking bots and spend a lot of time and money on tools, people, and ultimately hosting because the bots will always win; or you can offer the data for free/minimal fee and serve it with almost zero cost and cache it so you can do that with a micro sized server.
You can always lie about some of the prices if you want, but you will just encourage bots again.
Ethics are nice, but let's be honest, very lacking. Sometimes it's better to be pragmatic.
There's the problem right there. The prices and product data are publicy visible - because there is a target audience of /humans/ for whom the site is designed and intended to be used by. The site is not there to cater for a competitor's scrapers.
I don't care how much people couch their unethical behaviour in "the data is publically available", the basic fact is most if not all websites exist for human eyeballs to look at them. They do not exist for arseholes to DOS them by inundating them with scrapers.
But overall, information is one of those goods that has intrinsic properties like no other. It can be copied, infinitely. And we haven't yet figured out the dynamics of how to reason about it, so it feels like we're pretending they're physical goods.
Edit. Side note. I'd go further and say that some of the data is even worse, it's "offered" with the real intention being to confuse the users into performing non-optimally in the market. Look at Amazon/Ebay/AliExpress/Google listings for evidence of that. Just Google - Google is a ML and scraping power house, and the best they can muster is to be spammed with fake websites and duplicate/confusing listings.
And that's also the dirty secret behind the "attention economy": it's whole point is to make things as inefficient as possible, because if you're making money on people's attention, you need to first steal it (by distracting them from what they're trying to achieve), and then either direct towards your goals (vs. those of the users), or stretch it out to maximize their exposure to advertising.
--
[0] - Sometimes unintentionally. Unfortunately, the overall zeitgeist of UX design is heavily influenced by bad players, so default advice in the industry is often already intrinsically user-hostile.
This is exactly right.
There's a whole ethical subthread here of websites trying to making the experience for those humans miserable, and taking away the agency necessary to protect oneself from that. A browser is a user agent. So is a screen reader. So is a script one writes to not deal with bullshit fluff, when all one wants is a simple table of products, features and prices.
Your argument is perfectly valid and applies to offline activities as well (what stops a competitor from walking through the aisles of a Walmart or Costco?), but this is a battle that can't be won, there are too many parasitic actors. It is human nature.
That's a significant portion of Nielsen's business model.
Because otherwise the HTML will become the API.
Valuable public data is going to be scraped - this is inevitable. Even paywalled or signup protected valuable data is going to be scraped.
Why not sell valuable data for reasonable price then.
If an amateur can do damage to you, then I have some bad news for you...
I believe the point wasn't surprise that damage occurred at all, but frustration that damage can occur just out laziness/ignorance rather than malice.
I've witnessed a site being basically DOS'ed due to particularly greedy and aggressive mass scraping attempts.
It’s just selfish. If you’re going to take the product of other people’s work in a manner they don’t consent to, at least do it in a way that doesn’t cost them twice over.
I asked them if my customer could pay to access this data point via their API and they quoted 3600 EUR/month! Enter the scraper...
Does your API provide all the information that can be found on the site, or are they scraping because the API is incomplete?
We've once had to scrape Amazon product pages because they have a lot of API endpoints, but those didn't contain the data we needed.
In what universe is providing such a straightforward way of helping a competitor considered sane business practice?
It is a sane business approach when you are a pragmatic business who knows the limits that constrain your business.
Either the content company is going to build a simple API (could be just a static CSV file hosted on S3 or whatever) with useful information or try to monetize/hide this information and force scapers to use the website .
A bot is always going to win unless you want to make users also a lot of friction. In the era of deepfakes and fairly robust AI tooling the difference between bot action and humann action is not all that much.
If you are going to be agressive with captcha , IP blocks and other fingerprinting, users who get identified false positive.or annpyed would leave.
When the cost of losing those users is more than allowing access to scrapers,you would absolutely setup the API.
> We've once had to scrape Amazon product pages because they have a lot of API endpoints, but those didn't contain the data we needed.
...only a couple of comments up.
I’m asking this because I’m going through very similar situation and would love to see other opinions around this.
Just curious about the difference in value from using your API and web scraping as there is a cost to web scraping as well.
If you make your api client well, you don't have the problems of a scraper - but if the api owner decides to change rules for api and you can't do what your business is based on being able to do (think of api owner as Twitter) then you need to make a scraper.
Can't speak for the op but we have APIs and move the ones scraping and reselling our content to APIs. The majority are just a worthless suck on resources though.
Then graduated to JavaScript for surrounding logic e.g. data transformation
I had assumed I'd quickly give up and move to a headless browser, BUT I can't bring myself to move away from tiny CPU utilization of curl.
Throwing together a "plugin" probably takes me less than 20 minutes normally.
I'll probably have a look at using prowl to ping my phone.
And if I get more serious I'll look at auto authenticate options on npm. But I'm not sure if the overhead of maintaining a bunch of spoofy requests will be worth it.
Taking a stab at answering it: you scrape the data and build a business around selling it. Stock prices? But that's boring, plus how many others are doing it? I bet a lot.
Artificial scarcity - every week you release a "limited edition item", but if you do the math, it's not limited edition at all if you integrate over a year.
Prices (are yours high, low compared to competition?), reviews, locations of physical stores, search result placement (where does your widget show up when someone searches "widget" on your site?), just to name a few use cases.
I used to scrape websites to generate content for higher SERPs.
Ended up going into the adult industry lols. (https://javfilms.net)
I've always wondered, and since you're right here... how do sites like this make money?
It looks like you're probably crawling all the JAV vendors, finding free clips of today's releases, embedding them in your own site to draw traffic, and making money with affiliate links to buy the full content?
Am I missing anything? It seems hard to believe you'd get enough affiliate signups to make it worthwhile.
I can imagine your site as being a few hours a year of script maintenance and a money printer, or a 40hr/week SEO job with 1000s of similar sites across the adult industry.
I'd love to know anything you're willing to share about how the business works.
It should be a matter of a simple GET request to fetch plain html and parse the OpenGraph meta tags out if that. There are many open source libraries to do that for you depending on your language.
If bot blocks really are a problem, a SaaS solution like Microlink could probably do it for you.
Microlink is a good tip, thanks!
https://github.com/NikolaiT/Crawling-Infrastructure
And here I am writing about it (but its quite old): https://incolumitas.com/2019/08/31/web-scraping-puppeteer-aw...
(If I were writing something to be published, though, I would write "search-engine results page" instead of "SERP".)
SERP: Search Engine Results Page
That said, I do a lot of SEO work.
Still, it should be best practice to define any acronym or initialism the first time you use it