The Hard Economics of Selling Web Data
blog.databoutique.com
blog.databoutique.com
For one startup, which addressed problems for brands like counterfeiting and gray market diversion, we needed to monitor online marketplaces over time.
Off-the-shelf scraping products we tried, which supposedly targeted one of the biggest marketplaces, were very poor, even when operated in their manually-invoked modes (nevermind scheduled).
I ended up building a scraper that was correct and reliable, and captured all the data that we needed. And cost around $5/month to run.
(This superior capability was not only useful in general for our services, but I understand it helped with cold enterprise sales approaches, from a tiny startup, e.g., maybe something like: "Here's an analysis of impact of X to your brand, based on data we monitor with some of our technology. Would love to discuss with you.")
Controlling and understanding the software also gave us the flexibility to add capability quickly, as soon we had a need. For example, since the problems we were addressing were global, as soon as we needed to be able to see what different geopolitical instances of that marketplace were showing to shoppers from those same geopolitical areas, I was able to implement scraping of that.
This was remarkably stable/resilient over time, as well. The biggest risk I perceived was something I warned the team about: "The first rule of Scrape Club is: Don't Talk About Scrape Club." I didn't want to get on the radar of the marketplace's lawyers, given that they were profiting from counterfeits and gray market, and possibly have public information turn into something that burns up the remainder of our meager Covid-time runway on lawyer billable hours. (Having data get cut off due to lawyers also seems a risk if using a vendor for the data.)
Of course, there are often good reasons to buy data dumps/feeds. Just giving an example of when it wasn't the best choice.
If a stream of data is central to your business, then you might catch on eventually, but if you're just selling it along, it will most likely take forever to detect this type of quality problem.
Poisoning the data is much more insidious, since it doesn't let the scrapers know that you know. It means you can work to undermine whatever purpose they have. If they try to sell your data, you'll ruin their reputation with their customers. If they use your data for decision making, you trick them into making bad calls.
The biggest risk I perceived was something I warned the team about: "The first rule of Scrape Club is: Don't Talk About Scrape Club."
It seems someone broke this rule: https://www.google.com/search?q=the+web+scraping+clubThat's why people simply hire people, rent out proxies and store the data in their own internal databases, because then the companies have to figure out who the source of the scraping is and be able to prove it, which isn't a lucrative target for the scraping companies.
That being said, there are clear examples of companies that sell processed data that's been analyzed by data analysts and isn't the raw scraped data and they have successful and lucrative business models.
I came here to make substantively the same point. Scraping for your own personal use and re-use as a derived product is reasonably cool but can border into IPR concerns. Scraping to re-sell as a product, thats really stepping over a line.
This has been a problem since the days of the yellow pages phonebook. There are rules around when a catalog is in the public domain, page contents with data I don't think cross into the "its ok to just copy and resell" space.
I scrape several sources. World population and GDP stats are highly useful but finding good canonical sources was hard for a while there. I now use a Python API provided by the world bank, for most of them. Now I just have the problem that by economy, the inputs span 2019-2022 and I can't get a single authority for population estimates right "now" worldwide. But, scraping to resell that? I think it would be wrong.
Offering a brokering service to "share" access to the investment in scraping, has to be done carefully. If for example its brokering access to the page specific structure in some meta notation and you as a consumer scrape direct? thats actually quite cool. If its running a cache of the data re-presented in JSON form for you to consume, or monetising an Elastic Stack of the data and you read it from them? I think its getting iffy to be that seller.
I doubt that is true for all jurisdictions
In the EU there's a 'database right' that covers specific collections of facts. The US has categorically refused to issue any kind of "IP"[0] on facts, and doesn't have 'database rights', but in practice you can sprinkle a few fictitious bits of information in your dataset, retain copyright on the lies, and sue people who don't independently verify the factuality of every bit of information in the dataset. So the US has database rights.
No clue if people are sticking fictional data on their websites to sue scrapers with, but I wouldn't be surprised if they were.
[0] Intellectual property is the right to dictate to competitors the conditions upon which they are allowed to compete, if any.
Short ones, too.
Not a copyright lawyer and I don't know what the actual law is against scraping, but just pointing out that previous companies offering raw scraped data from a large platform were sued out of business.
All big players in the "information" economy found other ways to make money then sell data directly, for example Google's attention market for online ads or payment for order flow in trading.
On the other hand, extracting data from a couple of web pages was both more accurate for them, more reliable and less costly, assuming you can nail down the scraping/storage/filtering part. This latter part is a real challenge when people need data they actually rely on in a business process to make real world decisions, compared to "trends" or "insights" that are rather informative.
Can anyone provide advice on how I should package it and sell it or any advice towards adding extra value?
Do you have a sense for how people store, transform, organise, model the data once they have bought it? Is it typically data warehouse and dbt or straight into a data science pricing algorithm, or something else entirely?