How we learnt to stop worrying and love web scraping
nature.com
nature.com
Web sites change, web frameworks evolve, and just some subtle reordering of some <divs> or renaming of CSS classes, and your perfect scraping code from yesterday will break tomorrow -- maybe not leaving you empty-handed, but probably missing some data or delivering the wrong one.
If there is an API you can use, use it. If your budget allows to pay for API access, buy it. APIs tend to be more stable than scraping, and the data provider will probably inform you if it changes. Contacting them might even get you more interesting data, as not every column they have in their database might become published on the web site.
Conversely, overall document structure doesn't change much over time. I know it _can_; there's a social contract that APIs should change slowly while documents can change whenever, but that isn't what I observe in the wild. Even on fairly major redesigns, the overall structure has minimal edits.
A technique I've used before (wasted effort in hindsight since web pages are stable and I never have to update my scrapers) is to come up with several semantically different ways of accessing a piece of data on a page. It serves two purposes; you can recover from small page changes by having the different methods vote, and you can detect most kinds of page changes by noticing discrepancies, notifying yourself that the scraper needs to be updated soon.
Granted, but there are lots and lots of ways they can break scrapers in the pursuit of their core business, such as a website redesign. For example, moving from static HTML to a web framework would require your scraper to actually run the JavaScript to generate the DOM in the state that a reader might view it in, and this is quite a lot more complicated than walking the static HTML.
Or, as is often the case, the content is already there or fetched via an API in far more easily-consumed JSON format that you can use directly.
Granted, lots of APIs make it prohibitively difficult to authenticate such that it’s easier to simply scrape. Such is the case with just about every Microsoft product I’ve ever used, most recently the XBox Live API. I genuinely wonder what kind of nonsense goes on in Microsoft design review meetings.
Looking at this sentence, I have the impression that it is nowadays taken for granted that "web framework" means "front end web framework". I come from a time in which it was perfectly fine to generate static HTML via a (server-side) web framework.
Certainly more resource-intensive.
Those that don't use their own APIs almost always end up with an open API in the state you describe (except maybe the very big players like FB, where the open API is overall good).
The best explanation I've come up with is that as a naive developer, it's impossible to know the nuances of any sufficiently complex process or workflow.
All of our own websites are built on-top of the same public API that everybody else uses and scraping used to be a nuisance. It was also confusing because they would be able to get more data using the same free account just by using the API instead of scraping. Exactly like the OP mentioned we only show a small number of properties via the website but most scrapers never took the time to actually compare API vs website.
And the only thing we really wanted was article title and date.
People are usually hesitant to constantly change how a customer interacts. All to willing to change internal details.
We had to build a test suite just to verify that the API working as expected, because they would break it so often, and also the documentation didn't always match reality.
This was a paid API we used on a pay-per-use basis, IIRC, and had official support for.
In the beginning, we had a false sense of security about the version numbers and such. The first couple of breaks seemed to be "just this one time". Then we realized that it was happening all the time, and so the test suite. (I was a junior, so I can't take credit for this work, just a witness.)
API is often no more reliable than scraping human UI most of the time, with the added disadvantage of being second-level importance.
Personally, I've tried to combine human UI with API as much as possible. For example, I added a feature for being able to post via direct URL entry, like so: http://example.com/hello+world
most browsers will convert the spaces for you too, so you can just type your text into the address bar.
Understanding "web APIs", which did not exist when I first starting using the www in 1993, other than as a way to try to control and/or monetise scraping continues to escape me. I do like the increased usage of "endpoints" though, serving only data with no markup. Although XML and JSON are too bloated compared to something sensible like netstrings.
Similarly, on the client side, I fail to understand all the parsing tools and libraries and related promotion; it is just as easy to break any solution that depends on them and in many cases they are obviously overkill, more brittle than simple scripts using generalised text-editing tools.
One example is "jq". In many cases it is clearly overkill and is slower than sed.
https://stackoverflow.com/q/59806699
As a data source, the web is messy. "Standards" cannot be relied on 100%. Some people try to pretend the web is clean and can be tamed, or they "give up" because it is not "perfect" and things can break. Getting hands dirty works the best and most things do not break if kept simple, IME
That's just my own take. I've worked in environments with stronger versioning, and depending on your data needs and structures it can work. It's usually not worth it for most use cases though.
My best idea has been to simply maintain a collection of "reference" URL's (e.g. of different products or articles) and identify unique start/end text for those specific instances.
Then automatically extract as many possible different "rules" for locating the desired content (pure structure and ordering, class hierarchies, classes/ids, surrounding text, etc.) and find the ones that are consistent across different instances.
And then just use those rules until they break on the reference page... and when they break, develop new ones.
I'm curious if anyone's built this type of thing?
I'm having trouble finding those papers at the moment, but here are a couple commercial products that sound similar in spirit to what you're describing.
https://www.diffbot.com/ (kind of)
Edit: I hadn't searched recently enough. See the sibling comment recommending this library. Haven't used it yet, but at first glance it looks nice. https://github.com/alirezamika/autoscraper/
(1) Its wrapper generation code isn't much more advanced than that similar data will be similarly nested in similar parent blocks. It looks more brittle than I'd like.
(2) It has zero tests, comments, docstrings, types, or any other niceties so far (and minimal documentation).
(3) When things go wrong it strongly prefers returning no information and not throwing any errors. None of the examples in the README actually run (or rather, they give you a `None` response that's all but useless) without changes.
It works well enough. I tried a few other sites and had mixed results even when providing the raw HTML so that I knew its http logic wasn't the issue.
> maybe you didn't update the wanted list in the examples
Yeah, that was my only real problem with the readme examples. Those could just as easily be provided as local data (e.g. how `sklearn.datasets` works) so that the end user starts with working code, especially since there are no errors/warnings/etc when anything goes wrong.
> IMO the biggest problem is the lack of js enabled content support.
Haha, unless I'm seriously misunderstanding you this is one of the only things I don't mind :) Since you can pass raw html to the library, you can use your favorite headless browser to navigate (or in the happy case just load a non-interactive js-enabled site) to your content and pass it through to this library to do the data extraction. I rather like those features being decoupled and kind of wish this library didn't attempt to do any of the crawling itself. I know that's just a personal preference, but it's my account, so I'll say what I like about it.
It depends on the SLA, of course, but it's cheaper to check every few hours than on every request, and you get a couple of alerts instead of a constant stream of them.
I guess a lot of stuff I've needed to scrape are from old CMSes or sites where its viewed as part of a cost center they're unlikely to invest in and just maintain it
After 'scraping' some forums via their APIs for weeks, I ended up realising that the data and metadata given to me by the API was so restricted (e.g. providing 'recent comments' instead of all comments) that a pure vanilla web scraping approach became the preferred option.
I agree with all the points you mention about the shortcomings though and your argument is sound. This is my opinion in the other direction, APIs come with an element of trust.
We had a network of sites, and they all looked the exact same to the user, but to google, each site had a completely different structure, and it kept the network safe for years before a google employee (or we assumed google employee - @google.com email) signed up for the service without us knowing, and discovered the entire network by placing a large order which gave him links across the entire network. Within 1 week of them signing up, our entire network of 10k domains was dead, and everything they linked to was delisted from google. We had to shut down the network, and refund all unused credits from our customers.
We tried to just fly under their radar, and avoid any automated trigger that would arouse their suspicion.
All this seems like a nice illustration of how the web ecosystem encourages parasitic behavior on so many fronts. It's sad.
We were masking those domains from google because google penalizes selling backlinks to justify paying for their ads. My conscience is quite clear. When google delisted our network, we refunded our customers and moved on to a smaller invite only network that ran well for years. I left that company 6 or 7 years ago, but I'm sure they are still making some money off hosting and managing private blog networks.
The average American endorses slavery in their clothes, Christmas decorations and electronics. By comparison creating some bad links in a search engine is so low on my list of moral failings it doesn't even register.
Underperform, and then offer them a goodwill refund if they ask for it?
Sure, if you change the page structure enough you could defeat it, but it would require more than just adding a few divs. XPath easily lets you mix and match matching against not just CSS classes, but also the page's structure itself, inner text, attributes, and so on. As a result, you can get some really powerful queries without having any kind of complex post-processing of the results.
More complex scrape defeating measures I've seen are blobs of JS that need evaluating in order to generate URL parameters (all that needs doing is extract the JS and run it in a JS engine, if you don't want to drive a headless browser, with care of course!) or that need a captcha defeating (just buy some deathbycaptcha API calls).
Contacting the agency for the data is always a good step, but even if they are responsive, they may not be willing to email you on a daily basis with data updates.
But in the article they mentioned re-running their program to update their data, so it could be a long-term effort. And anyone planning a long-term project reading that article and taking their advice should at least be warned about this possible problem.
We ended up writing a scraper for the main unit. Once they have been installed and configured, they're likely not going to be updated unless absolutely necessary (such as when adding new, unsupported controllers to the bus).
The good thing with ciscos and most of the old technology is, that once you write that script, it works for years... commands never change, outputs never change, some perl, a regex or five, and you're done.
Doing the same with a webpage, where it's "pride month" today, "womens day" tomorrow, "day against aids" the day after, and each means a new div, a new popup, a new redirect, a backend update inbetween, to make the new banner possible, etc., is a pain in the ass.
I've found fully specified XPaths to be a mistake for example. It only takes one tiny change on the page to mess up the script. On the other hand, despite numerous warnings that it would be a disaster I've found I have a lot of luck maintaining regexes, even after major page reworks.
I know they mention maintenance, but having been in both fields what a web person means by maintenance and what a research person means by maintenance are an order of magnitude different.
Can you give any specific examples (sites and data needed from them).
Every few weeks or months, one of the comics would just not appear or not be updated (i.e. there was the same strip day after day). I had to update the code each time, checking the source code and finding the new page and location of the image. Later, some pages started using JavaScript to load the image file, and I lost interest in sophisticating my script. So, no single page comic collection for me anymore. :)
Set up a cron job to run `dosage @` every day, it will check for new comics and download them.
Though, you you get cases like celebritynetworth.com (no affiliation to me) that was featured on HN where Google wanted to APIize the data and when that was not available they decided to scrape it. It effectively killed the business.
I think if a company decides to offer an API or other structured data then fine, they've decided that their data can be shared in a way that makes it OK to collect en masse.
I wouldn't say we should "learn to love" scraping when there are sites that put a lot of man hours into the data that's seen on public html pages only for someone else to spend a few hours scraping it and repurpose it.
Now, if the major traffic drivers like search engines and social networks were able to differentiate who the true authoritative source is, that'd be great. But they don't.
The article says, "You might be able to use what you scrape, but it’s worth checking that you can also legally share it. Ideally, the website content licence [sic] will be readily available."
In practice, the data you want to scrape either has 1) no license/information about whether they are okay with you scraping it (and if they are okay with it, they usually offer the source data anyways), or more commonly they 2) strictly prohibit it in the terms of service, in which case it's not clear whether a scraper which mimics a browser falls under fair use.
It's an area where when you ask your university attorney, they exchange a few emails with you and then avoid making any decision. I think it's just because it's not well tested in court (at least when I encountered these situations), and depends on whether you'd be subject to a takedown notice or lawsuit.
Or what about a 'human scraper' a la Amazon MTurk?
Very strange fact.
Too bad I can't find a measure of English content written by country. https://en.wikipedia.org/wiki/List_of_countries_by_English-s...
Not part of the Berne convention, and not applicable in the US.
Here's a short example that scrapes HN Favorites. [0]
#!/usr/bin/env python3
import requests
from bs4 import BeautifulSoup
username = input('username: ')
# session uses connection pooling, often resulting in faster execution.
session = requests.Session()
base = 'https://news.ycombinator.com/'
path = f'favorites?id={username}'
while path:
r = session.get(base + path)
s = BeautifulSoup(r.text, 'html.parser')
for a in s.select('a.storylink'):
print(a.text, a['href'])
more = s.select_one('a.morelink')
path = more['href'] if more else None
[0] Python: https://github.com/gabrielsroka/gabrielsroka.github.io/blob/...[1] JavaScript, runs in browser with a UI: https://github.com/gabrielsroka/gabrielsroka.github.io/blob/...
Personally, I loved using BS for hobbies until the SPA era started, and then I had to either use headless (selenium back then was great) and/or monitor the network for their API.
[1] https://developers.google.com/web/updates/2017/04/headless-c...
[2] https://developer.mozilla.org/en-US/docs/Mozilla/Firefox/Hea...
And with Puppeteer (also Playwright) it's never been easier. Recaptcha solving, Ad blocking etc. in just a few lines of code[1].
I've built a business on the back of Puppeteer - https://simplescraper.io. Ten months in and we've just passed 100 customers so there's mucho opportunities in solving these kind of problems.
[1] https://github.com/berstend/puppeteer-extra/tree/master/pack...
name = Text(to_right_of='Name:', below=Image(alt='Profile picture')).value
Thank you!I totally agree with the general idea of sharing code and tools, in an open research community but incentivising someone from a field completely unrelated to CS to 'just code' sounds like pretty bad advice. This is very likely to lead to a lot of time spent, frustration, and underwhelming results.
Why not incentivise building a professional Software team at inter-departmental University level and have that team act as multi-department shared resource?
edit: missed words
First, almost all academic code is really simple from a software engineering perspective, but really complex from a subject matter perspective. Having a deep understanding of both the data and the relevant hypotheses is critical, and is often actually helped by writing the code yourself. Trying to communicate every feature requirement perfectly, and making sure every assumption is met, to a third party CS person might be possible but is definitely non-trivial.
Second, most of these projects are one-time use (code up some project over a few months, write a paper, never touch again), and so spending a ton of time + money making it robust and efficient is not really worth it. For things like open source tools that are expected to be used by a lot of people it's much more feasible to get engineers involved. The Chan Zuckerberg initiative is actually funding a program that essentially does this [1].
[1]https://chanzuckerberg.com/rfa/essential-open-source-softwar...
I’m sorry but learning to code very custom or short lived software is just not that hard. I think you’re making programming too precious.
Don’t get me wrong — writing good, maintainable, long lived software, and basically anything with a UI for non programmers, and certainly anything for sale, is much harder and usually something professionals or their equivalents will need to do especially for an institution.
But writing scripts in interpreted languages like python or ruby for the research purposes of an individual or small group is not something people need to (or more importantly can) hire a dedicated team of engineers to do.
IMO it’s good that programming has become more democratic. It’s an ongoing process. And one that has happened in basically all other creative industries to some extent.
[1] https://society-rse.org/ [2] https://www.software.ac.uk/
I'm keeping the codebase open on GitHub: https://github.com/umaar/learn-browser-testing/ so anyone who wants to follow along can do so for free.
In the 2-scraping folder, there's a bunch of scraping examples such as bypassing captchas, building an Amazon price checker, blocking CSS and JavaScript resources during scraping to make the process much quicker, and a few others. Hope it's useful.
Nowadays mostly I write custom scrapers use PHP (Guzzle, Curl, etc) and Python (I was introduced to Beautiful Soup, by the "Python for Secret Agents" book)
I've tried several commercial scraper tools and services over the years but few stuck for various reasons.
UbotStudio had great potential but in the end was buggy and painful, almost abandonware.
Scrapinghub.com is decent enough but bit expensive for my projects.
80legs.com is cool for massive scale but was overly robots.txt restrictive at the time I tried using it (for what I was scraping) and I don't like the syntax.
A scraper colleague likes using Winautomation however it's no longer for sale separately, because Microsoft acquired the company and rolled it into their Power Automate SaaS (RPA/RDA focused)
There is a new tool called RTILA that I used for the very first time a 9 days ago, which actually is the easiest way to create and run scrapers I have found in 20 years.
The RTILA software currently has minimal documentation (apparently is being worked on now), however the new features are being developed fast, all releases are here on GitHub (see the frequency of releases) https://github.com/IKAJIAN/rtila-releases/releases
Another user of the software has produced several video tutorials here showing how it works: https://www.youtube.com/channel/UCH6ov8LnB8-4ZF0yraxjw8Q/vid...
You download it from GitHub but also still need to buy a license key from here https://codecanyon.net/user/ikajian
The RTILA home page is here https://rtila.ikajian.com/
I am a genuine customer and have no connection in any way with the solo founder developing it, (except for a few emails and support forum messages). Genuine recommendation for writing a bot more easily.
Each website respond differently, some can just use requests, others would need selenium with pre-existing profile. To develop a common utility is almost futile.
I do have useful building blocks, but for each individual things I want to scrape I scale out using project specific code. It's never too slow either - the time it would take to fill in all of the required bits in a do it all tool would have been similar.
I recently sold a 'business' that does web scraping. Does anyone have insight into which industries need more web scraping 'experts'?
They checked for bots through useragent/screen size, maybe mouse movements, trends in searches (same area code), etc... (Can they really detect me through my internet connection headers, despite proxies?)
It was impossible for me to scrape, they won.
there are 2 approaches they use that make developing bots very difficult.
1. they detect device input. if there is no mouse movement, while the website is being loaded, they will consider it's a bot.
2. they detect the order of page visiting. A human visitor will not enumerate all paths, instead, they follow certain patterns. This is detectable with their machine learning model.
I really don't have a solution for #2
If you record, you can probably teach AI to emulate.
Think about websites that have every reason to stop you from scraping.
It's not reasonable to disect their huge obfuscated js code. So headless doesn't really work.
In the end, not hammering a site is key... also, on your own end, possibly hashing the main content, so you aren't creating duplicate entries on your own side are important.
As to fragility, that happens... in general, you need to update to match site updates, but most sites won't be dramatically updated more than a couple times a year if they aren't in active development.
Python Tutorial: Web Scraping with BeautifulSoup and Requests
I turn these requests down because it’s a ticking time bomb when you add in the element of login (password resets, 2FA, another point of change). On the other hand, I wonder how different this is from what Plaid does...
Although the caveats for this particular library[1] imply enough false positives and false negatives that it seems mostly useless. Sites that take this seriously must be doing something smarter.
[1] "Doesn't work if DevTools is undocked and will show false positive if you toggle any kind of sidebar."
https://spencerbaucke.com/2020/04/29/web-scraping-in-power-b...
b) what does "switching to WASM" mean for you? If it still generates a DOM or access a private API, many scraping techniques don't care.