Extraction => https://www.zenrows.com/blog/mastering-web-scraping-in-pytho...
Avoid blocking => https://www.zenrows.com/blog/stealth-web-scraping-in-python-...
e.g. good luck trying to get much out of youtube.com (or any other video site) without executing JS.
Imagine an async-loaded list, that continues loading more content as it comes in, until it displays all of the content available to the backend.
When would you know such a list is finished loading?
This sounds insane, but it's pretty easy and common for an ambitious UXer to key in on, and is something I've seen in production pages.
(In the event you are a UXer, please include some sort of status update! Even an overlaid spinner that disappears solves the problem.)
Usually the response is in JSON and you can ignore the original page. You might have to auth/grab session cookies first, but thats still easier than working with the HTML.
- session persistence
- dealing with cdns
- dealing with regional proxies
- dealing with captchas
- dealing with websocket data
- dealing with custom session handshake sequences
list goes on and on and on, but probably just edge cases haha