Usually you only need some subset of the data per page load if you invest some time looking at dev tools you can probably find the API call you need and save yourself a few MB.
They all offer scraping APIs now too that can be cheaper for certain use cases where you only need a subset of the data that is actually loaded. Like $1.3 per 1k requests: https://oxylabs.io/products/scraper-api/web/pricing or $1 per 1k: https://brightdata.com/pricing/web-scraper
https://faq.oxylabs.info/en/articles/8826164-restricted-targ...
Additionally I believe that $4/GB is an introductory price. When I went into the Oxylab dashboard, it showed me $8.
I’m curious if you can suggest a happy medium between curl-impersonate on VPS (dirt cheap) vs residential proxy ($8/gb)?
Personally I’m not trying to hit any of those common sites.
Most of them seem pretty reasonable?
"Entertainment & streaming" - who's trying to scrape netflix's library?
"Banking and other financial institutions" / "Government websites" / "Mailing" - seems far more likely it'll be used for credential stuffing than for "scraping".
"Ticketing" - seems far more likely that it'll get used by scalpers than for scraping
The main targets of scraping - e-commerce sites (for price comparisons) and social media networks (for user generated content) are fine to scrape. Is there some use case I'm missing here? Is there a huge contingent of people wanting to scrape ticketmaster or bank of america?
https://www.courthousenews.com/wp-content/uploads/2024/01/me...
"Residential proxies is based on traffic and purchase model. Pay as you go model starts at $7.35 per GB, and can be discounted as low as $1.84 per GB when purchased in bulk."
But I would guess this type of proxies is mostly use to send data rather than receive. You can access geo-locked sites through standard vpn.