I agree with another comment that called this "Abuse as a Service". It seems to me this product's design goal is nothing more than to circumvent measures site owners take to prevent abuse of their site and run a sustainable business.
I agree with another comment that called this "Abuse as a Service". It seems to me this product's design goal is nothing more than to circumvent measures site owners take to prevent abuse of their site and run a sustainable business.
How will well-behaved scrapers undermine the sustainability of a business? I guess adblocking is one, but we can already do that with uBlock and that's legal. Or adversarial bridging, but that only serves to boost competition.
In other words, the question is flipped; why would well-behaved (i.e. non-DDoSing) scrapers be illegal?
The question was literally,
> What are the legitimate (i.e. legal) use cases for a product such as this?
Nitter is an example of a service that explicitly disrupts Twitter/X's way to make money. If they can't make money then they can't provide the service, there would be no Twitter/X, and hence no Nitter. Of course they would try to prevent that kind of behavior and it should be obvious why. Resorting to using a service like this in order to continue using Nitter should raise some alarm bells. Sure you can still do it and rationalize it however you want, but you have to acknowledge you're trying to get the value of the service without paying for it.
Perhaps there are cases where there is a dissonance between a website's TOS and how they are blocking bot traffic? That sounds like a valid gripe. Otherwise, I don't buy the argument.
• I want to automate (or at least semi-automate) downloading bank statements. I've got ~14 accounts (checking, savings, credit card, IRA, investment, HSA) across 7 financial institutions.
It's tedious to go download statements from all of them manually.
• I want to save stories from FanFiction.net (FFN) for offline reading. FFN's terms allow automation as long as it doesn't operate faster than a human [1].
[1] From their TOS:
> You agree not to use or launch any automated system, including without limitation, "robots," "spiders," or "offline readers," that accesses the Service in a manner that sends more request messages to the FanFiction.Net servers in a given period of time than a human can reasonably produce in the same period by using a conventional on-line web browser.
Could you not shoot an email to those institutions asking for a copy of the documents?
They’ll respond within a few days, asking me to log into some web portal to prove that I am me, and then we’re back where we started
If there was an accessible API to do what I need, I wouldn't do this because scraping sucks. I have to write 100 JavaScript edge cases to handle all the times the host's servers fail in very weird ways. Plus, walking DOMs on these shitty sites with 10,000 nested divs is not fun. GPT helps with this.
It's net-positive for the host though, as I upload a lot of valuable content that their users genuinely like, but it sucks that I have to be sneaky to get the data I need.
I've used their product many times actually, and I'm shocked on Hacker News of all places no one's thinking of anything besides abuse. How often is it useful to get information from a webpage and apply it in a new context? Then think of how often said webpage is behind a Cloudflare bot detector.
They are completely in the right to block you though, you're not the owner of that data, you might be breaking their TOS.
In Europe, if the company is actually following the law, in theory yes.
> They are completely in the right to block you though, you're not the owner of that data, you might be breaking their TOS.
IANAL, but AIUI that's definitely not true in the United States and I suspect similar ideas hold elsewhere: https://en.wikipedia.org/wiki/HiQ_Labs_v._LinkedIn
- GDPR doesn't require it be a convenient export. Users want to paste a link on my site, a click a button, and have it magically appear. Not fill out a form, dump their entire account and sift through that.
- I never opined on the validity of blocking bots
- I never opined on if it's breaking their TOS
Abuse implies a harmfulness. Giving users a quick import option from already public data isn't harmful.
They’re not necessarily in the right to block you, if you’re the data subject or acting on their behalf.
Like if I want to programmatically unsubscribe from a subscription, why should I have to do it myself?
(and for that 1% of the cases where the address is not a spammer and user knows it, they can just hit "unsubscribe" manually)
Scraping is important for example, to monitor competitors' prices to see the opportunity to raise your own prices.
And let's not forget that Google does a lot more scraping than anyone else and has ridiculous profits from it.
For years I've used my own terminal UI player (di-tui) for di.fm. At some point in the not-so-recent past, di.fm added Cloudflare's WAF, which prevents me from using one of my app's features: managing channel favorites within the app.
To be clear, I'm a paying di.fm customer, and my app only works for paying customers. But now my preferred method of listening to di.fm is slightly hamstrung because Cloudflare's WAF sits between me and little string token available to every browser that accesses di.fm (even non-paying customers).
Ps: context why I need automation for such thing: those lessons are really popular and are announced at unpredictable time / there might be another spot when someone resigns
Data portability! Tools like this can be used to allow individuals to export their data from hostile web services trying to hold it hostage.
Legal in the EU, with GDPR.