We do have human emulation routines that helped avoid most detection, and that library is decoupled in such a way that we can edit behavior down to the individual site.
Some sites are just so damn good and detecting us and I just don't get it.
We do have human emulation routines that helped avoid most detection, and that library is decoupled in such a way that we can edit behavior down to the individual site.
Some sites are just so damn good and detecting us and I just don't get it.
The countermeasure would be to have a bunch of humans use the websites in any way they want, totally undirected, then use the totality of that browsing to facilitate your scraping probabilistically. It would be less efficient, but very difficult to catch.
Given that your run this division there is a good chance you are personally liable.
We've had this division for many many years, and before my time we paid another company to do this. There's no legal issues.
Computer security laws are very broad. It doesn't matter if it's just a website that the public can access. If you're accessing it in a matter that they don't want AND you're aware of that, then I struggle to see how your lawyers can justify it.
> Computer hacking is broadly defined as intentionally accesses a computer without authorization or exceeds authorized access.
https://definitions.uslegal.com/c/computer-hacking/
Hiding your user agent because you know they don't want automated retrieval of information is "without authorisation".
Don't think connecting a computer to a private network to suck up subscriber data is comparable to scraping publicly accessible internet content.
First, the accuser needs to, at least, send a cease and desist letter to the accused asking them to stop accessing the protected computer. Second, the accused needs to ignore that request and keep accessing the protected computer.
Is it possible to build a solid CFAA case when those two things do not happen? I cannot find any examples.
https://iapp.org/news/a/can-a-cease-and-desist-notice-create...