Gentoo bugzilla closed due AI bot scraper overload
social.treehouse.systems
social.treehouse.systems
Our largest offenders seems to be mostly limited to South-East Asia, so probably mostly Chinese AI projects, but that's speculation. I also don't recall ever seeing Grok IP ranges or a specific Grok UA, but that doesn't mean that they're hiding, perhaps they're just not interested.
> Our largest offenders seems to be mostly limited to South-East Asia, so probably mostly Chinese AI projects, but that's speculation
Ok. So you also don't know. Well, I don't know either, but I don't make a speculation by claiming x, y, and z companies to be exempt. In my book they are all responsible.
So, it's still possible they are also running a bunch of poorly-coded scraper bots through residential proxies, but it seems a little non-obvious why they would be doing this.
(It is worth pointing out that most of this traffic seems to be dumb: it's stuff like getting lost in generated link forests of some web apps or repeatly re-querying the same endpoint on a super-high frequency. This isn't exactly going to give a good return on investment for AI training data, especially since AIUI the main race for LLM performance now is in good quality training data)
why "non obvious"? The easiest explanation is gathering training data sets, and is mostly caused by AI companies guarding their pile of essentially stolen IP (given how little they care about copyright) from eachother, without sharing any competition need to get their own and re-crawl to update it too
do we have the same understanding of what fomo means?
From what I understand the pressure has created even Google to be a lot more aggressive than what it was before. Not entirely sure but I believe google has two categories of scrapers, the regular one and a new one for AI
> Residential proxies are everywhere, so why did proxy DDoS attacks mostly come from the U.S.? The answer is money. If you are committing fraud or circumventing content restrictions, a Russian IP address gets geo-blocked instantly. A fresh U.S. residential IP address (especially one behind carrier-grade NAT and harder to block individually) is “gold.” Customers pay up to $95 to lease a single U.S. residential IP address for 2 weeks (versus $0.30 for an Eastern European IP). Compare that to your own ARPU per subscriber and sit with it for a second. When an individual IP is worth more than the customer relationship behind it, you don’t have a technical problem. You have a market problem.[1]
1: https://www.nokia.com/blog/one-year-later-the-residential-pr...
Unfortunately you cant have it both ways which makes it hard problem to solve for everyone
This has zero to do with AI (it's just a convenient scapegoat that fits the narrative). Absolutely everything to do with those pushing for centralised control of the Internet.
It's going to change a lot the internet we knew, unless somehow we managed to agree on a common high quality dataset repository.
The status quo with those js PoW pages doesn't really benefit the server owner at all, it's wasted energy.
I mean proof of work is always wasted energy, but I figure it's better to kill two birds with one stone.
CoinHive was one example of this. (I think this is a correct link? https://github.com/cazala/coin-hive). Although I think ideally you would want to have some sort of browser plugin or app that runs on bare metal instead of a proof-of-work in the browser, because RandomX is designed such that it's slow when implemented in JS (https://github.com/tevador/RandomX/blob/master/doc/design.md)
Genuine question, I'm not up to date on how Cloudflare operates right now
I browsed a lot of those sites when I was a teenager and did not have a credit card or easy access to internet money.
- https://ourworldindata.org/grapher/share-living-with-less-th...
Another option I was thinking of would be a pretty inflationary (or demurrage) cryptocurrency in which you have some sort of RandomX or other CPU-bound PoW. A web server could act as a mining pool and use mining shares interchangibly with micropayments.
You could do this mining-share method with Monero right now, it's just that you have higher transaction size in Monero and no real analogue to Bitcoin's LN-based microtransactions. Also you would want the cryptocurrency to be more inflationary (or demurrage-based) to promote usage.
Monero's FCMP++ lays some groundwork for payment channels, but it still lacks the nessisary timelocks. Also there was DLSAG which could have enabled payment channels I think, but it's no longer relevant. I also insist that you would need to change the tokenomics to favor greater inflation (maybe you could make coinbase scale linearly with hashrate?), otherwise the miner reward would be economicially insufficient.
1) it requires you to open a lightning channel. So you need some amount of initial investment (which usually means you need a credit card and an account on a bitcoin exchange), not just any computer which is the issue that using mining shares solves. The initial investment is my main grudge, as it is way too much friction for 402 Payment Required applications. Also as bitcoiners like sztorc note, this isn't feasible for most of the world's population due to bitcoin's small block size.
2) It's not private. People are going to be linking these things to their identities on crypto exchanges. So now feds can basicially track everything you do on the internet that way. LN is only private in the sense that not every transaction is broadcast to everyone else on the network, which is a very low bar. Chainalysis is possible with the right connections.
To a lesser extent, 3) Centralization in the lightning routing protocol which contributes to the effect of #2.
2) it is at least as private as cash. If you withdraw cash from an ATM, the serial numbers are linked to your account. But you can simply spend that cash and obtain it from people you know and transact with, in which case it is private. So yes, if you acquire lightning from an exchange, it is not private. But those lightning transfers you receive outside from an exchance, are untraceable and private. And where you send/spend them too is always private and untraceable.
I am not opening libertierian and freedom-of-money discussions, it is all about being able to facilitate micropayments over the internet. In contrast to traditional banking and credit cards, lightning serves that purpose quite well because it is decentralized because it is based on bitcoin.
Re audience, the audience for Gentoo bugzilla went just to zero. If even 1% of the current bugzilla users would use the micropayments, that's literally infinitely more organic human visitors than today.
If it is the latter, then, perhaps, some p2p (torrent-like) content delivery system could help.
You're not wrong, it is time the scraper pays, but you're going to have to fund it still.
I don't claim to have a better answer.
I want a perfect solution that is free, putting pages up used to be free, now it isn't.
You mine crypto and my free service remains cost-neutral. Yes, that transfers some of the cost to you, but it is marginal.
You want to bang on my servers with the fury of a thousand madmen, then I'll scale up the servers and you can pay the marginal cost of your access.
Or micropayments, of course, but hard to get normal users to sign up for micropayments. Micropayments could of course be the way to bypass the crypto mining gate.
This is not only unfair to legitimate users having to pay.
It also needs to explain what actually happens when the Ddos succeeds (I know, you are talking about scrapers, but what’s the difference really?). Does that mean that the attacker just gets to shrug it off.
Don’t take my statements as facts, I just want to outline a few reason I could come up with that show a purely technical solution might not be enough.
Akin to “just use Cloudflare, it’s free”.
Yes it works but why conceding defeat and say “oh from now on you depend on a business to publish a server”?
I mean, if residential proxies are tools used to DDOS websites, why is Google advertising them? And if you have Google ads for them, can we really say “ah, if we only knew who these guys are?”.
But in the case of anubis it's not even used for crypto. It's just wasted.
>You mine crypto and my free service remains cost-neutral. Yes, that transfers some of the cost to you, but it is marginal.
No, the problem is time wasted. I don't care about the electricity cost. Spending 10s to solve a challenge on a 5W SoC translates to 0.0005 cents (yes, cents, not dollars). Meanwhile if I click on a link and it doesn't load in 5s, I'm seriously questioning the value of your blog or whatever, and will probably just close the tab.
We direct scraper traffic to a bot-specific server using Cloudflare's load balancer, slowly analyzing traffic and adding conditions one at a time. No accidental scraper DDoS in a long time.
Most scrapers are relatively honest in some way shape or form.
Did you miss a "dis" in there?
Scraper "attacks" don't take down our robot-specific server very often; it's safe for us to take heavy-handed approaches that sometimes redirect users there. 99% (made-up high number) of the time, the misdirected users don't realize anything is amiss.
Start by analyzing your traffic, specifically user agents. Look for "robot" or even "bot" in the user agent and load balance those to a robot-specific server. This can all be done within Cloudflare. The only code is the user agent condition. Note: I'm very open to input here if anyone reading notices that we're shooting ourselves in the feet. Based on our analysis, the remaining traffic is a good picture of our human users.
We have loads of other conditions, mostly balancing specific IP ranges for entities when we know exactly who they are, but this is a good start.
What's the point of this compared to letting cloudflare handle everything automagically? Presumably whatever heuristics they come up with are going to be better than you can, given limited time and budget?
To your point: I think you're simply mistaken that they offer a good scraper-detection solution, but I'm all ears.
That's it, I guess?
Also: where exactly are AI companies incentivized to be anything but shitty 'neighbors'?
Spam, DDoS attacks and other network abuse used to cause your hosting company or provider to call you and tell you to knock it off or you'll get disconnected, if your provider was reputable. If your provider wasn't reputable, it was likely a matter of time before they would get a nasty call from their upstream provider.
Now it just gets you a thank-you from the sales team for all the bandwidth you bought.
Meanwhile, do any of the cloud providers have any incentive to do anything about this? Hell no. They're making money off you having to ramp up extra or bigger instances. They're making money off the bandwidth. They're making money off the people doing the crawling, too. They're incentivized to do the exact opposite of effectively help you with your AI bot problem.
I get what you're asking, and I'm wondering the same. Not all sources are created equally and we see the results all the time. LLMs outputs nonsense all the time, like Flock cameras containing 5 grams of gold and ounces of copper, because they are completely on critical of their sources. Perhaps there's some weights that says: Kernel mailing list, MariaDB documentation and Microsofts Learning sites are 100% trust, Reddit 50%, 4Chan 10%, but I doubt it.
Anthropic might care a little bit, seeing as they scan books, but again, is it just all books? Because other than some flowery language I don't really see the point in scanning a 1970s paperback only spy novel.
This is not a tech problem. This is about what should or should not be legal.
Nor it’s a question of having time to implement solution X or Y.
If someone attacks you yes, you should have better security but you also need to have legal recourse, or it will never stop.
Ddos is already illegal.
I am not a lawyer so don’t ask me for exact resources, which vary by country anyway, but stop treating scrapers as an inescapable force of nature.
If you create a law that says you have to honour robots.txt files, what do you do if an IP from another country fails to do so?
Then you report a crime and send the logs. And maybe report to some politician that is tough on crime. And you are not alone.
I am sure you follow through, Pls check also my other comments.
Laws without enforcement are meaningless, without jurisdiction they are less effective.
Maybe yes, some poor guy will find a letter that says he needs to pay 50 USD or EUR to some faraway country because the puzzle app in his tv has weird ToS. Or whatever, a cease and desist letter? A “you are part of a hacking circle please uninstall this app” letter, signed by a judge?
I don’t advocate for people getting thrown in jail, I just think that the whole “residential proxies” issue might be a charade that would not last longer if citizens and businesses would claim rights instead of play victim.
I also really dislike grasping for a tech solution because I don’t want to start to mine crypto to read an article or to depend on some “free” cloud service to host an open bug tracker.
It just feels so weird: why is it that no-one is doing it? I feel like I an one “comment away” from someone explaining to me how naive I am, and yet it’s never coming.
In whatever country they are.
I am not saying this is the magic bullet that will make it all go away. Some will, some will not.
Enforcement’s aim is not to completely eradicate crime: make it anti-economical for the majority and it’s a good deal enough.
Go to Berlin and torrent a movie: you will be fined, for a few hundred euros. Can you VPN and get away with it? I suppose so, but your average Joe is just not going to do it: why pay 10 euro per month on a VPN and risk? Just rent it.
DDOS is also, in general, a criminal charge. What would happen to prices once you get a few inducements? If it’s a crime and enforced, what are the roles of App Stores that let users download residential proxy? Are they also committing a crime? Residential proxy can be used for legitimate purposes? I don’t think that line of defence worked well for Pirate Bay, no?
And yes pirate bay is still around, some scraping will always be there.
Stablecoins have a great use-case and you get paid for bots to access the site with the humans living in peace without the site getting botted.
Job done.
Though, checking, I see that it requires a login but is still up. I misunderstood what 'closed' meant. This seems fine.
A headline would be great if those AI companies would close down. I hold them all responsible for this.
However, it is also easy to get large numbers of v6 addresses cheaply.
Maybe Taler could help ?
You can link to them in 10 million HTTP 429 responses, they will still ignore them.
The biggest thing about what's happened today though, is how a "janitor" on the infrastructure team can shut down the entire distros ability to work on bugs. The decision to shut down should have been left to someone else like robbat.
Mgorny has consistently caused issues for Gentoo, and many people have left the community after having to deal with him, but it seems comrel/devrel can't/won't remove him for whatever reason.
Edit: it looks like bgo is up now. Hopefully it stays that way.
Would fronting it with Anubis have helped at all here?
(please do not ramble about how requiring javascript is a violation of your human rights)
I've taken #Gentoo Bugzilla down, because it was unusable anyway.
No point in feeding the #LLM scrapers that are using thousands of different IPv4 addresses, with no obvious patterns I can see.
Now you might try to interpret what he said as the LLM scrapers made it unusable, but apparently it was usable to the scrapers as he turned it off to stop feeding them. If he can actually attribute the traffic to LLM scrapes, then he can use that same method to identify their requests and block them or shape (slow) their traffic.He then says not to tell him how to fix it because he doesn't care.
Then you've got good old cloudflare which is free to use
2. Considering the spike of complaints over the past two years and the fact that they range from small time site operators to larger orgs (Linux Kernel, GNOME, Duke University, etc), do you think that maybe you're missing part of the problem space?