Let us serve you, but don't bring us down
blog.archive.org
blog.archive.org
I was keeping an eye on it, because we are hard-capped at 100 QPS to our provider, beyond that and they start dropping our traffic (and it is an outside provider, bundling domain registries like verisign and stuff), which makes regular users break if their traffic gets unlucky.
Anyway, after a week of 40qps, they start spiking to 200+, and we pull the plug on the whole thing: now each request to our endpoint requires a recaptcha token. This is not great (more friction for legit users = more churn) but it is successful. IF they had only kept their QPS low, nobody would have cared. I wanted to send some kind of response code like, "nearing quota".
FTR before people ask: it was quite difficult to stop this particular attack, since it worked like a DDOS, smeared across a _large_ number of ipv4 and ipv6 requesters. 50 QPS just isn't enough quota to do stuff like reactively banning IP numbers if the attacker has millions of IPs available.
I first thought it was something like "quadrillion bytes per second" or some newer over-the-top data measurement :)
If they’re ipv6 address wouldn’t they be safe to block across large ranges?
It is also trivially “easy” to get past reCAPTCHA, but it costs more. My guess is that a domain name checker tool isn’t worth the cost per request to bypass reCAPTCHA (approximate 0.02 cents per session)
It’s hard to “throttle” a single isolated request from a lone IP.
There's only so much you can really do when your underlying resource is so limited. Luckily the value of the query is lower than the cost of a recaptcha solve, so the attackers moved on to some other target.
Ironically I could now turn off the endpoint protection (or have it responsive to traffic load), until the attackers return. I shall not go into too many details, no need to give people a map.
The webmaster doesn’t need to worry about it, the anti-bot services handle who gets what difficulty of challenge. But the webmaster can specify whether they’d like to be more or less strict/difficult than usual.
The main idea here is that at some point just leaving the scrape running in a way that didn't overwhelm my backend would have resulted in me not caring enough to do that. But now they get nothing. Even if you're borrowing bandwidth and not paying, you should be a good neighbor is all.
Though thinking about it, I wonder if there is a hybrid: start with a difficulty that's just a few seconds for CPU/WebCrypto and ramp up quickly, but also support WebGPU where possible so that web users on abusive connections may still succeed? I am not sure though, I guess this depends on the feasibility of using WebGPU and etc.
I commented in the past about it:
"At first glance, yes, we can create intentionally expensive computations without relying on a blockchain, that would serve the same purpose. In reality we cannot. Special computer hardware (ASICs) could generate much cheaper PoW annotations than general purpose computers, and sell it to spammers. Blockchain economic incentives ensure that ASICs will be used by the miners first and foremost."
So instead, what, you want people to buy BTC and send small amounts of it to websites?
Problem is, given varying income levels across society and across the world, one person's expensive micropayment is another person's almost free micropayment.
Very true, that micropayments could vary on their relative cheapness across the global population. Let's put a number, is 0.01 cent affordable by most people on the planet, for every http request?
An obscure coin my users cannot purchase through major exchanges is a really poor solution. Sorry, that makes no sense.
But again, economic incentives need to work their way into the system.
In case you run out of proof_of_burn a friend of yours might send you just a cent, and you are good to go, for one hundred thousand http requests more. There is no need for a normal person, to hold more than a handful of dollars for every year's internet use.
That renders exchanges almost useless. Not totally useless, but much less relevant than they are today.
I personally wouldn't care less, if there is one bitcoin/blockchain which reaches that sweet spot, or if they are a hundred, including litecoin, ripple etc. PoW is meant to be used for practical reasons. Blockchain however constitutes an economic system, of suppliers-miners and consumers-users, it is more than just a software program. It will evolve in the future, and it requires some crucial time.
For the moment there no API which provides the kind of PoW service to be very useful, and some hacks might be required to mitigate side effects of relentless scraping. These are just hacks, useful today, but the real solution is coming soon. Web3 some people call it, or Cryptocosm is another name of it.
Right now, if I was to ask my users to do a proof of burn for $0.0001 of BTC most of them would just close my site as they don't have any BTC. The process of setting up an account on an exchange, waiting several hours/days for KYC checks to clear, adding a credit card, buying BTC, sending it to a browser extension, and then trying the signup again is a *significant* initial hurdle. If we were in a world where I could assume all my users already owned BTC, that's a different story, but we haven't seen adoption of cryptocurrency anywhere near that level. I don't expect PoW schemes to drive that adoption either, so this seems like a poor solution today.
Do you happen to know if a hCaptcha/mCaptcha-like micro-proof-of-burn tool exists already? I'd be happy to be proven wrong.
That brings us back to options that exist today, which includes every user computing their own PoW. Looking at mCaptcha some more, it uses a SHA256 derivative so it's compute-hard and vulnerable to GPUs/ASICs. The author mentions some of those concerns in an earlier thread [1]. I wonder if a different proof-of-work algorithm would be better, like a memory-hard PoW, proof-of-space, and/or proof-of-wait. I'm skeptical of those too, unfortunately.
In my calculations, with millionth of a cent per transaction, even paying for torrent blocks (64 KB), not torrent pieces (16KB) will soon become profitable.
Unlike blockchain systems, the implementation details of mcaptcha are also totally hidden, and you don't have to maintain compatibility at all.
Someone creates an asic? Great, you can make it not work anymore without affecting any of your users for real.
10+ million down the drain.
I added FriendlyCaptcha to some of my sites, and stopped 100% of abusive traffic. Open source, user friendly, accessible to people with disabilities.
Most of us are not running amazon.com here.
If i may add, in case someone desires a little bit of revenue from a website, one very popular solution is to put advertisements in some places. Well some people consider that a security hole, including me. So i guess the definition of security varies, but the security mania goes on for many decades. I am one of those security maniacs, and any tool to enhance security is important, blockchain is one of them.
I've found that black and white thinking in security is very dangerous, as you often end up with very "secure" controls that have terrible UX, which users bypass completely via byob etc... And pwnage ensues. UX is a primary pillar of security.
We have pretty good guides to what's an expensive computation to everyone. That's how password-hashing algorithms work, much better than BitCoin does. Besides, the existence of specialized methods of compute is hardly determinative. We're not trying to stop TLAs. It just needs to be expensive enough to stop 99%, and at most we'll later update again.
ASIC's take tons of time and money to design.
Unlike blockchain, where the PoW algorithm is part of basic compatibility, it's not at all in mcaptcha - the users see a checkbox. The rest is implementation details.
If someone creates an ASIC, you can break it very easily by changing the PoW algorithm a bit, and no users are affected.
Even if someone has the money to keep up with you, which is remarkably expensive, they will be too slow right now.
That would be hard to change
The genius of bitcoin (not BTC) is that it provides an organized and practical way, for PoW to be used by everyone on the planet. Some people find it strange, because there is an imaginary token created out of pure nothing which represents the PoW. It would work fine, without that imaginary token i.e. bitcoin, just by dollars of euros, but it wouldn't be internet native. The point is always to just send a minimal PoW with every http or tcp/ip request.
Bitcoin is an evolution of that idea applied to digital cash.
IMAP isn't used for sending email. You're thinking of SMTP.
Hardware specialization breaks the economics behind a CAPTCHA. To fight that you need to use a PoW that hasn't been ASIC'd yet, and be willing to change PoW functions at the drop of a hat. PoW functions that stress memory or cache are also helpful here, though you run the risk of browsers flagging you as a cryptominer (which is technically correct, even if economically wrong).
Seems like it should be the same cost (barring the friction of having a wallet, etc.) for both sets of people since Bitcoin is just a commodity and the value is the same to everyone, miner or not.
For example, a miner should value some fraction of a BTC the same way anyone else does, since they can sell or buy it at the same price a normal user can. The fact that they can profitably mine BTC just means they have a profitable business on the side, it doesn’t mean they should prefer to pay for services (or emails) in BTC.
So in order to moderately inconvenience said spammer, you have to make each and every ordinary user wait hours mining a few satoshis' worth of hashes in order to be let in. This is the exact opposite of what you want.
But even if you pick a novel function that doesn't have special-purpose hardware, spammers can still optimize their setups to lower the effective unit cost below what a legitimate user faces, by doing normal miner activities (picking hardware, scaling up, moving to where electricity is cheap, etc.).
Since your goal is to maximally discriminate between legitimate and spam use cases, you'd want spammers to face at least the same per-unit cost as legitimate users.
What's one way you can do that? Well, how about charging actual currency, whether Bitcoin or fiat? Money has the useful property of having the same nominal value for everyone, and not being amenable to further optimization.
In short, forcing users to actually run PoW themselves doesn't really make economic sense. Even if you're avoiding existing hash functions, it's mostly worse than just charging money because spammers have better ability to optimize against it.
And if charging money doesn't work, switching to local PoW is unlikely to be better.
FWIW this thread inspired me to implement Cloudflare Turnstile on one of my pages - highly recommend. As compared to reCAPTCHA your users are never wasting their life away clicking on traffic signs, and as compared to mCaptcha you don’t have to host the server yourself.
As the problem with captchas, is rather who profits from them and who could track users, it seems that with an OCR system to be protected, it is most reasonable to actually give OCR tasks to humans. IMHO this is far more sustainable and people would understand the value: there is nothing bad in improving ML in general. Maybe someone could even define a sensible PoW task for OCR but I doubt it...
I think this whole thing is a big hurdle just because I'm unable to solve visual puzzles. Besides, having a company collecting email addresses of people who are disabled in one way or another and giving them an identifying cookie is a privacy/data disaster waiting to happen.
That being said, I think the audio alternatives for visual CAPTCHAs are also unacceptable. Even if you can hear them, they may be hard to solve especially if they are not provided in your mother tongue. I think we can and should be able to do better by now
If you're going to waste my energy to do proof-of-work anyway, I'd rather you use it for something useful (even mining crypto-currency to pay for server costs) rather than let it go to waste.
eg. in your case this could mean if traffic is above eg. 75QPS then captcha is enabled, and if it's below that it's disabled.
I don't know what tech stack you are using, but nice trick that i figured out was to abuse rate limiting to detect global traffic (doing if branch with rate limit with const as client id)
For a while, we had to just set off pagers when global traffic exceeded a threshold and manually toggle the extra hardening, but eventually it became a lot more reactive.
I guess your tool is asking the registries if domains are registered.
host example.com
If this returns with an IP address, no need to talk to a registrar. Only if there is no IP address, I go with whois example.com host -t soa example.com
which should give you domains that have any DNS record at all, not just an A record.If the registrant creates any DNS records at all, then the SOA will need to exist for the zone to be valid; but I don't recall whether the registrar is or isn't required to publish a zone for a registered domain that otherwise contains no records. (Also whether the registrar is required to inform the registry of authoritative nameservers for the domain at all times from the moment that the domain is registered, and then whether that information would lead to the synthesis of a SOA record.)
I guess I can either try this (if I could find a no-frills-enough registrar that doesn't create any records at all for add-on hosting services or domain parking) or try to take some more ICANN coursework to find out the answer.
Or maybe someone else reading this thread knows whether we can have a domain in practice that is registered but has no published SOA.
This reminds me of the RMT driven botting problem in WoW (World of Warcraft). Instead of fighting the neverending game of cat and mouse against botters, Blizzard just decided to supply the long reprimanded demand for in-game currency by creating the WoW token, and they make money while they're at.
He knew that was the reason for the intentional DDoS due to messages (along the lines of “think it is funny to poison my information do you?”) in the query string of the bulk requests. Like unpleasant fools making a big noise in shops because they aren't served immediately after cutting in line or some such, the entitled can be quite petty when actively taken to task for the inconvenience they cause.
Passive defences are safer in that regard, assuming you don't get into an arms race with the scrapers, though are unfortunately more likely to mildly inconvenience your good users.
Though I did have success once stopping hot-linked images by serving up images that didn't fit well with the sensibilities of the internet forum that was hotlinking them. It had the advantage of looking like the people in the forum posting had deliberately chosen to post the image. Serving different images depending on client ip made for fun too, as they argued with each other about why they posted "that".
Evil. I like it.
If somebody wants to access endpoint you might send him a challenge first. Random text. The client must append to the text some other text chosen by him, so that when you calculate sha256 on concatenated text, first byte or two of it will be zeros.
To access your actual endpoint client needs to send that generated text and you can check it if it results in the required number of zeros. You might demand more zeros in times of heavier load. Each additional bit that you require to be zero increases number of random attempts to find the text by a factor of two.
To make stuff easier for yourself the challenge text instead of being random might be a hash of clients request parameters and some secret salt. Then you don't have to remember the text or any context at all between client requests. You can regenerate it when the client sends second requests with answer.
Honestly I don't know why this isn't a standard option in frameworks for building public facing apis that are expected to be used reasonably.
Generating challenges is really cheap. Calculating a single SHA256 or sth out of a querystring+salt (or better yet 128bit SipHash). Generation and validation can be done on separate layer/server so requests without valid PoW won't even register on your main system.
I actually do like your puzzle challenge idea from an obfuscation standpoint of making an API less attractive for clients that aren't your own, though. A challenge system + an annoying format like Protobufs instead of JSON is too much work for a lot of abusers.
Biggest problem with DDOS though is that if the volume is even reaching your application server, you're probably hosed.
It means that on this hardware they can make one request per two seconds not thousands per second.
And it affects all of the hardware under the control of an attacker, so his attack becomes thousands times less dangerous.
> Biggest problem with DDOS though is that if the volume is even reaching your application server, you're probably hosed.
That's why I'm suggesting that challenge generation and validation can be done on separate machines. So your application servers can be safe.
Obfuscation is a valid point too.
I thought about it for a bit, made some experiments. Now I think the challenge should be completely random but the server needs to keep track of recently issued challenges and solved challenges to prevent reusing of already calculated solutions. I think bloom filters would be perfect for that because some small percentage of false negatives doesn't matter.
How about they actually fund something really worthy of preservation? Of course it is archive.org role to reach out first. For example EU could fund an archive.org mirror in the EU (with certain throughout etc).
Of course opponents of public/government funding have a very good point in that many organisations when they get public money, they find a way to burn through all of it in a lot less efficient way. This can be mitigated by attaching concrete conditions to the grants. One example is a mirror in a specific location.
Since they are a U.S. 501(c)(3), they also publish annual reports, which can be downloaded e.g., at the ProPublica Nonprofit Explorer
https://projects.propublica.org/nonprofits/organizations/943...
They may be "of very questionable value" to you but your solution to remove funding from them to channel it to a much wealthier organisation in a wealthier country is neither ethical, legal or practical.
There are indeed lots of cultural projects that get funding, that I consider not that important in comparison, but of course those people involved would think different. (Opera for example is heavily subsidized)
For real? I am not against subsidicing art and culture, but I am against selective subsidicing. For example in germany there is a strong divide into "serious art" like opera and classical music that gets lots of money direct or indirectly - and trivial art, everyone else. Getting allmost nothing. So it boils down to taste and the favourite culture of the establishment. But there is so much other good music and performers besides the mainstream out there, who gets categorized into "entertainment" and have to struggle on their own.
So back on topic, I would be fine with taking money from opera to give it to internet archives. But of course, I rather would have more money for everyone involved in arts and culture.
The suggestion was well-motivated but presumably guided by a lack of understanding of the existing digital landscape of international publicly funded projects and their obligations and constraints.
1: https://blog.archive.org/2017/04/17/robots-txt-meant-for-sea...
The torrent protocol was meant to relieve this level of server load in mind.
By "loudly", perhaps actually redirect to the .torrent instead, for those who trip the rate limit?
BitTorrent also supports web seeds and they don't even really have to keep a full client running, just embed an HTTP link into the .torrent file.
This approach allows the internet archive to effectively rate limit hosts without making content unavailable entirely, and allows others to help carry the load.
As far as I know the proof-of-data-availability folks (whether using a blockchain or not) are the only folks trying to solve this problem with a serious technical solution at a cost scale lower than "rent your own server and have deep pockets for bandwidth" at the moment.
It's not about "willing to spend money", as if there is only one threshold to pass.
Torrents are used largely because they diffuse the cost of hosting, so that people (and organisations) don't need deep pockets to distribute large files to many others. There is an inevitable infrastructure cost but it's spread out more fairly among users.
The proof-of-data-availability stuff is similar, but takes it further so that a wider range of data stays publically available to whoever wants to download it, whenever they want, than it otherwise would be. It is really just a more fancy version of torrenting that is less prone to the excessive per-file popularity fluctuations that torrents suffer from. The objective is to lower costs compared with the current state of the internet, not add more. And it does not specifically require a blockchain.
If you know of something else tackling the problem I'd be love to hear about it. Torrent seed sites don't qualify, as they don't solve the problem: Most files aren't available on them either, and you have to pay for the more obscure content they do have.
You are proposing a system based on financial rewards for hosting. Who pays those rewards, for those files in which there is no interest? If archive.org is to pay for it, we are back to square one. They are already very good at hosting content, within the limits of the resources they have. If people with an interest in downloading the files pay for it, no go, files with no interest go away. If you propose that people currently abusing the free service of archive.org, to the point of bringing it down, would pay a fee per download, you must be joking.
No, I'm not. You are incorrectly assuming that blockchains are necessarily financial or that cryptoeconomic incentive structure involves net pay to someone.
> If you propose that people currently abusing the free service of archive.org, to the point of bringing it down, would pay a fee per download, you must be joking.
I'm not proposing that.
People "pay" for hosting by participating in some amount of upload to offset their download, in order to be granted higher download rates. That is the same principle as BitTorrent has used since its inception: Upload is measured and download is traded for upload to ensure users choose to upload for a while.
The difference is that information is split and diffused in a different way, which ensure that some amount of upload bandwidth and temporary storage is available for less popular data as long as there are people participating in the network, mostly when they are downloading something more popular and providing some upload in exchange. The network power law helps by ensuring the long tail of less popular content needs relatively little "extra" bandwidth so it is not onerous on the users who, most of the time, are downloading and storing popular data.
No money needs to be involved.
Probably some financial things will emerge much like paid Torrent sites do at the moment for enhanced access which some people prefer. You can't prevent them, but they are not required for the network's operation nor required to access it.
So as not to piss them off (and so they don't try to block me), my script will take about 6 hours. Between each page fetch it sleeps for a small, random amount of time. It's been working like that for years.
Funny story - at work we once had a huge spike in requests from a single IP. We all crowded around, thinking it was some malicious hacker from France. How exciting - we're now interesting enough to warrant a DoS! Turns out another team in the company was just pulling all our data into Algolia to improve search. They were clearly not very courteous!
So on the other end (building APIs) I certainly do pay attention to traffic and have Grafana alerts set up around it.
In general I would rate limit by IP anything connected to the internet.
Can you provide an actual example? I see it come up a lot in these conversations, but I'm really skeptical that anyone actually does this analysis. It seems like rate analysis (either requests or bandwidth) would achieve the same result in a far simpler manner, so I suspect that is what actually happens.
Merely for your consideration, they actually do a great job of indicating in the response how many more requests per "window" the current authentication is allowed, and a header containing the epoch at which time the window will reset: https://docs.github.com/en/rest/overview/resources-in-the-re...
I would suspect, all things being equal, that politely spaced requests are better than "as fast as computers can go" but I was just trying to point out that one need not bend oppressively over the other direction when the site is kind enough to tell the caller about the boundaries in a machine-readable format
I am sorry, admins! I won't do it next time.
It just strikes me as surprising that sites are still dealing with problems like this in 2023.
The idea that a site is still manually having to identify a set of IP addresses and block them with human intervention seems absolutely archaic by this point.
And I don't think different sites have particularly different needs here... basic pattern-matching heuristics can identify IP addresses/blocks (plus things like HTTP headers) that are suddenly ramping up requests, and use CAPTCHAs to allow legitimate users through when one IP address is doing NAT for many users behind it. Really the main choice is just whether to block spidering as much as possible, always allow it as long as it's at a reasonable speed, or limit it to a standard whitelist of known orgs (Google, Bing, Facebook, Internet Archive, etc.).
It just strikes me as odd that when you follow a basic tutorial for installing Apache on Linux on Digital Ocean, rate limiting isn't part of it. It seems like it should be almost as basic as an HTTPS certificate by this point.
Obviously if someone is attempting a large-scale DDoS your servers can't handle it and you'll be using CloudFlare for its scale. But otherwise, for basic protection against greedy spiders who are even trying to evade detection across a ton of VPN/cloud IP's, this strategy works fine. It's exactly the kind of thing that I would expect any large website to implement.
If there isn't an open-source tool that does this, I wonder why not. Or if there is, I wonder why IA isn't using something like it. But heck, IA wasn't even using a simple version -- it was just 64 IP addresses where basic rate-limiting would have worked fine.
Obviously judges aren’t going to have to worry about reasonable rate limits but if these DDoSes are rare, I’d much rather they dealt with them on a case by case basis. Without some complex dynamic rate limit that scales based on available compute and demand, rate limiting would be a blunt solution that will necessarily generate false positives.
[1] https://www.theregister.com/2018/09/04/wayback_machine_legit...
Virtually every major website has per-host inbound request limits. This is completely standard practice and Archive.org is the odd one here.
And DDoS isn't the only concern. Legitimate users that run poorly written Python scripts that make insane numbers of requests can hog server resources, and rate limits with appropriate error messages linking to resources documenting efficient access patterns can improve the experience for everyone, and drastically cut costs for the service operators.
What an old fashioned concern! Don't you know that the US justice system uses ChatGPT as an archive retrieval system nowadays? /s
They don't care. And yes I've heard stories of people "finding a service" that does something basic and just dumping their whole traffic onto it
(of course they play the victim once they're found out)
I get about 20,000 queries per day that may be human.
> The search engine is currently serving about 36 queries/minute.
To:
> The search engine is currently serving about 36 real queries/minute, and deflecting 1806 bot queries/minute (please don't).
* If they had a better API (a simple non-synchronous API would be enough, one where we could send a list of URLs would be even better), one could have made a lot less calls.
But failing that I'm likely to implement a web proxy that utilizes the wirefilter library that is already open sourced by Cloudflare (but no longer updated).
The tools to stop these attacks are reasonably trivial.
Most of the art of stopping them is observability.
When you can see every dimension of a TCP/UDP connection, the TLS handshake, the HTTP communication... Then the dimension by which an attack is being conducted is glaringly obvious.
Once you have the very obvious correlation, then you only need a blunt instrument of a tool that can block/deny/nullroute the attack.
What's hard isn't the block rules, What's hard is instrumenting enough observability to see the obvious correlation of an attack.
Even on my personal nginx a single request logs about 20 things. But there are literally hundreds of properties one can log if you have access to the entirety of the stack, and if you have access to log on it, then you have the ability to carry that context to a point in the code where you can block it.
Also, 10k qps really isn't a lot. So this should be treated as a warning sign, this attack was low volume even for amateur booter services.
over 50MB in 1minute (6MB/hr).
http://dumps.wikimedia.your.org/enwiki/20220820/
They also seed torrents
$ curl https://dumps.wikimedia.org/enwiki/20230520/enwiki-20230520-pages-articles-multistream.xml.bz2 -o/dev/null
% Total % Received % Xferd Average Speed Time Time Time Current
Dload Upload Total Spent Left Speed
2 20.4G 2 476M 0 0 4387k 0 1:21:38 0:01:51 1:19:47 4393k^C
$ curl https://wikidata.aerotechnet.com/enwiki/20230520/enwiki-20230520-pages-articles-multistream.xml.bz2 -o/dev/null
% Total % Received % Xferd Average Speed Time Time Time Current
Dload Upload Total Spent Left Speed
0 20.4G 0 45.9M 0 0 1953k 0 3:03:23 0:00:24 3:02:59 2257k^C
$ curl https://mirror.clarkson.edu/wikimedia/enwiki/20230520/enwiki-20230520-pages-articles-multistream.xml.bz2 -o/dev/null
% Total % Received % Xferd Average Speed Time Time Time Current
Dload Upload Total Spent Left Speed
1 20.4G 1 344M 0 0 35.2M 0 0:09:56 0:00:09 0:09:47 36.8M^C
maybe you should check your own internet connection first?Well, that's ironic.
... Anyone have an archive.org link of the page?
On the other hand, Cloudflare has quite some experience with DDoS protection, so they definitely have the right infrastructure in place to stop this kind of abuse. What is the bill going to be, though?
10,000 requests/seconds is nowadays an average number for load testing in the companies I work with, for websites / mobile apps deployed at this scale
(1) Limit your load / parallelism: do not use more than 4 threads, and if the site takes more than a few seconds to respond, use fewer.
(2) Limit aggregate load: after each request, sleep/wait at least the amount that the previous request took to get served.
(3) If you need more than that, ask the site owners for direct channel.
This way if multiple crawler happen to crawl the same resource-limited site, the site may have a fighting chance.
I recall tripping the rate limiter at least once by accident - I was clicking around the archived site for reference while working on creating a restored version.
1. whatever interface you expose to the public, some will abuse it
2. rules/laws/countermeasures are then added and enforced, often universally (though sometimes inconsistently)
3. which then bites, annoys, insults or otherise burdens or adds to the prices paid by ALL the OTHER non-abusing users pf that same interface
always. eventually. every time
It's infuriating.
For example: https://aws-new-features.s3.us-east-1.amazonaws.com/update/2... (see the `NEW_ip_prefixes` and `ip_prefixes_REMOVED` keys)
It should teach them a lesson to have their AWS account banned.
They don't even respect robots.txt. So content creators can't even opt out of that. Not that copyright would have copyright holders having to opt out of copying in the first place.
How have they not been sued out of existance yet?
Copyright law specifically allows for libraries and archives to make copies of copyrighted material.
Without such laws, without libraries, knowledge could not be guaranteed to be shared freely among the public, resulting in ever growing knowledge and education gaps between those with means and those without.
Edit: Since you asked for the legal ground, here it is specifically:
https://www.law.cornell.edu/uscode/text/17/108
And here’s is further discussion by the copyright office itself:
https://www.copyright.gov/policy/section108/discussion-docum...
Congress and the copyright office have made it an important part of copyright law to protect archival, library and fair use doctrine.
I don’t think it’s helpful to flag things people disagree with, as long as they don’t attempt to spread misinformation, or trolling etc. The parent phrased the topic as a question, meaning I believe they were open to understanding.
I also think it’s relative to the topic posted, as we’re talking about either an attack on archive.org, or a massive recopying of archive data by an unknown.
Talking to people with opposing views is important. Let’s not just shut people down if we don’t agree with something, especially when someone is asking a question.
Since archive.org have now updated that the problem scraper is now evading countermeasures and bringing them down repeatedly, and has been identified as an "AI" company, it could end up being an existential risk: if companies start using archive.org as a large-scale commercial IP theft proxy, they will likely face even more legal challenges than they already do.
That said, personally I do use Wayback machine about once a month for genuine search of historic content.