The Internet Archive is under a DDoS attack
mastodon.archive.org
mastodon.archive.org
I have a `wget-mirror` shell function invoking wget with all the trimmings that takes care of 99% of sites. I’ll edit the full command into this comment when I get home if anybody else wants to start doing the same :)
(Oh, never mind YouTube videos that I once added to playlists ... that later disappear leaving only holes in my playlists.)
It is probably easiest to save the render as a picture and then store text separately for searchability?
Chromium's MHTML "Save as…" and the SingleFile WebExtension should both save copies of the rendered DOM.
Apparently Safari has WebArchive and Mozilla had MAFF for similar use cases.
I think WARC is supposed to save enough data about network streams for dynamic pages to work. At least on the Wayback Machine, infinite scrolling and "Load More" buttons do kinda work sometimes. You may have to load the archived pages in a browser and try to use each dynamic feature at least once, to trigger requests for needed resources.
SingleFile: https://github.com/gildas-lormeau/SingleFile
LWN on WARC, tools: https://anarc.at/blog/2018-10-04-archiving-web-sites/
Self-hostable web archives: https://awesome-selfhosted.net/tags/archiving-and-digital-pr...
Wayback Machine addons, bookmarklets: https://help.archive.org/help/save-pages-in-the-wayback-mach...
wget-mirror() {
wget --mirror --convert-links --adjust-extension --page-requisites \
--no-parent --content-disposition --content-on-error \
--header="Accept: text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8" \
--user-agent="Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:129.0) Gecko/20100101 Firefox/129.0" \
--restrict-file-names="windows,nocontrol" -e robots=off --no-check-certificate \
--no-hsts --retry-connrefused --retry-on-host-error --reject-regex=".*\/\/\/.*" $1
}
Some notes:— This command hits servers as fast as possible. Not sorry. I have encountered a very small number of sites-I-care-to-mirror that have any sort of mitigation for this. The only site I'm IP banned from right now is http://elm-chan.org/ and that's just because I haven't cared to power-cycle my ISP box or bother with VPN. If you want to be a better neighbor than me, look into wget's `--wait`/`--waitretry`/`--random-wait`.
— The only part of this I'm actively unhappy with is the fixed version number in my fake User-Agent string. I go in and increment it to whatever version's current every once in a while. I am tempted to try automating it with an additional call to `date` assuming a six-week major-version cadence.
— The `--reject-regex` is a hack to work around lots of CMS I've encountered where it's possible to build up links with an infinite number of path separators, e.g. an `www.example.com///whatever` containing a link to `www.example.com////whatever` containing a link to…
— I am using wget1 aka wget. There is a wget2 project, but last time I looked into it wget2 did not support something I needed. I don't remember what that something was lol
— I have avoided WARC because I usually prefer the ergonomics of having separate files and because WARC seems more focused on use cases where one does multiple archives over time (as is the case for Wayback Machine or a search engine) where my archiving style is more one-and-done. I don't tend to back up sites that are actively changing/maintained.
— However I do like to wrap my mirrored files in a store-only Zip archive when there are a great number of mostly-identical pages, like for web forums. I back up to a ZFS dataset with ZSTD compression, and the space savings can be quite substantial for certain sites. A TAR compresses just as well, but a `zip -0` will have a central directory that makes it much easier to browse later.
Here is an example of the file usage for http://preserve.mactech.com with separate files vs plain TAR vs DEFLATE Zip archive vs store-only Zip archive. These are all on the same ZSTD-compressed dataset and the DEFLATE example is here to show why one would want store-only when fs-level compression is enabled.
982M preserve.mactech.com.deflate.zip
408M preserve.mactech.com.store.zip
410M preserve.mactech.com.tar
3.8G preserve.mactech.com
Also I lied and don't have a full TiB yet ;) [lammy@popola#WWW] zfs list spinthedisc/Backups/WWW
NAME USED AVAIL REFER MOUNTPOINT
spinthedisc/Backups/WWW 772G 299G 772G /spinthedisc/Backups/WWW
[lammy@popola#WWW] zfs get compression spinthedisc/Backups/WWW
NAME PROPERTY VALUE SOURCE
spinthedisc/Backups/WWW compression zstd local
[lammy@popola#WWW] ls
Academic DIY Medicine SA
Animals Doujin Military Science
Anime Electronics most_wanted.txt Space
Appliance Fantasy Movies Sports
Architecture Food Music Survivalism
Art Games Personal Theology
Books History Philosophy too_big_for_old_hdds.txt
Business Hobby Photography Toys
Cars Humor Politics Transportation
Cartoons Kids Publications Travel
Celebrity LGBT Radio Webcomics
Communities Literature Railroad
Computers Media README.txt
Some of this could stand to be re-organized. Since I've gotten more into it I've gotten better at anticipating an ideal directory depth/specificity at archive time instead of trying to come back to them later. Like `DIY` (i.e. home improvement) there should go into `Hobby` which did not exist at the time, `SA` (SomethingAwful) should go into `Communities` which did not exist at the time, `Cars` into `Transportation`, etc.`Personal` is the directory that's been hardest to sort because personal sites are one of my fav things to back up but also one of the hardest things to try and organize when they reflect diverse interests. For now I've settled on a hybrid approach. If a site is geared toward one particular interest or subsulture, it gets sorted into `Personal/<Interest>`, like `Academics`, `Authors`, `Artists`, `Goth` (loads of '90s goths had web pages for some reason). Sites reflecting The Style At The Time might get sorted into `1990s` for a blinking-construction-GIF Tripod/Angelfire site or `2000s` for an early blog. Some times I sort personal sites by generation like `GenX` or `Boomer` (said in a loving way — Boomers did nothing wrong) when they reflect interests more typical of one particular generation.
I have encountered "GnuTLS: The TLS connection was non-properly terminated. Unable to establish SSL connection." multiple times, and retry options seem to be useless when that happens. Some searches suggest it could be related to tls handshake fragmentation, but nonetheless wget could retry if related options are used. Manual retry seems to download the missing URLs, otherwise mirroring jobs are randomly incomplete.
I donate to The Archive. More people should too.
Plus for as great of a service as Wayback Machine is, it can be very unpleasant to actually browse. I dislike how it injects its own toolbar into every page (yes I know how to massage the URLs to get the raw page data, but it isn't browsable that way). Have you never encountered sites in Wayback Machine where certain pages were just randomly not archived? Or when you click a link and get a page from years earlier or later than the one you came from? Never encountered a page or an entire domain that was blocked from Wayback Machine? Why do you think I would get started doing something like this in the first place if I didn't find it more fun to browse my own archives than Somebody Else's?
https://en.wikipedia.org/wiki/WWWOFFLE
https://ftp.netbsd.org/pub/pkgsrc/distfiles/wwwoffle-2.9j.tg...
The way the www is going, it seems like downloading a copy of libgen, i.e., nonfiction books, and scimag, i.e., academic journals, via torrent, would be more valuable than archiving websites, in general. These primary sources are part of the material used to train so-called "AI" anyway. The problem is that this so-called "AI" also includes all the garbage from the www.
Worst case is eventually these books and journals will again become publicly inaccessible but "AI" will be offered as a bogus substitute; a future where few people will do research using primary materials anymore, they will just submit questions to a remote "AI" server. Truth will be decimated.
https://blusharkmedia.medium.com/the-ongoing-battle-against-...
https://techhq.com/2023/09/can-libgen-shadow-library-survive...
https://www.twitter.com/theshawwn/status/1320282152689336320
https://qz.com/openai-books-piracy-microsoft-meta-google-cha...
https://qz.com/shadow-libraries-are-at-the-heart-of-the-moun...
https://goodereader.com/blog/e-book-news/authors-file-lawsui...
When asked about whether this was true, they refused to answer based on confidentiality concerns, then said they had deleted all copies of the dataset, stopped using it, and no longer employed the individuals that compiled it:
https://www.businessinsider.com/openai-destroyed-ai-training...
We do know for a fact that the (non-OpenAI-controlled) "Books3" dataset is just "all of bibliotik":
https://www.twitter.com/theshawwn/status/1320282149329784833
https://github.com/soskek/bookcorpus/issues/27
And we also apparently know for a fact that this was included in the datasets used to train LLAMA:
https://en.wikipedia.org/wiki/The_Pile_(dataset)
https://aicopyright.substack.com/p/the-books-used-to-train-l...
https://aicopyright.substack.com/p/has-your-book-been-used-t...
https://news.ycombinator.com/item?id=40258584
https://arxiv.org/pdf/2005.14165.pdf
https://www.wired.com/story/battle-over-books3
https://www.washingtonpost.com/technology/interactive/2023/a...
https://www.theguardian.com/technology/2023/apr/20/fresh-con...
https://storage.courtlistener.com/recap/gov.uscourts.cand.41...
See 40-45.
https://storage.courtlistener.com/recap/gov.uscourts.nysd.60...
See 87-116.
I would love that. I have a little for parameter version, but I feel yours is more tried and true.
wget \ --recursive \ --mirror \ --timestamping \ --page-requisites \ --html-extension \ --convert-links \ --restrict-file-names=windows \ --no-parent \ $url
For the sake of argument (maybe not true), let's say that all techies are aware of archive.org, and consider it beneficial, probably using it themselves.
Why don't they instead demo against a target that will be proof of capability, and one that someone won't pay them to do (no freebies), yet one that they perceive as bad or deserving in some way?
Probably improper to suggest "better" targets here, but I really wonder what's going on when some relative do-gooder gets attacked.
Similarly, ransomware attack on a children's hospital, of all places? Doesn't that get you uninvited to criminal mastermind dinner parties?
As Omar of "The Wire" told us, a man's gotta have a code.
LockBit was so successful partly because they didn’t have to hack anyone themselves. It was basically something advertised “Got SSH or RDP access? Let’s make a bunch of money.”
This attracted hackers who might not trust themselves to do the extortion part safely, as well as people who didn’t actually hack anything but hated their boss, wanted a payday.
Perhaps they intentionally attack targets that are generally seen in a positive light, to prove to potential customers that morale is not an issue.
Oh, you want me to DDOS a children's hospital? No problem.
(Googling "Jason Scott TIA" gives me "Dr Jason Scott is a Senior Research Fellow in the Tasmanian Institute of Agriculture" which doesn't explain much to me)
TIA = The Internet Archive (i.e. the victim of the DDoS).
>The user you're responding to is Jason Scott of The Internet Archive
I am shocked that any HN reader could be ignorant of this fact. Their director is a (controversial) Turing Award winner.
This could even be as simple as "Some aspect of the attack pattern is inconsistent with such a motive", or "We spotted the perpetrator credibly gloating about it". But just from IA's public statements, the pattern ("launching tens of thousands of fake information requests per second") is quite consistent with simple denial.
And not taking on the job of police when you don't know as much as you think you do, such as the speakers whose speech you presume to police.
You might not know the significance of "textfiles says no", but you do know that in general it's a thing that on HN, sometimes the rando is no rando, and you do know that you're not dang.
That's all it takes to avoid looking like a douche. And soon enough some comment or other would fill in the significance, from someone else looking like a douche and having it explained to them.
Jason could have added "Internet Archive here, it's not that." But he would have to say that in front of every comment he ever writes, which I think would get old for him and probably no small number of other people would criticize that too "yesss we know you work for IA FFS get over yourself..."
I think it's fine for him to just speak and let everyone else take care of themselves.
Sometimes. More often than many other places we could mention. But in the majority of cases, even here, a rando is a rando. And if the randos see the accepted conventions being ignored without comment it might encourage them to do it more.
> Jason could have added
I'd argue should have.
> But he would have to say that in front of every comment he ever writes
Only comments where it is significantly relevant, or in this case where his experience and proximity to the issue at have might be considered enough to ignore the standard commenting conventions.
So in other words, anybody can carry out a DDOS for basically no cost. So trying to analyze the purpose, let alone suspects, is probably not going to be fruitful.
https://robindev.substack.com/p/cloudflare-took-down-our-web...
So in other words, Cloudflare noticed the author was running a gambling site, they decided that this was negatively impacting the shared IPs and the author would therefore need to upgrade to a plan that included BYOIP because they would need to use that feature to continue using Cloudflare and they likely insisted on prepayment for the annual plan because gambling sites have a reputation for being flaky and prepaying would have demonstrated the liquidity necessary to continue operating the site at that plan.
Again, Cloudflare could have communicated this better (and maybe they did in parts of the correspondence the author didn't share) but this all seems perfectly understandable, especially given how the sales team kept referencing Trust and Safety (implying the alternative is ending the contract for violating the ToS).
The issue of tainting shared IPs would indeed have suddenly gone away had the author brought their own IP (which would have required an Enterprise plan to do while staying on Cloudflare). Instead the author feigns ignorance arguing they don't even need the features of the Enterprise plan and doesn't acknowledge the issue with sharing IPs while sheepishly mentioning that maybe they're accidentally invading bans of their domain in certain countries by having alternative domains which they of course don't actually need because most traffic comes from their main domain yet somehow having these alternative domains is critical to running their business.
What are you even trying to argue here? The author is being deliberately dishonest in how they frame the incident and Cloudflare's motivation is perfectly understandable. The only thing to take offense with is the communication style which we can only judge based on a select few messages the author shows us. We have to rely on their word after they have already demonstrated dishonesty.
“We tried saying that we don't need any number of the 14 features that are included”
Which, to me, is the crux of the issue. Is it fair for Cloudflare to say “You are breaking the terms of service if you do not change your set up in this specific way, and also the way you need to chance your setup is locked behind a significantly more expensive pricing.” Being able to bring your own IP does not, to me, seem like something that should require a plan that is orders of magnitude more expensive than the standard. It seems much more to me like something that is more fundamental, and should be included as an option in a lesser version of the product Maybe I’m wrong, and there is actually significant overhead to Cloudflare for letting customers bring an IP. But as is, it feels very much to me like a situation where something vital was locked at the most expensive tier to force certain kinds of customer to pay more.
Yes, BYOIP as a feature does not seem complex enough to warrant paying for an Entperise license. But the kind of customers who need BYOIP (especially if they need it to avoid harming your IP reputation) are likely to be at a higher risk of being flaky or otherwise painful so this is very much a tax on running that kind of business (just as porn sites often find it hard to find payment processors because of the high risk of credit card fraud).
As a freelancer I have absolutely made offers at 10x my going rate for client I did not want. The idea is that if they really want me to work for them, at least I get reimbursed for the suffering that will entail. This kind of pricing structure is no different.
Zero ingress puts the upfront bandwidth cost onto the attacker. Because... you actually may succeed to defend and stay up. Their success is not guaranteed, they might be shouting into the void.
Attack success (as in, "impact on you") is guaranteed if your ingress is chargeable.
This is a sampling of currently available services and who they use for DDoS protection:
stresslab.app - Cloudflare
maxstresser.com - Cloudflare
sunnystress.com - Cloudflare
tresser.io - Cloudflare
ip-stresser.net - Cloudflare
hardstresser.com - DDoSGuard
zdstresser.net - Cloudflare
starkstresser.net - Cloudflare
stresserhub.org - Cloudflare
nightmarestresser.net - DDoSGuard
Just for fun head over to Cloudflare's abuse reporting site and try to figure out how to get one of these taken down. https://abuse.cloudflare.com/Now they hide behind Cloudflare who will refuse to turn over any information so that security folks can get them taken down. Unfortunately Cloudflare has grown too large that we can't just block all of it or depeer them like we would any other network that provided services to bad actors.
Most of the listed domain names are under US jurisdiction. That means the authorities should be able to take them down. If Cloudflare is found to have been knowingly enabling crime, it could face fines, and the CEO and other key people could end up in prison. The Cloudflare services have probably been paid using means that are under US jurisdiction. Those payment accounts can be closed and the people behind them tracked down and potentially charged with crimes.
Or at least that's how things work in the real world. The internet is still apparently too new for the authorities to understand how to deal with it.
Npm has been under pretty severe attack for ~6 weeks now. I forget who else.
The scariest thing to me is what we might do in the face of persistent online attacks. If this stuff gets rolled up into western nations rolling back privacy & liberty? That's an theonion.com "bin laden plan to sit back and enjoy collapse" situation. Freak out & let cyber security paranoia reign & destroy free communication & connection.
Given its benefit to the lay persons I recommend everyone who use their services give a small amount once for a while. I already did so but if not for family issues I'd donate way more.
To combat this you need to buy enough pipes to the internet for your regular internet traffic, as well as an extra 500 Gbps or so. That is a lot of unused bandwidth to be paying for every month. Then once the packets arrive at your datacenter you still need to buy dedicated appliances to scrub out the bad and let the good flow.
Google is constantly under attack, but their normal daily traffic volume (multiple Tbps) is large enough that just the extra capacity they keep on hand to deal with traffic spikes the World Cup or a popular YouTube video is larger than what most attackers can muster.
Cloudflare describes that policy as a commitment to content neutrality rather than extortion, and I think that's more or less sincere (since they've protected many other unpopular sites that didn't give them such a benefit, with a few high-profile exceptions). It does work out very conveniently for them, though.
But we know that's not true. Point out problems with a very controversial blogger and they'll cancel your service.
Because they are a publicly traded company
There are open source tools to mitigate DDoS, but all of them will have some marginal cost to run, and they will all be significantly worse than Cloudflare as they benefit from neither Cloudflare's data moat or scale.
Secondly, considering Cloudflare would MITM all traffic, it would make a data good source for the NSA, thereby violating all user privacy.
This seems like a weak argument. Should we just take down anything widely accessed because it might be used by the NSA? What about AWS?
This had a high risk of getting Cloudflare's limited ipv4 addresses to a blacklist - affecting ALL of their customers.
All CF did was ask them to switch to an Enterprise plan and bring their own IP-addresses. They refused to do either and rather cried on the internet claiming CF to be bullies. It's not like the price they asked was even a fraction of the profits an online casino brings in every DAY.
If a customer's action is truly illegal, the customer's account should be terminated, either immediately or after a reasonable warning to fix things. Under no circumstances should money be sought to support the illegal activity.
CF is engaged in a pig-butchering scam whereby they lure customers, then when the customer is all fat and happy, they get asked to pay up or lose their business.
In this case, CF destroyed the customer's business as soon as CF got word that the customer was going to move to Fastly, considering that the protection money was not paid. It's an open and shut case of racketeering.
Nice.
Secondly, considering Cloudflare would MITM all traffic, it would make a data good source for the NSA, thereby violating all user privacy.
Or any proof of work proxy that delays the ingress traffic. If you only have one server there is very little you can do except maybe redirect to a static page or kill the DNS entries.
Even if you go all out and buy a bunch of huge IP transit links, you are not gonna be able to stop the IXP 800 miles away from getting congested and blocking your customers from accessing your site anyways. You need access to a backbone to route traffic differently to avoid those kinds of issues, which is why DDoS scrubbing services will partner with a T1 ISP to do most of the work.
It is the same problem with email spam. What's stopping someone from sending billions of spam mails?
If we suppose that: a blockchain exists which is fast enough, cheap enough, and spread out enough on the globe (to mitigate latency), then there is no reason, for a tcp packet to not carry with it a small money transaction, in the order of a millionth of a cent. Information gets served back, only when the transaction is confirmed.
In that way, any request with no transaction gets discarded, and only requests with a small cost pass through. Suddenly by sending requests one after another and no end in sight, DDoS attacks and mail spam start to cost money. It is the serving of request that makes DDoS attacks and mail spam to be effective.
The problem however, is that no blockchain is fast enough and cheap enough as of today. But there will be one in a handful of years.
(also, you're arguably just moving the problem to DDoSing the payment processing / firewall mechanism)
These ideas indeed exist for decades.
> The big problem isn't a technological one, it's that it presupposes some "sweet spot" price (negligible for legitimate users, yet prohibitive for abusers) that has never been shown to exist in reality.
Advances in technology, software and hardware, make it easier and easier for that sweet spot to exist. That sweet spot, didn't exist in the past, certainly, but we are close right now.
One example that i think is useful here, is aluminum cans for fizzy drinks. Aluminum, a strong metal compared to cardboard or plastic or glass, is better at withstanding pressurized gases without exploding. The downside, is that it's more expensive. When manufacturing prices dropped down a lot, then it was feasible to drink half a liter of liquid and just throw away the metal. Aluminum still not free though, but the small price did worth it. Huge waste of energy as well to smelt all that metal and throw it away after 10 minutes of drinking, but it is economically viable.
One could manufacture Titanium cans, and drink even more fizzy drinks. But that's not economically viable as of today.
> you're arguably just moving the problem to DDoSing the payment processing / firewall mechanism.
Yes, the problem is moved elsewhere, that's the weak link in the scheme i described. The thing is that a flood of transactions still costs money. Blockchains cannot be flooded just with requests, they have to be flooded by transactions. Take a look at the article [1] which outlines some ideas. I don't agree with a lot of things in there, but it states the problem and gives some numbers.
The theory when it comes to blockchain deterring DDoS attacks (and other kind of attacks), is that there are not bad guys in general, just rational economic actors who use dirty tricks. When a dirty trick starts to cost money, and profit disappears from an attack, then the rational economic actor will stop the attack. The bad guy will resume the attack regardless of profit, but that's one of the axioms of the theory, that there are no bad guys.
[1] https://www.dlnews.com/articles/defi/ddos-attacks-are-an-inc...
If you need PoW for every connection, it's going to end up being very expensive for an attacker to saturate the connection. And the captcha is probably a different server to the main site.
I thought a gateway would work for that purpose.
Are you sure it wouldn't? In that case there's no defending against ddos.
The filtering/dropping of packets has to be upstream of your connection to be able to protect your connection - additionally somewhere where the available bandwidth is greater than the attacker's bandwidth.
"Make em offer they can't refuse."
https://mastodon.archive.org/@brewsterkahle/1125141764988452...
@internetarchive is back up!
this is a back-and-forth with attackers. Sees weekends and holidays are popular.
we made adjustments, but we will see.
at least Happy Memorial Day!looks a bit more broad than i wanted. as is its like a thin wrapper for IPFS.
including all media conglomerates (obviously) and all scientific, literary, etc, publishing houses.
also, there's a global war, so it well may be a fog-of-war technique or like somebody else also mentioned: someone needs something to stay quite for a little bit as part of some larger operation
I think that's also the wrong question to ask. "Who's doing it?" is less interesting than "What's enabling them to succeed?"
Lots of people want to rewrite or erase history.
Quoting a story I wrote about this a few years ago:
"Everything you speak, all ideas, all things, all thoughts, they are all of the past. Society and knowledge is a composite of the shadows of former presents.
When people lie or misrepresent knowledge they speak of a past they wish to change.
What if people who have the most to gain from deceit had a tool to actually change the past and make these lies the truth?"
Here it is if you're curious https://kristopolous.medium.com/stephen-hawking-had-a-time-t...
This is indeed very tough to resist.
oh wait
This is really bad for CC media.
Internet Archive is now working for me
Example:
https://wayback.archive-it.org/all/20240506083041/https://ar...