A million ways to die on the web
wiki.archiveteam.org
wiki.archiveteam.org
Of course redeploying every single thing would not be seamless because of course, there might be some configuration stored in services, or something similar, but I'd say that ~90% of our automation is stored in Git.
maybe they could even use a relatively inexpensive colo/baremetal provider to simply mirror the bigtech deployment on a smaller scale (would need to be quite flexible/vendor-agnostic to make that work...)
The main problem is that would outgrow single hard drive so would need NAS. Also, the transfer speed could be an issue as database gets bigger. Even if don't store all customer data, it does make sense to store all the configuration, keys, and secrets.
If anything, it increases the need for 3-2-1 backups: the original copy of all of your files are on somebody else's computer that you have no control over. Hopefully they're keeping it backed up, and hopefully they don't go belly up and pull the plug all of a sudden. So you can use a primary backup in another cloud service from another company that hopefully won't kill their product at the same time as the other one (again, you have very little knowledge or control of the way they run their data center). Ultimately, it's a good idea to have a copy of your data that you have control over, maybe in a big drive (or set of drives, tapes, etc) in the safe, rotated daily/weekly/however long your company can cope with losing in a major SHTF situation.
Excessive? Maybe. For what it's worth my shop is locally hosted with both local and cloud backups. I have never regretted having at least one backup of anything and it's saved my bacon (or my coworkers', boss', etc.) a number of times. I've been fortunate to never need to rely on a secondary backup, but I sure wouldn't bet the company on it.
As Fred Brooks said in Mythical Man-Month: Show me your flowcharts and conceal your tables, and I shall continue to be mystified. Show me your tables, and I won’t usually need your flowcharts; they’ll be obvious.
At their scale lots of their stuff is custom and needs to be ordered at least 18 months in advance.
The fact that they can do capacity planning 2/3 years in advance and have very limited misses in a way that people are astonished that they have capacity misses is a testament to how good they ate at it.
'We had a lot of discussions along the lines of, 'Yeah, that's a great idea, but we can't just autoscale out of bad config or planning' or 'We can't just reboot a host and get a new one, we have to plan to take care of the ones we have'
It's been pretty fun
https://siliconangle.com/2011/08/01/third-largest-bitcoin-ex...
(Valued at 220k USD then, or 700M USD today)
Reading this from 2011, and reading the latest on Web3IsGoingGreat.com today, is vaguely staggering. They never have learned.
>Altern.org is a free web hosting service created in 1992 by Valentin Lacambre and disappeared in 2000. From its origins to the closure, Valentin Lacambre, a pioneer of Free Internet in France, had to permanently close the free hosting service in early July 2000 following numerous lawsuits. This closure was due to the laws of the time, which placed the delicate obligation on hosts to act as judge, censor, and, by default, guilty, as it was deemed difficult and contrary to his principles to control the 21,893 sites that existed on Altern.org at the time of closure.
http://yavista.com/98/1f/981fa5fe.html
https://fr.wikipedia.org/wiki/Altern
(Valentin Lacambre went on to be a co-founder of Gandi.net.)
Gandi.net which recently (last year) was purchased by some other corporate entity and subsequently raised all the prices + made previously free features paid. I only mention this as I discovered this week that my Gandi bill suddenly got a lot bigger.
> The acquisition was completed on June 23, 2023.
They can use their influence on the people who having voting rights, to actually push the hot-potato to public investors.
That's one big reason why you public companies purchasing somewhat useless companies, or acquihiring them at insane valuation.
I was a paying subscriber to a great little site called magweb ~20 years ago. They scanned old and new (military) history and (war) game magazines and posted HTML versions, with permission of original publishers. It was really nice and explicitly allowed users to print or save copies of articles. No DRM. Perfect example of how a site for reading magazines online should work. Everything just basic HTML that even worked great to read in phone browsers 20 years ago.
Then Hurricane Katrina hit and the server that was apparently running out of the owner's basement somewhere in that area was flooded, with no working backups. I still have not found any traces of most of the 40000+ articles, other than the few I had saved, that used to be on that site. Since it was paywalled only a tiny part of the site is available on the wayback machine.
https://web.archive.org/web/20050529083811/http://www.magweb...
(No, I can't imagine running a business and scanning magazines for 9+ years and not make sure to have backups of everything.)
https://www.datacenterdynamics.com/en/news/ovhcloud-ordered-...
But make sure you check where the other provider is hosted: I read it one user (individual, not a startup or other company) who had backups of stuff on a shared host with a second shared host. It turned out that they were both running on servers in that one OVH DC and they were cheap hosts with no better backup/dr plan themselves...
Since 1999, New York City’s Office of Emergency Management, charged with coordinating all aspects of the response, had occupied permanent headquarters in Seven World Trade Center, on Greenwich Street, just north of the landmark twin towers. A vital communications link was the radio repeater system based on the ground floor of One World Trade Center, the north tower. The loss of those facilities – and key personnel working there – significantly hampered the response.
<https://theconversation.com/disaster-communications-lessons-...>
<https://www.computerworld.com/article/2510996/9-11--top-less...>
I seem to recall an instance (I thought it was the NYC OEM, may have been another) in which both primary and secondary data archives were both within the WTC complex. Best practices now are to locate secondary / backup services at least 100--200 km from the primary. Preferably within different watersheds, seismic regions, etc.
Widespread disasters are comparatively rare, but can affect considerable areas, and locations sufficiently proximate might well be affected.
Renting names that serve as resource identifiers, locators and trademarks all at the same time is just not a good idea.
Habbo Hotel: It was marked that sexual predators were using the service to groom and instead of tackling the issue, they applied a worldwide mute. You could walk but not talk.
Habbo still exists, but almost killed the whole fanbase. https://en.wikipedia.org/wiki/Habbo#Moderation
And that, you required Shockwave.
0: https://github.com/gauteh/lieer
1: https://notmuch.readthedocs.io/en/latest/man1/notmuch.html
For security you should assume someone recorded it indefinitely.
For archival you should assume nobody recorded it including the original creator.
Best source I seem to find is an HN posting from 4 months ago:
<https://news.ycombinator.com/item?id=37295238>
AzureDiamond/hunter2: RIP
On a side note, the current form of digg is interesting, yet somehow poorly done. Go there, sort by year: get only things from 2024, since no way to get 2023 or a rolling year. Not to mention being able to choose all or certain time periods
I started going there regularly about a year and a half ago, mostly because they would feature articles that I wouldn't find in my other typical websites. But in the last six months they have been optimizing to death, cutting anything that's not an instant success and publishing variant after variant of anything that's mildly successful (yay, another article on Twitter memes!). Clickbait titles are there too, along with a new-ish comment section that's 90% spam.
It's been a sobering lesson on what happens when you put growth above everything.
In 2019 Digg posted a job listing for a links curator, and I cheekily applied, noting that I'm already doing the job anyway, so they might as well pay me for it. They didn't take me up on it, but like magic, the poaching went away.
I also have a small group of well-read friends who make an effort to send me stuff, that helps a lot too.
This is not mentioned here https://wiki.archiveteam.org/index.php/Tripod because I think the event precedes Archive Teams formation.
It's a sad list. Even Google, after it acquired YouTube, forced me to change my YouTube login to a gmail account.
An example is BBSpot.com. It's still up and going, but very different from many years ago.
I miss SatireWire. I think it's dead now (ERROR ESTABLISHING A DATABASE CONNECTION).
I still have a CD-ROM from them (back then, everybody and his brother published a CD-ROM).
https://www.lighthouse.storage/
To me, they seem like the most useful stuff coming out of the blockchain industry.
The political environment the site operates in turns hostile to website content or method of publishing. The operators face costs of compliance, loss of scope, or personal risk in continuing.
See: The UK Online Safety Bill [1] [2] and especially [3] "Ofcom's >1,500 page consultation on the Online Safety Act 2023, and why small companies don't have a chance"
[1] https://www.theverge.com/2023/10/26/23922397/uk-online-safet...
[2] https://www.eff.org/deeplinks/2023/09/uk-government-knows-ho...
[3] https://decoded.legal/blog/2023/11/ofcoms-%3E1500-page-consu...
GitHub: https://github.com/linkwarden/linkwarden
Website: https://linkwarden.app