GitHub buries 21 TB open-source code in Arctic vault for 1k years
github.blog
github.blog
Worth considering swapping the submitted link out for the original source, per the HN guidelines. I would suggest the original blog post: https://github.blog/2020-07-16-github-archive-program-the-jo....
Still, the main question remains. Computer programs are extremely volatile by nature, they are written for the constantly changing media, and need huge amounts of context to be useful for future generations. Will a snapshot of the Linux kernel be useful in 100 years? 1000 years? After some global disaster? Even if you don't intend to run it (which requires archiving the current state of the supported platforms and much more), how much additional work will be needed to simply read it for studying?
>After the Challenger disaster, a hunt for the blueprints of the abandoned Saturn V rocket ensued. They were largely recovered, thanks to the work of archivists.
Yet nobody can build another Saturn V today with these blueprints. Any complex product is much more than just a set of blueprints. It's an organized process that depends on a certain team of people, supply chain, existing infrastructure, and much much more. Useful products only exist when they are supported.
This is so true. It is not easy to replace humans. Especially if it included a group of experts who all had in depth knowledge of not just the technology but also the process involved. It is not easy to document everything and we have to accept the fact that somethings are lost in translation. While documenting we typically assume that some of it is common sense and omit explaining the steps/process in depth. But over time, things change/evolve and what was common sense back in the day is lost due to passage of time.
Archived for a few hundred years that could be pretty interesting for future historians.
Even with how digitized everything is, there is definitely a lot of things that are undocumented (at least consistently in one place) or left out as too obvious.
We still have some of these same issues today, but given the rate of change in technology, we’re talking years instead of centuries or millennia. What was the average life like for a startup employee during the Dotcom era (1995-2001)? I have no idea because I was too young to experience it and I’ve never been able to find any books or accounts of the average daily work life even 20 years ago.
The problem is that history books/podcasts/shows are not intended to reveal that level of granularity about life. They exist to provide a summary within a specific set of boundaries (printed pages, run time, etc...).
If you want to know the nitty-gritty, you need to start with the source material. There's more source material on daily life available in the world's museums and university archives than you could ever digest in a lifetime. That's why it's summarized in the above-mentioned forms.
It’s not that the medium isn’t granular enough, it’s legitimately that no one wrote this stuff down so it’s just gone forever. There is no source material to start with.
The code itself is a manifestation of human creative expression and part of our cultural heritage even without executing it.
> Does it also contains OS, dependencies, compilers, linkers etc?
If they are free/open-source software (in this specific case on GitHub), then yes. See also their partner https://www.softwareheritage.org/mission/ which archives more than just GitHub.
You need to know how tar, lzma, utf-8 and git work. Judging by how hard it is to decipher ancient texts, I'd be surprised if anybody was able to read code without additional information on these subjects. Especially storing the binary git repo as-is will make it very hard.
One of the fascinating parts of this, to me, was this whole thing was about permanence... and at the same time, Nat's home burnt down.
I now have permanent, inaccessible backup of my code. Also wild.
This sort of seems cynical on Google's part. The right to be forgotten has just been irreversibly denied to some of us.
I know for a fact that a recently deceased friend wanted his repo's nuked. I wonder if this is going to be taken into account somehow. One can assume that laser can be used to burn nulls where needed, too ..
I hope in the year 3020 they have invented the advanced technology that allows a company like Github to sort their repositories by name.
And because ads.
Short video about the project:
My source code is worthless without it.
what about water chips?
I'm currently indexing thousands of GitHub repos for analytics reasons, and if you don't play nice (fetch infrequently), GitHub will throttle/terminate your connection.
You just need to use one of the many scraping proxy services which will give you a HTTP proxy that cycles through millions of IP addresses. It's impossible to block since they are largely residential IPs controlled by malware.
It definitely works but you are encouraging them which isn't a nice feeling.