GitHub Archive Program
archiveprogram.github.com
archiveprogram.github.com
Someone will write a GNU Emacs mode for that
Better check in your node_modules directory the night before archive day so it’s included.
> The snapshot will consist of the HEAD of the default branch of each repository, minus any binaries larger than 100KB in size. Each repository will be packaged as a single TAR file.
100KB is a very low limit, many repos will be useless
I wouldn't call that a problem.
CockroachDB
The thing is, it's really impossible to predict what information will and will not be a critical key to unlocking understanding of future generations. Just keeping it all, history, comments, and all, will be a huge boon to future historiographers trying to figure out what developing in the early 21st century was like.
I can see the argument for being forceful about prompting you to write down backup codes or whatever, but fixing that after-the-fact is something they should absolutely not been doing.
I do think that sites should offer better options for recovery. NearlyFreeSpeech do a really good job of this, offering seven methods of recovery and letting you decide how many you need to fulfil to be given access and which you want to configure. However, things like checking photo ID and more "offline" options are expensive to support, so I get why that is rare.
> However, things like checking photo ID and more "offline" options are expensive to support, so I get why that is rare.
For you to characterize my efforts with Github support as "talking my way past it" is absurd. I was never once asked to provide a photo ID, etc. something I would have gladly complied with.
The entire point of 2FA is that it's a second factor and no way around it without a second factor that is verifiable without doubt.
The most relevant issue for me being the lack of U2F (now WebAuthn, I guess) support. It is also really annoying you can't have restricted SSH keys to allow for automation that is locked down to single sites.
They have had those issues on their feature voting for years.
After that experience I switched to Authy, which does an encrypted cloud backup of your TOTP secrets.
No, it just means you use a different 2FA. Here's some examples.
- Gov ID
- Pushing to a private repo
- Checking if IP matches historical records (hell, since they're fingerprinting everyone, why not use that?)
There's problems with 2FA. Mainly that it is on your phone. Phones are pretty valuable devices and it isn't unlikely that they get stolen. Even if Google backed up Authenticator to Drive, how do you get Authenticator back on your new phone?
It isn't hard to come up with a hundred scenarios where things fall apart. This is why there are different levels of security. But my GitHub isn't so sensitive that if I lose access I don't want anyone to ever get access. Yet at the same time it is sensitive enough (and GitHub is attacked enough) that I don't think just a user name and password is sufficient. Where's the middle ground? Even with YubiKeys I have to have multiples (as a backup in case I lose one). There's such a thing as "acceptable security levels."
We're trying to make things easier and safer for humans. But that doesn't mean they aren't still going to be human.
Gandi handles this nicely with a checkbox in the settings that explicitly tells support to never ever restore the account if 2FA is lost
The main problem is with targeted attacks. That's why SMS not secure
> This is why there are different levels of security. But my GitHub isn't so sensitive that if I lose access I don't want anyone to ever get access. Yet at the same time it is sensitive enough (and GitHub is attacked enough) that I don't think just a user name and password is sufficient.
Yes, there is stuff people would rather have destroyed than get into the wrong hands. But I personally can't think of any. But I do think we all agree that a password alone is not secure enough. What I'm saying is that there shouldn't be two choices. There should be middle grounds. SMS is not a great recovery tool for 2FA because it is trivial to fake. Specifically with my GitHub I'm okay with my security not being invulnerable to nation state actors. But I don't want it to be hacked because of a dictionary attack with some modifications. If someone has a copy of my passport then there's much bigger problems that I have than my GitHub being hacked.
The key part here is that there should be varying levels of security. The two levels we have (no 2FA vs 2FA) is too small.
* as to the IP: I was suggesting that it was being used in combination with other stuff. Not a standalone verification. Just like your phone number should never be a standalone verification. We're talking multifactor.
GitHub is a social site, and the users accessing repositories are just as important as those that own them. You might not care if your account is compromised, but if other people trust your repoisitories and they get attacked through your compromised account, that is a problem.
Sounds reasonable. What's important is that you can securely recover the account at all, not that it's cheap.
If you still have the keys, try:
ssh -i ~/.ssh/your_linked_privkey -T git@github.com verify
Edit: I previously wrote to pass your public key to the ssh client rather than your private key. Of course, that was incorrect.
[1] https://help.github.com/en/github/authenticating-to-github/d...
Heck, even in this very scenario, if I haven't used an SSH key with GitHub in many years, and then GitHub receives an artifact signed with that key saying "I lost my 2FA token and backup codes, please reset account auth so I can log back in", I very much do not want GitHub to trust that artifact. If I haven't used the key in years, that probably means I don't have it anymore and either never got around to removing it from GitHub or forgot it was there.
Of course, someone might still have removed those keys. IDK.
The entire point of 2FA is to avoid someone taking over your email and then being able to access _anything_ tied to that email.
This is probably the most popular current dogma in security circles. This is a policy issue not an intractable mathematical problem. In theory maybe it could be as rigorous as math but in practice security is always relative.
The same can be said for strong encryption, where losing the keys means you have no chance of recovery. I'm not an advocate of defaulting to things like full-disk-encryption for that reason (and know people who lost a lot due to that.) I guess the underlying problem is "humans are fallible, strong security is not", and that risks "in the other direction" aren't often mentioned; would you rather have the risk of being hacked, or of losing access to your data forever?
You enabled 2fa, you lost your 2fa, and you did not have any recovery codes. Now you are asking for them to bypass the 2fa, and they are refusing.
Again, that sucks, but when I compare this to what the cell phone companies are doing with sim swapping, it increases my respect for Github.
Isn't one of the biggest selling points of 2fa that if one account gets hacked you don't lose all accounts?
In this case, an email account is the second factor. Your Github password is the first factor.
If Github is doing something different than most websites, that needs to be very very clearly communicated at setup time, otherwise people will assume Github behaves like other websites.
> Warning: For security reasons, GitHub Support may not be able to restore access to accounts with two-factor authentication enabled if you lose your two-factor authentication credentials or lose access to your account recovery methods.
https://help.github.com/en/github/authenticating-to-github/r...
>Treat your recovery codes with the same level of attention as you would your password!
Most people forget passwords all the time and treat that as no big deal because they can just reset the password. If people use that "same level of attention" for the recovery codes, then they will be lost just as easily, except this time people won't be able to recover.
Edit:
After enabling I got an email with a much more cautionary tone:
> Recovery codes are the only way to access your account again.
> GitHub Support will not be able to restore access to your account.
It's interesting they call this out in the email but not during the sign up process.
Given how we all have placed GitHub front and center of our lives, there is no excuse for any of this. Even if I'd no recovery codes, I would expect GitHub had some process to get back access back. GitHub account should be looked upon with same reverence as bank account. The idea that you need to forget about money you have in bank account because you lose your phone and recovery code is mind numbing. They could charge you $500, run professional background check on you, get your credit card/bank info and verify it's really you, wait for 30 days, remove private repos and then grant you access again at least for your public repos. That's what I would expect an efficient customer obsessed organization to do.
2FA with SMS has the problem that companies offer support by human and don't have a tight system for changing numbers to another phone, as proven again and again, resulting in accounts being compromised.
While the current solution of MFA aren't perfect, it's hard to come up with other solution that would be as safe or safer and prevent most to all mechanisms used to compromise accounts, like phishing, social engineering and other possible remote attacks. Giving you the possibility to save the codes somewhere physical has its downsides, but an important upside is that it allows _you_ to keep in charge your own security in most cases.
- You log with your 2FA into a device. - You set this device as trusted. - In case you lose your phone or 2FA device, log onto the trusted device and disable 2FA. - Then set up 2FA again with new phone or device.
It's not perfect, but it's workable. For instance, while I won't enable trusted device on my laptop, having my desktop stolen is a way rarer occasion, so I enable the "trust this device" options. It's just a matter of thinking on the threat model and where you can place spots for recovery while loosing as little security as possible.
You can have multiple second factors with Github. I currently have two Yubikeys and one authenticator app enabled. If I lose one I can still log in with another.
If a human can’t give me my account back through tech support I’m not very keen on trusting my account a gadget that can break or get lost.
The risk of losing a phone and the backup codes is probably several orders of magnitude larger than the risk of being the target of a sim swap attack for the vast majority of users.
What difference does it make unless everyone you trust is gone or has lost everything? At that point you have larger problems than logging into online accounts.
>≥ without having to have them present at every registration
For example, I have given a token to a family member in another country, for proper utility I need that token back each time I register on another site..
For one, no one is forcing you to only have one TOTP device. You can scan that QR code as many times as you want. Have them on multiple devices.
Depending on your threat vectors, putting them into a password manager that supports it (like Bitwarden) might also be smart. Less secure than fully offline, but definitely better than SMS.
As for the backup codes - one big encrypted text file synced to the cloud of your choice should do the trick, but if you prefer the "scary men with guns" kind of security, safety deposit boxes were literally made to store this kind of stuff (bonus points for on-paper encryption).
Cite: https://www.eff.org/deeplinks/2017/09/guide-common-types-two...
As an extra suggestion: if you use an Android phone for OTP, [andOTP](https://github.com/andOTP/andOTP) supports exporting directly into a PGP-encrypted JSON file which can then be either imported back into the app or converted back to QR codes with a script.
Since it allows you to trigger the export using a Broadcast Intent, I have it set up to do that as a part of my weekly backup Tasker script (of course, you could also just use any other sync solution and manually export when you add a new code).
Now instead of accepting the result is the consequence of your own poor choices you are trying to shift the blame to GitHub.
You clicked it, failed to know the ramifications of 2FA (which GitHub does spell out), and didn't secure your backup codes.
Take some responsibility instead of blame shifting.
Nevertheless the anxiety of losing the physical device with all my 2FA logins is what prevented me from enabling 2FA on most of my accounts until I was referred to Authy (authy.com) where you can sync your 2FA across multiple devices including your PC, which other than being very convenient, the effortless syncing + redundancy gave me confidence to enable 2FA on all my accounts as the redundancy ensures I'll still be able to access my accounts if one of my devices is broken/lost.
You may have forgotten that it doesn't matter how you store your password, but the problem is that it is a single factor. Once compromised, one can gain access to anything within that account. You may be compromised by phishing, keylogging or other means. 2FA can help with making these types of attacks more difficult, although not impossible.
Luckily I only had public repos, so I created a new account and forked them all. Support had told me that if there isn't any login activity in the blocked account for a period of six months, they would delete the account and release the username. And yes, I had to follow up after six months were up to make that happen.
That just sounds like a fun project.
But also a great engineering exercise. I wouldn't be surprised if this exercise leads to lots of valuable improvements for GitHub users in the here and now. Trying to solve such a grand challenge forces you to develop a vocabulary and understanding of your current systems that can lead to more immediate improvements. I think it unlikely any of these archives will actually be accessed, but simply building them could lead to great side effects.
It's also great marketing, as I now believe that Microsoft/GitHub takes the job of not losing user data extremely seriously, more so than if they had spent an equivalent some of money buying an ad that says "We take not losing data seriously".
www.longnow.org
The second question is what will enter public domain first, business-owned software or individually-written software. If the latter, which open-source developer has died the earliest?
Can we just take a second to appreciate how downright inhumane copyright expiry is? It basically encourages cheering for people to die because of copyright expiry.
https://github.com/eeeeeeeeeeeeeeeeeeeeeeeeeeeeeeee/eeeeeeee...
The archive team should be able to get a very good compression ratio on that repo.
So, "share only code no-one cares about and don't work on it much" seems to also be an option.
There’s a big difference between “public” and “stored in a glacier forever”.
So you can feel better?
So there's a lot of reason to believe the 2020 snapshot won't also be online.
My intuition is that archiving data for long term historical use is different from datamining a [meta]data to maximize invasion of privacy. Also there is a difference in accessibility, stored inside a glacier very few people are going to actually read it.
I believe that if mass complaints from all over the EU emerged it would be a different story. But this does not look like the activities the GDPR was created for
You can't exactly GDPR request deletion of your information from a printed book, so I'm curious how GDPR applies to such physical archival mediums.
It's more akin to "hiding" - to me, deleted means unrecoverable, by anyone, at any time.
Even your OS doesn't "delete" files, until the actual sector on the drive is overwritten (depending on media used, of course).
You probably meant this as a rhetorical question, but I'd argue that yes, (for public available data at least) it probably should be. It'd enable solutions to a lot of problems we have with the current web, not least archival and broken links.
Yes, probably
I'm going to go against the HN zeitgeist and say no.
If I have the right to publish something to the web, I should also have the right to edit and delete it if I so choose.
"What happens on the Internet stays on the Internet. Forever."
> As today’s vital code becomes yesterday’s historical curiosity
shouldn't be "tomorrow's historical curiosity"?
Also archeologists study rot.
Why not use a much more accessible medium like M-DISC, which claims a lifespan of at least 1,000 years? https://en.wikipedia.org/wiki/M-DISC
just wondering how an encoded digital format is more accessible than something printed on a transparent film that you can see by looking through it.
I bet people in 2200 can't wait to figure out what version of node they need to run a specific gulp version from an obscure error, caused by an underlying dependency's reliance on an deprecated built-in (true story).
After figuring out that issue, they will need to debug why a specific package's version is broken, only to find out the maintainers 200 years ago hard-coded the century.
Plus, people are going to reinvent the wheel no matter what.
I think that’s what GP means, that if you don’t include the “context” (interpreter, cpu architecture, power supply), you’re only burying the most malleable tip of the iceberg.
It reminds me of the Zork source code that was published here some months ago. It was written in a language for which the compiler appears to have been lost. People have tried to develop compilers given the source code provided, but we can’t get it to work just right, because some crucial bit of context appears to be missing.
There will probably be people who maintain LISP and C for decades more, but imagine trying to write a Java or Haskell compiler with only the source code for Kafka or Pandoc as your guide?
Not impossible of course, but hey, maybe you should check the source code for a JVM or ghc version (and maybe add a Linux distro and gcc for good measure) in with your source in time for the cut off date. :)
Point is, once you enumerate all of the dependencies, you realize the only solution is what GP said. Constant maintenance at all levels is the only thing that keeps the ship running.
I don't think anyone is going to be hoping to use this code for their own purposes, but rather to study programmer culture and development. For instance, when did the idea of "immutable by default" start, and how did it spread from esoteric academic languages into mainstream languages like Rust? How did the decades-long specter of Perl 6 affect the course of Perl 5, and how did the rename to Raku lead to its resurgence (or lack thereof)?
I'm no expert in history, but I have read a reasonable amount; and what's often fascinating is when there are two groups who think that they are in complete disagreement, but in fact share a common assumption which you, from the perspective hundreds of years in the future, do not share.
There are things that are so obvious to us now that we don't bother writing them down or explaining them, which will be a complete mystery to people 1000 years hence; and it's basically impossible to imagine what those might be.
At face value, this seems like a really odd choice. I don't understand why you would choose QR encoding unless this was being printed. I feel like I'm missing something here.
It is. QR codes printed onto microfilm, essentially.
It seems prudent. We may be entering a period of global upheaval and panic where large parts of human civilization will not make it past the next 100 years. It makes sense to attempt to preserve what we've accomplished in a way that will last for 1000s of years.
"Si vis pacem, para bellum"
Its very important to protect the civilizatory process from fallout, as severe or black-swan conditions might swipe what we have acomplished so far, and history is full of examples of advanced civilizations that were vanished and we had to dig out from the mud slowly without a chance to learn, starting all over again from scratch.
The civilizatory process is fragile and is always menaced by all sides, all the time. Its good to be always vigilant and prepared to anything.
What is the issue of archiving like this may I ask.
It's like these strange place full of dead elephant in africa.
Observe: resultant web page is unusable.
(I am participating here solely out of curiosity about the web; am not trying to make any point or argue any position.)
We know plain text will be around.
I was surprised to find the website promoting this is tricked-out, which isn't consistent with very long-term vision of the project.
Chris Pedrick's Web Developer Toolbar, an extension that was once the greatest thing, does this and many other useful and insightful things.
My fave, I think, is Information -> View Document Outline, which is pivotal for gaming Google SEO, by telling you how your Heading tags semantically convey the page's story.
uMatrix can disable CSS for any source, including 1st-order references.
There are extensions which will disable CSS as well.
Alternatively, you can view under a console mode browser (lynx, links, elinks, w3m, etc.), and see what appears.
So the acid test is, have the people implementing this really thought this through?
Based on the web pages they are using, the answer appears to be, no.
Complexity is the enemy of reliability. A principle first stated in 1958 in The Economist.
https://books.google.com/books?id=aDsiAQAAMAAJ&dq="complexit...
I would consider it pretty obvious that the permanence of usability/readability for archiveprogram.github.com is orthogonal to the actual goal of this project. Do you think future generations will learn about the archive from archiveprogram.github.com?