GitHub Source Code Leak
resynth1943.net
resynth1943.net
GitHub hasn't been hacked. We accidentally shipped an un-stripped/obfuscated tarball of our GitHub Enterprise Server source code to some customers a couple of months ago. It shares code with github.com. As others have pointed out, much of GitHub is written in Ruby.
Git makes it trivial to impersonate unsigned commits, so we recommend people sign their commits and look for the 'verified' label on GitHub to ensure that things are as they appear to be.
As for repo impersonation – stay tuned, we are going to make it much more obvious when you're viewing an orphaned commit.
In summary: everything is fine, situation normal, the lark is on the wing, the snail is on the thorn, and all's right with the world.
Some people think RIAA’s DMCA notice is not legally valid, arguing RIAA is not the copyright holder and there is no infringing material. DMCA takedowns are for taking down works you own the copyright to; not for enforcing any arbitrary aspect of legislation.
It’s my understanding that service providers do not need to comply with illegal requests.
For example, if I DMCA’d <an oil producer>’s repository on accused violations of environmental protection acts, I don’t think it would be taken down, would it?
If GitHub was an independent company advocating for open source; would it have acted any different?
Note: Microsoft is a member of the RIAA.
Apple made waves and built lots of favour for resisting the FBI and challenging quasi-legal processes. They took risks and demonstrated their principles (Suing the FBI over a terrorist’s iPhone is unlikely to be the first recommendation from their legal counsel).
This smells like a qausi-legal process, and it would look great for GitHub/Microsoft if you do.
Let's pretend for the moment that the original youtube-dl DMCA had been valid, or that you removed youtube-dl due to an innocuous mistake. If I post youtube-dl to MY account, you have NO reason to take it down until you receive another takedown request from the RIAA for my repo. You certainly have no reason to ban users. There is nothing in your ToS which this violates.
I work on education projects which use youtube-dl in legal, non-infringing ways. I don't think the RIAA has a legal leg to stand on for reasons I'm not going to get to in this post.
Until github starts following DMCA processes properly, I CANNOT respond to the existing takedown request, since I have no standing. It's not my repo.
The right course of action for me would be to:
(1) Consult my lawyer and figure out if this is a fight I want to pick. I'm pretty sure I'd win in court if this went all the way, but I might go bankrupt first.
(2) Post youtube-dl to my repo.
(3) Wait for a DMCA takedown notice.
(4) Respond to it with a counternotice, and litigate with the RIAA.
Because github has decided to act as an arbiter on behalf of the RIAA, rather than a neutral third party, I cannot follow this process. github short-circuits this process at #2 by threatening to remove the repo and ban my account.
I'm sorry that you've chosen to side with the RIAA against the Internet. I'm gradually moving my business to gitlab. This is approximately what people thought would happen as a result of the Microsoft purchase.
That's how section 512 works. The RIAA letter referenced section 1201, not 512. There is no copyright infringing material to identify. The letter relates to distribution of copyright protection circumvention technology.
Maybe Github needs a new page explaining section 1201 takedowns.
https://cdn.loc.gov/copyright/1201/1201_background_slides.pd...
Just curious, what would the effects be if one were to use multiple accounts to automate the submission of DMCA takedown notifications for all <content> hosted on <content provider>? Does <content provider> honor takedowns only from or in preference to blessed accounts? Could one DoS <content provider> in such a manner? If a human has to review all DMCA complaints, would a flood of false claims DoS the human reviewers?
Asking for a friend.
https://docs.github.com/en/free-pro-team@latest/github/site-...
mentions:
> The DMCA requires that you swear to the facts in your copyright complaint under penalty of perjury. It is a federal crime to intentionally lie in a sworn declaration. (See U.S. Code, Title 18, Section 1621.) Submitting false information could also result in civil liability — that is, you could get sued for money damages. The DMCA itself provides for damages against any person who knowingly materially misrepresents that material or activity is infringing.
That's interesting; is US copyright law enforceable everywhere?
The DMCA works in much the way that its authors, the telcos and the MPAA and RIAA intended it to. To indemnify ISP's in return for their becoming enforcers for rights-holders ridiculously over-broad "anti-circumvention" clauses[0] which lead to outrageous abuses of the law (including anti-trust violations, attacks on the rights of consumers, academics, etc).
Now, Microsoft's lobbying machinery must have been in its infancy back then so the blame can't entirely be laid at their feet. But Microsoft don't seem to be doing anything to help either.
Fundamental to the problem is that youtube-dl (and many others) seem to be obvious candidates for exceptions to DMCA 1201. But the process around those exceptions seems not be working at all. Something which Microsoft appears completely tone-deaf and oblivious to[1].
So, with respect, I suggest you... get a grip to how you guys are going to be being perceived in this situation.
[0] Fritz Attaway, policy advisor MPAA. https://www.wired.com/2008/10/ten-years-later/ [1] https://beta.regulations.gov/document/COLC-2015-0012-0054
There is a way forward, Nat. You can reinstate that repo today, and tell the RIAA that they cannot use your online tool. They have to send your legal representation (in Alaska to slow it down) a certified, hand-signed letter through snail mail. Make a big public show of this process, and get public mindshare on your side.
It seems to me that this could be used to cause a lot of damage – target a popular open source project with a totally bogus DMCA notice, even if they instantly file a counternotice they still get made unavailable for 10+ days.
(Also, why 10-14 days? Why not just 10 days or just 14 days?)
> GitHub Asks User to Make Changes.
I think many would argue that the takedown notice wasn't "sufficiently detailed", especially when you consider the 1201 vs 512 issue.
>With potential damages multiplied across millions of users, cloud-computing and user-generated content sites like YouTube, Facebook, or GitHub probably never would have existed without the DMCA (or at least not without passing some of that cost downstream to their users).
Talk about backwards logic
Thank you.
I am going to believe that. Github CEO wanted a reason to open source it, and used a rogue leak in a win-win situation.
Why he didn't sign it to prove it was him? because the desktop client doesn't even have this basic git feature implemented ¯\_(ツ)_/¯ ...and everyone knows managers only uses GUI, Q.E.D.
Unless you’re referring to the generic “our hands are tied” tweet that said nothing: https://twitter.com/natfriedman/status/1321221940774723584?s... in which case, I suppose we’ll agree to disagree on what “all within their power” actually means.
[0] you failed to read the 1st paragraph of the linked article :)
By the way the serious design flaw where GitHub forges signatures on merge commits I told you about when you joined as CEO... Still not fixed.
The fact a commit can be shown as "verified" in the interface when I didn't sign it with my Yubikey is totally broken.
One issue is that you are loading profile images and creating links based on unverified emails (if I click the little picture next to the commit message I get to the impersonated profile). I mean I get that a proper solution might introduce unacceptable friction, but you can't really blame users for misunderstandings in the current state either.
Maybe add a checkbox in your profile like "Specifically mark unsigned commits" or even "don't associate unsigned commits to my account" as well.
> (!) This user usually signs their commits, but this commit is not signed. [Learn more]
Is what I've been surprised there isn't something like in the past.
> This is GitHub.com and GitHub Enterprise
It also contains linting config, ci workflows, dockerfiles, and other build related files that you probably wouldn't put in an "un-stripped/obfuscated tarball of our GitHub Enterprise Server source code"
The readme seems clearly aimed towards developers of GitHub.com
Instead of making conflict messages clearer and easier to work with using local files, contributors keep thinking the users are too dumb and adding (and changing) merge resolution hacks.
This boils up to github, as can be seen by teams who do not understand the very basic about git commits, and enable "squash commits by default" on their repos. With these teams, git commit history cease to be bit sized changes in a larger changeset, and become useless displays of the author interacting with the remote server while they upload small changes to tests to make the continuous builds get green.
Funds are safu?
But Github uses the same single Git repository for all forks, and they have an issue where you can access a branch/commit of a fork from the main repository if you know its hash. They should probably fix that at some point.
Ah I see, thanks for the explanation. I didn't know this was the case. I thought each fork would have its own `.git` folder. Seems like this approach could allow forkers to mess with the original repo, but maybe Git is designed in a way that this is mostly safe.
bit of a Wodehousian twist there. appreciated.
Robert Browning, Pippa Passes (1901)
A nice one. I didn't know it.
https://www.goodreads.com/quotes/314320-the-year-s-at-the-sp...
Insane.
(Reading docs...)
The desktop client explicitly does not support this, why is that?
Git does not make it trivial to impersonate commits. http://www.linuxjournal.com/content/signing-git-commits
As you yourself mentioned, very, very, very few projects/people sign their commits. Even fewer actually verify them.
Are you aware of the fact that it is irony in the original work?
https://en.wikipedia.org/wiki/Pippa_Passes
Have you read the newspaper in the last months?
I suspect irony on your side and if it's true you are kind of funny...
You "hacked" yourself. A majority of commits are not "verified", and a majority of users don't know to "look for" the verified label. Why didn't you make signing mandatory if you recommend it?
As for repercussions to your mismanagement, I will certainly stay tuned.
In summary: you're fucked.
Also, if users don't know the very basics of how Git works, they probably shouldn't be using it, and certainly not trusting it.
Security by oh yuck it's Ruby.
No, it is not.
This URL has been excluded from the Wayback Machine."
Apart from the code linked above, I don’t think you’re missing any significant information that was on that archived page whose URL is now excluded. The screenshot at the top of this article matches what I remember of that page: https://arstechnica.com/information-technology/2020/11/githu....
To describe the screenshot in words, for the sake of those who prefer text to images for whatever reason: the commit was an orphan with no parent; it had no commit history. The message was as SethTro quoted it: “felt cute, might put gh source code on dmca repo now idk”. GitHub’s interface suggested the committer was https://github.com/nat, but there was no “Verified” label. The interface described the commit as having been created “1 hour ago” at 2020-11-04 05:00:26, the time of crawling.
I remember reading somewhere that once You include some url in robots.txt archive.org even though can have it archived it will stop showing it to the public.
https://blog.archive.org/2017/04/17/robots-txt-meant-for-sea...
They have some interesting nuggets in their robots though like `/ExplodingStuff/` and `/account-login` both of which seem to be some accounts.
Or probably more possible - they just got in contact with the archive.org people.
EDIT: Anyone looking to try doing this, please support open alternatives instead: https://gitea.io/en-us/
This is GitHub.com and GitHub Enterprise.Most likely support. GH probably doesn’t want to support an open source version (triaging issues, reviewing 3rd party pull requests, having an open roadmap). Likewise it would probably be bad PR if they just dumped the code base and were really slow (or didn’t) respond to bug reports.
Being open source requires a lot more than just the source code being available.
The GitHub way of working has only been around since GitHub launched. Accordingly, the free software and open source communities that predate GitHub had their own ways of working before GitHub came along. Despite the widespread perception that GitHub makes doing open source easy, it comprises a set of practices that can be and are frequently more taxing than the alternatives. If GitHub is all you know, though, or you've forgotten, or you've just not noticed and never measured it, then it's easy to think that the GutHub way embodies the essentials of development in the open, even though its workflows are pretty bloated.
Well, you don't necessarily have to do those things. See for instance Apple's XNU Kernel. :)
Well, "maximum engineering secrecy" would be not releasing the source code to XNU. Apple is very secretive overall, to be sure, but not in this one respect.
I'm glad that XNU's source code is available—it lets you do a number of neat things. I wish more was available, but I'll take what we can get.
By extension, I don't support the idea that there's no point in releasing source code if you can't also release documentation and review outside pull requests. Making a tool available to the public is always better than hoarding information. All the other stuff is even better, but code is code.
---
All of that said, I recognize that for Github specifically, releasing code and not engaging with it might be a bad look, because their product is a code sharing platform. I don't think that applies to most companies, though.
I'm not trying to be mean or sarcastic or anything. Just look at how maintainers are treated for a week and you'll see exactly what I mean.
Just because something is open source, does not mean you have to engage with the "community". Slapping GPL on some code on a git repo somewhere, with a big sign saying if you don't like it you have the right to fork, so please fork off, is also open source, and a totally ok thing to do if you don't want to develop a community [Open source maintainers don't owe the world anything beyond what they freely want to give it]
why does anything needs to be open source in the first place?
Open-sourcing something should add value. Github doesn't see any value in doing so (and i would agree). It's not like github has any secret ingredient that makes github source special - gitlab has replicated most of github's functionality, and so has many open hosting platforms.
The value of github is mindshare, rather than anything code wise.
Not advocating FOSS approach for anything, and I'm not strictly speaking of GH in that context, but generally, malicious code loves being hidden into closed source software since it's the best effective way to keep it hard to control and correct. When privacy, security, safety, accountability become important, the FOSS approach might solve a problems or two. Of course the downside is that everyone can look at it, so building a billions worth business out of a fork + 10 line changes + renaming of a FOSS project cannot work without violating the license by keeping it closed.
I'm also genuinely curious how many people actively review all the code they actually run. I doubt anybody but the very largest tech companies and high-end government would actually be able to afford and resources such a feat, and even then they would have DMZ-type areas to detonate unaudited software.
Not help you learn.
- If it is so easy to shoplift, why don't stores just give out everything for free?
I generally like to open-source all my work. I'm working on a closed-source app, right now, but I think that it should be made open-source, once the embargo has passed. The backend is already open-source, as is the SDK.
I like to use the MIT license, which says that you can use the code how you like, but don't come whining to me, if you pooch it.
But I will, sometimes use a license that says "Here's the code for you to look at. It's not authorized for copying."
I think that it's a good idea to have it available. I seriously doubt there's anything in my stuff that is so proprietary that I'm afraid it will get ripped off. I do the same stuff everyone else does; maybe not as cleverly.
I just feel that it's good to show folks what's under the hood, if at all possible.
It depends. Figuring out how to design a large system still takes an appreciable amount of time even if you only end up gluing a bunch of other stuff together at the end of the day. And chances are you have something original in there somewhere. If the market is particularly cutthroat it might be a bad idea to risk giving your competitors even a slight edge.
I do appreciate the sentiment of proprietary source available projects though!
It's not that people do not pick locks because locks are so secure. It's because lock is polite way to say "do not enter here" and almost everyone respects that. If someone wants not to respect it, there are a multitude of YouTube channels showing how trivially easy it is to bypass it. But then, the legal aspect kicks in and punishments follow.
The same with stealing. It is trivially easy to steal. Moral code stops most people. Legal code punishes the rest.
Anything in our world can be "so easily acquired" if one does not care about laws and ethics. In that sense, the question posed in the original comment seems utterly bizarre to me. If anything that could be easily acquired were to be released for free, almost everything in the world should be released for free.
I've used it as a local-network git remote and generally enjoyed my experiences with it. It's not quite as developed of a web interface as Gitlab or Github, but as a git remote with a web frontend it's very usable and quick to deploy.
Zipped source (140 MB) mirror https://web.archive.org/web/20201104050247/https://codeload....
Similar things happened with Sony over Other OS. Sadly I bet there will be further attacks and leaks as time goes on here.
After the youtube-dl event many people became aware of those "hacks" because someone used it to "upload youtube-dl into the DMCA's repo".
Since those hacks are known by GitHub but they won't fix it, someone thought that the best way to "protest" against the decition was to push GH's source into the DMCA's repo impersonating GH's CEO.
That's my theory, a protest.
edit: DMCA
Come out against the RIAA if you must do something. Better still to let the process work, you don't win legal battles by committing crimes.
It's unreasonable to expect that the individual who took this action could go to court against the RIAA and extract their deeply embedded claws from the legislative, executive, and even cultural environment in which they've entrenched themselves.
Instead, it's reasonable to expect that somewhere in a conference room at say, Gitlab, Sourceforge, etc., sometime in the near future, some lawyer is going to mention "We got DCMAs from X for repo Y, and foo for repo bar, and from the MPAA for the popular repo (insert software tool here), obviously we should remove all of those repos." Maybe someone familiar looks up and says "that sounds a little silly, isn't that tool just a general-purpose media player?" And the lawyer might respond "We don't care-we don't have the budget to determine if the DCMA notice looks valid or not, if it turns out to be valid we have to do it and if it's not valid it costs us nothing to take it down". Actions like this provides a reasonable response of "But, as we saw with GitHub, some of our users and maybe even some of our own employees will resent us for this. It could cost us users, cost us bad publicity, or cost us real money. Let's put a couple people on this for an hour or two to estimate the pros and cons."
Github was just a middleman, responding rationally to their existing financial, legal, and moral incentives, in a battle between their little users and the RIAA. Like a little flock of oxpecker birds riding a rhinoceros, we wouldn't even register in a fight against a tiger, we can only warn our larger intermediary of the danger we perceive. The RIAA doesn't care that you're against them. We need to change the incentives for companies like Github that have a chance to be heard.
That being said, it also seems like GitHub is well within their rights to choose not to host a project accused of violating the law.
Tell that to the Civil Rights movement.
GitHub (via their owners Microsoft) is a member of the RIAA. They give the RIAA money to do stuff like this.
Quite possibly one of the most incorrect things I've ever read.
What's sad about this?
Regarding it's current, although waning, top choice for hosting open source software, I certainly do think we are starting to see the trade offs the open source community has been making.
There are tons of hosted GIT solutions. I've started using sourcehut [1] which seems like an rising platform that is also opensource. Hopefully youtube-dl ends up there instead of risking it on another big provider.
[1] https://sr.ht/
They don't have a mantra of all information wants to be free, and indeed their entire business model relies on hosting private repos.
I understand people wanting to host their open source on a server that's also open source. There's an argument it makes business sense if the closed source setup creates a long-term drift to GitLab or elsewhere. Nevertheless, I don't see the closed source as a contradiction.
You can love open source while producing a mix of open and closed source. Building a sustainable business often means being selective about it and seeing open source as a tool you can benefit from (enlightened self-interest), rather than an ideology you must adhere to at all times.
Which at the time of writing read as follows "It's not open source because the open source "community" is a liability and you want them far away from you at all times.
I'm not trying to be mean or sarcastic or anything. Just look at how maintainers are treated for a week and you'll see exactly what I mean."
[1] http://theorangeduck.com/page/reproduce-their-results#source...
This is not a bug, it's a part of how Git fundamentally works. If you want to mitigate it you have to sign your commits. GitHub could only attribute commits in the UI if they're signed, but I suspect that this is considered too much friction to enable.
https://www.reddit.com/r/programming/comments/jnpufo/using_t...
Using the same trick as the one with youtube-dl, I uploaded the entire GitHub backend source code to GitHub's own DMCA repo. Maybe now not only GitHub can have the chance to fix the "bug", but the entire community as well? ;)
Of course Drew DeVault thinks this way. He's trying to monetize his own github-like product, the sourcehut, so less people using GitHub means more people using sourcehut.
2) They most certainly did not "impersonat[e] Nat Friedman using a bug in GitHub's application"; they impersonated him using a design feature in Git.
I believe this isn't an existential threat to the company by any means.
https://pbs.twimg.com/media/De17PIKXUAE27W6.jpg:large
...and show that making it work in any browser, even text-based ones (as far as possible), is not hard.
I'd love to hear more about that.
I'd love to study it but not if just viewing it is a gray area
You can read bomb-making instructions, for heaven's sake. You can certainly look at this code. Just don't base a product on the ideas you got from looking at it.
Are there any verbatim copies of the original in your project. Did you sign an enforceable ip ownership or non-compete agreement with GitHub about trade secrets learned while working on their source code?
If the answers to those questions are no, there is a very long, very steep, uphill battle facing GitHub.
How many people can actually push to that repo? I wonder if it would be easy to figure out who actually did it...
Github does the thing where a cloned repo shares the object space with all repos of the origin, so you can use the same commit sha1 on any of the repos. All someone would have to do is clone the repo, make the commit to their clone, then share the link with the sha1 in the origin repo.
The sha1 blob may not actually appear in the main repo timeline.
For example, if I were to 'git clone' a project currently on gitlab, create a github repo for it, add that origin, and push... Well, that pushes commits authored by every single person who ever committed to that other repository. Do I have to make all of them re-push only their own commits in order? Can robots not mirror repositories anymore? What about commits authored by people without github accounts?
There's also the obvious issues of me merging a coworker's commit into my branch, or cherry-picking, or rebasing a third-party contributor's commits to update them before merge.
I think the case of "push an existing repo with N authors to a new repo" is a really compelling reason though for why that sort of "you can only push commits with your email" thing would not work.
All of the commits in that branch I pulled, regardless of who committed them (not me, presumably) should still be attributed to their original authors, and that's what will happen on the fork I push to GH.
This is fundamentally necessary to how the distributed nature of git works. If you want to assure others that commits really came from you, you need to sign your commits. But so few people do that, so the default is just to trust that commits are from who they say they are.
Perhaps GitHub could have a feature whereby you could toggle a setting so they won't link a commit to your GH user account unless it's signed by you. That still comes with its own problems (like say you submit a PR to some project, but the maintainer rebases master onto your branch before merging, which will kill the signatures).
But still, signing every commit is not really necessary. I personally only sign release tags, which implicitly cover all commits leading up to those releases.
Specifically, the way GitHub "embraced and extended" git (even before the Microsoft acquisition) is to have quite a few things outside the repository - like issues. You can take your source anywhere you want, but a project with 30,000 issues is going to have trouble migrating those issues.
Microsoft are running Bartertown and and are doing a pretty good impression of Master Blaster. Ironically youtube-dl just got dragged into the thunder dome.
You can trade outside Bartertown but it’s going to have low visibility.
People are afraid that they will do that again after relying on them as a foundation for their product or experience based on some trite promise of "we changed". So when something happens, the attribution and a hypothesis is already formed.
Unfortunately, based on my experiences attempting to fight their universal telemetry, the promise they made is invalid. So why should I trust any of their other promises? I think a lot of people are there too.
Or is there something specific you are accusing Microsoft of?
To give an analogy, I think the Internet is great. But I dislike the individuals who exploit it to send spam and propagate worms etc.
Capitalism offers a reward for those who are willing to put in the work.
Microsoft has helped billions of people around the world find employment, build software, use a computer, save time, money, and achieve something or another across pretty much every field in existence. For that, the capitalist system lets them dominate the market and gain all kinds of advantages. They deserve it as far as I'm concerned.
I'm wondering when one of the 1000s of services with write access to 100s of thousands of github run repos gets hacked or tokens expropriated and lots of repos suddenly get malicious commits.
I saw the headline and assumed this was a leak of someone else's source via stolen tokens.
https://web.archive.org/web/20201104050247/https://codeload....
mirror: https://anonfiles.com/Jax980m9p6/dmca-565ece486c7c1652754d7b...
c51717e6755ac0efdf22f7421e372f5b061724d61. is not equal to making it open to collaboration.
2. does not mean they can’t keep making money from it.
Viewing classified top secret nsa documents is not illegal unless you have agreed to never viewing them when gaining classified clearance. Anyone who does not have classified clearance is free to look at them.
You are not liable for what you read. Those posting it may be liable for infringement, and you could be liable if you infringement upon it. But reading is not yet illegal in the usa.
Reading something isn't inherently illegal itself, true. I'm fairly certain that intentionally navigating to a website that you know contains pirated content is a violation of copyright law though. (Of course no one is going to bother prosecuting you for it, but still.)
The laws surrounding classified information in the US are quite different from copyright law. They aren't relevant here at all. (https://en.wikipedia.org/wiki/Classified_information_in_the_...)
You haven't licensed the material so it's copying or possession by you is an illegal act in most jurisdictions. Your awareness of the nature of the content and its licensing status establishes intent.
Of course, if you don't redistribute it you shouldn't have any problems (at least in the US) because an IP address alone isn't sufficient to pursue someone in court here.
I really don't think any of that applies here though. This is an isolated leak of materials which you very clearly do not have any rights to access, copy, or possess in any manner whatsoever. I just can't see how "research purposes" helps you here. (Unless you happen to be an established researcher with an ongoing project that involves surveying source code leaks over time, perhaps? But again, you'd really need to consult a lawyer before doing something like that.)
Click "Code". I'm not suggesting you download it, that would be naughty.