GitHub Private Repos Considered Private-Ish
tylercipriani.com
tylercipriani.com
https://docs.github.com/en/get-started/privacy-on-github/abo...
I assumed up to this point they could only use public ones but this wording suggests otherwise.
GitHub aggregates metadata and parses content patterns for the purposes of delivering generalized insights within the product. It uses data from public repositories, and also uses metadata and aggregate data from private repositories when a repository's owner has chosen to share the data with GitHub by enabling the dependency graph. If you enable the dependency graph for a private repository, then GitHub will perform read-only analysis of that specific private repository.
If you enable data use for a private repository, we will continue to treat your private data, source code, or trade secrets as confidential and private consistent with our Terms of Service. The information we learn only comes from aggregated data. For more information, see "Managing data use settings for your private repository."
It seems pretty clear to me that this means they're allowed to use private repos to train copilot, etc.
I wonder if any researchers have tried putting fingerprinted source code into a private repo, and then (after it is retrained) getting copilot to suggest stuff that could only have come from the injected supposedly-private source code.
That would make a nice paper. I hope someone does it.
Maybe we have different definitions of the term "aggregated"?
This suggests to me that GitHub need to extend that text to explain what they mean by "aggregated".
I think GitHub need to clarify this themselves.
Whether that would hold up is another question. But yeah, I agree with the conclusion that they need to clarify this.
The clause seems to mean “we can do whatever we want with your data as long as we violate many people’s privacy at scale at the same time”.
Definitely needs clarification, though somehow I suspect this is all by design.
I'm sorry but when was Microsoft ever reputable? They have a long history (and reputation) of being merciless in every single way they can, and have for as long as I can remember.
However, me thinks this relates to the times before Github became an offering by Microsoft. But the deal was just too hard to miss, getting this massive army of minion coders who all pray to the octocat and now do the Balmers dance.
Oh so much fun, now it turns out, that all feed the new AI overlords.
It made a lot of sharp business choices in that decade, but it also left a LOT of money on the table for developers, as part of a strategic goal to grow the platform.
Then the 00s came, platform growth slowed (because they were already running on everything desktop), and the "vs linux" decisions started coming.
Microsoft was disreputable in the 90s to the point they were almost broken up several times.
Microsoft in the DOS days was still fighting hard for market. Microsoft in the Netscape days? Eh... less of a competitive claim. Post ~2005? No claim.
For me, the definition of "reputable" changes when you're a competitor among equals vs when you're a monopoly.
Please note I am not attempting to address the reputability of GitHub pre-acquisition. That is a separate matter.
They went on an open source charm offensive a few years ago. "Oh, we've turned over a new leaf" etc.
A lot of people believed that they'd had a legitimate change in heart because of the change in strategy.
More realistically, Linux had driven them into near irrelevance in the server market and just pushed them from "extinguish" or "extend" to "embrace".
Their dubious anti-Linux tactics via leaning on OEMs in the desktop market remained more or less unchanged.
and with vscode, copilot, and wsl2 they're doing a terrifyingly good job :-/
I really hope people don't let their guard down.
They're still probably the most indirectly trusted company on the earth
Why? because almost all enterprises do run some non-trivial amount of MS code
Either Windows, Azure / Azure AD or AD at all, Teams/Outlook, VS Code, or anything else.
And please, let's do not start arguing that some startup made of 30 people uses Macs only.
you said: I'm sorry but when was Microsoft ever reputable?
nobody said Microsoft had ever been reputable.
GitHub is the formerly reputable corporation here.
GP comment doesn't even make sense without that.
Just code is not an IP.
Your product or a specific algorithm is.
And in my opinion patent on algorithm should be illegal.
There.is no inherent problem hosting Code on GitHub.
You are not doing a good job if you move companies away from working setups due to this.
And they haven't had high security requirements anyway because everyone else normally hosts GitHub Enterprise or gitlab themselfs
But sure have your opinion but at least try to bring your issue actoss
Can you imagine how much intel Google Docs, GMail, Salesforce, Profitwell, etc have about company performance and plans?
I’m sure nobody is using any of that data to insider trade, to give just one example. Nobody would do that.
The top of the document is:
> GitHub aggregates metadata and parses content patterns for the purposes of delivering generalized insights within the product. It uses data from public repositories, and also uses metadata and aggregate data from private repositories when a repository's owner has chosen to share the data with GitHub by enabling the dependency graph. If you enable the dependency graph for a private repository, then GitHub will perform read-only analysis of that specific private repository.
> If you enable data use for a private repository, we will continue to treat your private data, source code, or trade secrets as confidential and private consistent with our Terms of Service. The information we learn only comes from aggregated data. For more information, see "Managing data use settings for your private repository."
Taking a single paragraph out of context isn’t worthy of the top-voted comment.
There are many reasons to distrust Microsoft. The wording of this particular paragraph explaining how data gets used in accordance with the linked terms of service (which are the actual governing documents, not the page you’ve linked to) is not one.
I agree that the particular wording is not sufficient to specify much of anything, but does it sure doesn't shut the door on the possibility either.
Unless you can meaningfully show that Microsoft is actively applying a subsidiary relationship (that is, where it directs OpenAI’s product direction), I have to disagree with your base notion.
At this point, I reiterate that your original claim is 100% FUD and disinformation.
I’m not asking you to trust GitHub or Microsoft, but legal terms have meaning and the terms do not support your assertions.
- Enable mandatory 2fa within your Github organization (if you don't use an organization, you probably should)
- Disable the ability to fork repos in your organization
- Configure and enable mandatory SAML authentication. In combination with mandatory 2fa, this makes phishing and even key leakage less likely (specific keys need to be double authorized for SAML, so that random private key you have on a random CI/CD platform from a few years ago can't access repos in SAML organizations even if it's leaked)
- Disable the ability to make repos public at the organization level
These are some quick setting changes that should address at least 3 of the 4 main issues in the article.
Additional suggestion:
Enable branch protection on master. Require at least 1 peer review to merge anything in it. Enforce branch restrictions to include repo admins so restrictions can't be bypassed. This should stop obvious mistakes like accidentally committing .git or random credentials.
Edit: if the author sees this, feel free to add these things to the article!
Of course it’s still possible to download the code and upload it to a separate repo (but then it’s not a fork).
It also breaks GitHub's normal protection against accidentally creating public forks of private repos (and the feature where it auto-deletes your private forks when you change jobs).
There are probably other ways in which disabling private forks undermines your organization's security, but those are pretty obvious.
The only one I can think of is what you mentioned (testing CI workflows that match a branch name) but at least at our company that can usually be done in a PR.
I don’t see any obvious reason why an employee would ever need to fork a company’s private repo to their personal GitHub account.
I don't see it as well, but Github docs seem to suggest that forking a private repo inside the same org is allowed/forbidden by the same setting?
It’s a typical workflow for inter-org codebase.
I do not see the benefit of forking private repositories and have disabled it in the company org that I manage.
Clone = copy of code that you can sync with the remote on GitHub
Fork = copy of code that isn’t connected to the original remote. Of course, every fork involves a clone. But not every clone is a fork.
I could be wrong, but “fork” isn’t really a concept native to Git.
> - Disable the ability to fork repos in your organization
If someone can read it, they can trivially fork it (clone locally, then republish as new repo). The only thing you're preventing with this advice is the free discoverability and tracking of forks which you get with forks created with the GitHub "fork" button. The forks are still there but now you have a harder time finding them.
Importantly when this is discovered you can fire the person who did it, and they can't say it's was an accident, while the fork button to the wrong place is potentially an accident.
If you want a collaboration repository with forking disabled, public is the default.
As far as I know the article's example of 'a developer forks a private repo and makes it public' is not possible in github.
Every team I've seen do this had their productivity drop by over a half when it was implemented. YMMV, but my normal heuristic is to see if I'm making 2x market value (due to stock vesting or whatever). If not, I dust off my LinkedIn profile and start catching up with old colleagues once this gets turned on.
The mindset that leads to enabling that policy implies a few bad things:
- The team has probably seen growing pains, but did not switch to feature branches for each team, which means it needs to have a release manager, but doesn't know what those people do.
- The product has inadequate testing, and there is no QA organization, and the mismanagement is creating revenue headwinds.
- Management doesn't trust prior hiring decisions, and has decided to treat the dev team like children (treating them like a cost center comes next).
- The company could be chasing revenue via regulatory compliance. There is a backwater of companies that adopt some viral compliance standard that mandates all suppliers also comply with the boneheaded compliance rules. This last reason is the least ominous of the possible reasons, and usually comes with checkbox style implementation of master protection (e.g., admins can override, and over half the team members are admins.)
Admittedly, branch protection requiring peer review into master was something we started for SOC 2 compliance.
But, it’s actually great if implemented well.
Some suggestions:
- Limit the use of “master” branch to code currently deployed to production (don’t use master to stage code that hasn’t been deployed yet). This branch should have the strictest restrictions (no admin override) because merging anything into this branch only happens immediately before a deploy. This also allows your team to assume that any pushes to master should/could trigger an automated deploy workflow
- Have a release branch where all PRs are merged/squashed into. Use this branch as the base branch for all PRs. At our company this is simply the “staging” branch. This branch can have looser restrictions (allow overriding restrictions). The release branch is anything that’s going to be deployed in the next release. Merged into the release branch can trigger automatic deploys to staging for QA team to do any final functional testing before code gets to production
- Or for a smaller company, don’t have any restrictions on the release branch. You’ll still be in compliance with all frameworks because, at a minimum, no code gets to production without a peer review (because the entire release branch has to be peer approved before it can be merged to master)
- Still allow hotfix PRs into master if you need to deploy something urgently without deploying the release branch. But hotfixes ideally should be rare (if they happen all the time it means you’re deploying a lot of buggy code to production and probably need to do better testing/QA)
There’s always a balance between security/quality control and productivity.
At the absolute minimum, you really should enable branch restrictions even if the only restriction is requiring all merges come from a PR. This will block developers accidentally force pushing their local master and potentially overwriting master branch completely (this has happened at our company prior to enabling restrictions, and it required another developer force pushing their, more up to date, local master to resurrect the correct state)
The goal should be to block actions that are obviously bad in all cases (e.g. pushing a commit directly to master without a PR).
Whatever your branch restrictions are, they should align with whatever your company’s internal code review processes are. Then, the branch restrictions are simply acting as a fall back in case developers make a mistake (e.g. avoids merging/pushing to master by accident)
For a small startup, all this advice is irrelevant because you probably care a lot more about how fast you can pump out changes and care much less about bugs getting into prod. And that’s fine.
These suggestions are mostly relevant for mission critical code bases where the restrictions align with your QA process. The branch restrictions should be enabled after documenting a QA process. Branch restrictions should not be arbitrarily enabled if the reasons for the restrictions don’t align with your existing internal processes/workflows.
I didn’t say anything about peer reviews.
Regarding the rest of your comment:
- if you are committing to a branch directly that then immediately goes to production you are doing something horribly wrong
- having a formal process for code reviews is like making sure there is a formal process to make sure employees wear pants to work. If this sort of thing has to be policed by security settings in the dev environment, then something deeper is wrong.
I mostly work on safety critical code, so I know what it means to ship a mission critical code base.
Most of the things you allude to suggest that you’ve never worked with a competent release manager or qa organization.
Edit: also:
> There’s always a balance between security/quality control and productivity
This is only true for definitions of “productivity” that exclude security/quality control, which usually mean that things are so out of whack, it is time to find a new job. By definition, the “productive” people aren’t worrying about product quality, and are being promoted for it, so soon the organization will be run by people that sabotaged the business.
You’re right, I oversee an 8 person product team for a 20 person company. Not a huge organization. QA is a shared function and we don’t have a “release manager” on staff.
I’m genuinely confused about what you’re trying to say. Sounds like you’ve had a bad experience with branch restrictions. It would be interesting to hear what the workflow was and how branch restrictions, etc, got in your way.
When we moved our code from privately hosted SVN to GitHub 10+ years ago, folks at GitHub were quite clear that we should not trust private repos with secrets like private keys or passwords.
So why have private repos at all? It is to allow control of collaborators. GitHub originally allowed everyone to see and fork every repo—a true open source approach. Private repos were added basically so corporate codebase managers would not have to spend tons of time rejecting random PRs and responding to random people.
For a LOT of private software projects, everything secret can be easily stored in the database, environments, or a dedicated tool for managing secrets. GitHub private repos are great for those. If my entire codebase is a highly sensitive trade secret, I would not use GitHub private repos, personally.
It's not really "private" if the backyard has a one way glass pane on one wall. Github should not be calling repos private if they're not.
Edit to add: would you store your bank password on a piece of paper you always leave sitting in your backyard? Think carefully about your metaphor.
Not that I'm advocating for checking passwords in git repos, Github or otherwise, but you metaphor depends on how rich you are and how bad is your neighborhood, I would think.
My mom has definitely done exactly that — her medical conditions make it very hard to use bank apps without full daylight and also hard to bring the notebook home after every use — and I don't think that her password management choices, however frowned upon they may be here, have ever been a serious problem in practice.
Go out in your fenced yard and assume you won’t be seen, but if you forget to close the gate someone might come back there anyway
For example, those repositories contain a lot of privately identifiable information, it is not that easy to get such a baseline ready for that _"should be treated as if they could be exposed to the world at any time any way."_
Depending on jurisdiction this can affect sensitive information that requires much stronger controls in place when you (rightfully!) expect the repository to become public despite it is a private one.
This seems to miss my point - I have no idea why PII is in a code repo.
The real solution to this is to have a business model that is about more than just some source code in a repository, or "things one person could undo". Concerns like years-long relationships with customers (aka trust), platform-style lock-in patterns, etc. How many times have various parts of AAA game studio or Microsoft codebases been compromised? I am still waiting for that hacker edition of HL3 and my free copy of windows that doesn't suck.
I actually spent a few minutes walking through a hypothetical where 100% of our latest code is lifted and taken to our biggest competitor. I think it would probably cause them more harm than good. Complexity is a hell of a thing. Just having a point-in-time snapshot doesn't really give you an advantage over someone who has been at it continuously for years. Perhaps we can strike a consulting contract with whoever steals our code...
I am not against reasonable measures (i.e. MFA, GH Enterprise, VPNs), but I won't go into paranoid-tier (Citrix-style desktops) over stuff like this ever again. Code is a cheap commodity in 2023. Enshrining your repository as if it is the actual vehicle of business value is indicative of poor leadership. Customers, relationships, execution, etc. are way more challenging and important.
> So, if you’re worried about it: stop putting sensitive data into private repositories.
Most of the issues mentioned in the post (misconfiguration, phishing, mistakes, zero-days) apply to all software, including non-cloud software. So the above advice is equivalent to "stop putting sensitive data into computers". It's run-of-the-mill popular-security nonsense that conveniently ignores that the alternatives come with their own risks, including security risks, and doesn't even attempt to perform a cost-benefit analysis.
Less of this stuff, please.
If your system relies on secrets being present in repos, it is a poorly designed system. There is no scenario where this is necessary, other than one not wanting to put in the effort to inject secrets sensibly.
Yes, "the alternatives come with their own risks", but those risks are where the "cost-benefit analysis" equals "not worth the effort".
Repos are databases that track changes in files, nothing more and nothing less. There are millions of repos on GitHub that don't contain application code.
Some people put their entire home directory in a Git repository. Home directories almost always contain secrets. That doesn't mean it's a bad idea, it just means the repo isn't meant for the eyes of others. In other words, it's a private repository.
...are directly attributed to the cited GitHub incidents. I only scanned the article once but I saw no FUD in it.
Give those tokens the minimum permissions possible. Make sure the impact of needing to rotate that secret doesn’t mean other secrets are exposed (looking at you: azure storage account primary access key!) Automated rotations make keeping them up to date in git untenable anyway.
Some services like Vercel and Firebase encourage bad practices in this regard. Firebase requires you to download a plaintext JSON of kingdom keys in order to do anything useful! Vercel KV storage encourages you to run an SDK command to download prod secrets into a local env file.
Attempting to transfer ownership of a repository to another user was aborted if the user had a repository of the same name—even if it was private.
Public GitHub doesn’t seem to have this issue with the transfer request system, though. Maybe it did at some point?
Excuse me?!
/edit:
Also what someone considers a secret and then not, is often not well defined. If management has no clue what this is about, it is often better to only commit and push on direct and simple work order, because these need to be well understood and you have the paper trail (as that author also suggests blameful retrospectives - IMHO hilarious).
Or do we have forgotten about the basic rules sending data over the interwebs to other people computers?
The topic of discussion shouldn't be how to secure your desk from spying eyes, but about why having post-it notes with passwords is bad practice and just a bad idea overall.
If your private github repo accidentally goes public, the response should be "that's annoying but ultimately harmless", anything else is misguided.
The problem is not keeping those passwords in a secure location, treat it like a stack of $100 bills.
Or we return the metaphor to github repos, having a separate cabinet is like having a secret vault so that secrets are not directly in plain view in the repo itself, which is exactly what you should be doing.
The reason secrets scanning even became a thing is because of how often secrets get committed to git. Some of them even lead to intrusions.
Uber (2016) – Attackers gained unrestricted access to Uber’s private Github repositories, found exposed secrets in the source code, and used them to access millions of records in Amazon S3 buckets.
Scotiabank (2019) – Login credentials and access keys were left exposed in a public GitHub repo.
Amazon (2020) – Credentials including AWS private keys were accidentally posted to a public GitHub repository by an AWS engineer.
Symantec – Looking at hardcoded AWS keys in mobile apps, discovered they had a much wider permissions scope and led to a significant data leakage.
GitHub – Over 100K public repositories on GitHub were found to contain access tokens.
I've worked at companies with developers who didn't know that once committed, the secret remains in the history even if a subsequent commit removes it. It's not trivial, and involves rewriting the history[1]. There's also no way to fix clones of the repo, and there are a handful of other ways secrets can still leak.
The most secure way to deal with secrets accidentally committed to git is to rotate the secret.
1 https://docs.github.com/en/authentication/keeping-your-accou...
You can use this and crenedials rotation, amazing, I know...
https://docs.github.com/en/code-security/secret-scanning/abo...
https://docs.gitlab.com/ee/user/application_security/secret_...
https://circleci.com/blog/detect-hardcoded-secrets-with-gitg...
https://www.gitkraken.com/media/events/azure-spring-clean-20...
edit: used config/secrets management for as long as I can remember doing this cloud stuff (several years), so the excuses are very, very poor imho.
Sure, but in computer programming “secrets” is also industry jargon for small strings of characters that enable authentication, like passwords or private keys, which have much higher standard of secrecy than the rest of the codebase.
1. Write code v1 2. Add secret 3. Write code v2 4. Rotate secret 5. Oops, some kind of problem, let's go back to known-good and redeploy (2). Broken because it tries the older secret, not the rotated secret.
Just don't store secrets in version control.
Moreover, the process of securing approval for NAT inbound rules, particularly for integrations with services like Jira, Slack and so on, often turns into a labyrinthine ordeal. It usually involves navigating the differing viewpoints and interpretations across multiple departments, each with its own unique stance, which further exacerbates the complexity of the task.
Of course then there is the embarrassment of azure ipv6 so it's perhaps somewhat forgivable:)
On a side note, I'm trying to imagine what "sensitive" code would be read, incorporated into an LLM such as Co-pilot, and somehow have any meaningful impact to me once incorporated?
I wrote a Bash script to copy our ~25 repositories to a new organization using the GitHub REST API. Turns out that the default when creating a new repository was to make it public. Within a few minutes I was getting emails by third party services who had downloaded and parsed our code, discovering a long abandoned AWS credential hard-coded in an older codebase. (I think GitHub now does this for you)
I was lucky that nothing important was actually compromised (we had done a decent job of keeping secrets out of our repos and the one exception was an account that was long gone), but it was an eye-opening experience. If these services had found and downloaded our code in minutes you can assume any repo made public, even for a few moments, has been downloaded and cataloged by a potential bad actor.
Trying to develop on a code base without history and a one time snapshot seems quite hard.
https://github.com/AGWA/git-crypt has been solid for me
Nevermind, I guess they can be web developers :p
Why not inject the credentials at runtime, from a system meant for this, with support for auditing and key rotation etc?
This is essentially a shameless admission and they not even hiding it any more:
Private repository data is scanned by machine and never read by GitHub staff. Human eyes will never see the contents of your private repositories, except as described in our Terms of Service.
This is another great reason to self host instead of using services like GitHub openly reading your private code, which I have been saying for years.
[0] https://news.ycombinator.com/item?id=34258431
[1] https://docs.github.com/en/get-started/privacy-on-github/abo...
Every gitops platform on the planet supports not pushing secrets to version control. There is no valid reason to have credentials inside a git repo ever. If the repo itself is what's secret, you should have known the risk when you pushed it to someone else's server. There is no excuse when gitlab is free and a docker container away.
Do we? Putting secrets in git is against policy almost everywhere.