https://docs.github.com/en/get-started/privacy-on-github/abo...
https://docs.github.com/en/get-started/privacy-on-github/abo...
I'm sorry but when was Microsoft ever reputable? They have a long history (and reputation) of being merciless in every single way they can, and have for as long as I can remember.
However, me thinks this relates to the times before Github became an offering by Microsoft. But the deal was just too hard to miss, getting this massive army of minion coders who all pray to the octocat and now do the Balmers dance.
Oh so much fun, now it turns out, that all feed the new AI overlords.
It made a lot of sharp business choices in that decade, but it also left a LOT of money on the table for developers, as part of a strategic goal to grow the platform.
Then the 00s came, platform growth slowed (because they were already running on everything desktop), and the "vs linux" decisions started coming.
Microsoft was disreputable in the 90s to the point they were almost broken up several times.
Microsoft in the DOS days was still fighting hard for market. Microsoft in the Netscape days? Eh... less of a competitive claim. Post ~2005? No claim.
For me, the definition of "reputable" changes when you're a competitor among equals vs when you're a monopoly.
Please note I am not attempting to address the reputability of GitHub pre-acquisition. That is a separate matter.
They went on an open source charm offensive a few years ago. "Oh, we've turned over a new leaf" etc.
A lot of people believed that they'd had a legitimate change in heart because of the change in strategy.
More realistically, Linux had driven them into near irrelevance in the server market and just pushed them from "extinguish" or "extend" to "embrace".
Their dubious anti-Linux tactics via leaning on OEMs in the desktop market remained more or less unchanged.
and with vscode, copilot, and wsl2 they're doing a terrifyingly good job :-/
I really hope people don't let their guard down.
They're still probably the most indirectly trusted company on the earth
Why? because almost all enterprises do run some non-trivial amount of MS code
Either Windows, Azure / Azure AD or AD at all, Teams/Outlook, VS Code, or anything else.
And please, let's do not start arguing that some startup made of 30 people uses Macs only.
you said: I'm sorry but when was Microsoft ever reputable?
nobody said Microsoft had ever been reputable.
GitHub is the formerly reputable corporation here.
GP comment doesn't even make sense without that.
Just code is not an IP.
Your product or a specific algorithm is.
And in my opinion patent on algorithm should be illegal.
There.is no inherent problem hosting Code on GitHub.
You are not doing a good job if you move companies away from working setups due to this.
And they haven't had high security requirements anyway because everyone else normally hosts GitHub Enterprise or gitlab themselfs
But sure have your opinion but at least try to bring your issue actoss
Can you imagine how much intel Google Docs, GMail, Salesforce, Profitwell, etc have about company performance and plans?
I’m sure nobody is using any of that data to insider trade, to give just one example. Nobody would do that.
I assumed up to this point they could only use public ones but this wording suggests otherwise.
The top of the document is:
> GitHub aggregates metadata and parses content patterns for the purposes of delivering generalized insights within the product. It uses data from public repositories, and also uses metadata and aggregate data from private repositories when a repository's owner has chosen to share the data with GitHub by enabling the dependency graph. If you enable the dependency graph for a private repository, then GitHub will perform read-only analysis of that specific private repository.
> If you enable data use for a private repository, we will continue to treat your private data, source code, or trade secrets as confidential and private consistent with our Terms of Service. The information we learn only comes from aggregated data. For more information, see "Managing data use settings for your private repository."
Taking a single paragraph out of context isn’t worthy of the top-voted comment.
There are many reasons to distrust Microsoft. The wording of this particular paragraph explaining how data gets used in accordance with the linked terms of service (which are the actual governing documents, not the page you’ve linked to) is not one.
I agree that the particular wording is not sufficient to specify much of anything, but does it sure doesn't shut the door on the possibility either.
Unless you can meaningfully show that Microsoft is actively applying a subsidiary relationship (that is, where it directs OpenAI’s product direction), I have to disagree with your base notion.
At this point, I reiterate that your original claim is 100% FUD and disinformation.
I’m not asking you to trust GitHub or Microsoft, but legal terms have meaning and the terms do not support your assertions.
GitHub aggregates metadata and parses content patterns for the purposes of delivering generalized insights within the product. It uses data from public repositories, and also uses metadata and aggregate data from private repositories when a repository's owner has chosen to share the data with GitHub by enabling the dependency graph. If you enable the dependency graph for a private repository, then GitHub will perform read-only analysis of that specific private repository.
If you enable data use for a private repository, we will continue to treat your private data, source code, or trade secrets as confidential and private consistent with our Terms of Service. The information we learn only comes from aggregated data. For more information, see "Managing data use settings for your private repository."
It seems pretty clear to me that this means they're allowed to use private repos to train copilot, etc.
I wonder if any researchers have tried putting fingerprinted source code into a private repo, and then (after it is retrained) getting copilot to suggest stuff that could only have come from the injected supposedly-private source code.
That would make a nice paper. I hope someone does it.
Maybe we have different definitions of the term "aggregated"?
This suggests to me that GitHub need to extend that text to explain what they mean by "aggregated".
I think GitHub need to clarify this themselves.
Whether that would hold up is another question. But yeah, I agree with the conclusion that they need to clarify this.
The clause seems to mean “we can do whatever we want with your data as long as we violate many people’s privacy at scale at the same time”.
Definitely needs clarification, though somehow I suspect this is all by design.