Genuine question, do you not use GitHub for things other than copilot? It seems to me either the privacy issues of copilot are overblown or the privacy issues of GitHub itself are underblown, because they both end up with basically the same data.
Genuine question, do you not use GitHub for things other than copilot? It seems to me either the privacy issues of copilot are overblown or the privacy issues of GitHub itself are underblown, because they both end up with basically the same data.
In this context I think it's important to note the distinction between "Copilot for Individuals" and "Copilot for Business", because for twice the money you potentially get a lot more privacy:
> What data does Copilot for Individuals collect?
> [...] Depending on your preferred telemetry settings, GitHub Copilot may also collect and retain the following, collectively referred to as “code snippets”: source code that you are editing, related files and other files open in the same IDE or editor, URLs of repositories and files path.
> What data does Copilot for Business collect?
> [...] GitHub Copilot transmits snippets of your code from your IDE to GitHub to provide Suggestions to you. Code snippets data is only transmitted in real-time to return Suggestions, and is discarded once a Suggestion is returned. Copilot for Business does not retain any Code Snippets Data.
say, if training is determined to be fair use
https://docs.github.com/en/site-policy/privacy-policies/gith...
Maybe I rely too much on the community to red flag these things. I don't install questionable extensions. I wouldn't have tested copilot at all if it hadn't received such enthusiastic support (here, among other places). The fact that it isn't sandboxed to the document you're working on should make it an absolute malware pariah, and this is the first I'm hearing of it.
self-hosting solves those problems: if done right, it gives me a thing which is static until i explicitly choose to upgrade it, so i can actually integrate it into more longterm workflows. if enough others also do this, there’s some expectation that when i do upgrade it i’ll have a few options (models) if the mainline one degrades in some way from my ideal.
LLMs are one of the more difficult things to self-host here due to the hardware requirements. there’s enough interest from enthusiasts right now that i’m hopeful we’ll see ways to overcome that though: pooling resources, clever forms of caching, etc.
though that too comes with its own set of privacy issues, especially when done at a larger scale with strangers on the internet as opposed to just with your personally known and trusted buddies… and rare are the people for whom the latter happen to contain enough people with strong enough hardware that that pool would be sufficient.
Personally, yes, I do host with Github, bit previously that was a very well-defined boundary - only what I deliberately committed and pushed went to them. Now, it's potentially anything in my editor, so I make sure to use another program for scratch files. I probably don't have anything I'd really care about them getting their hands on, but still, it's better not to get in a bad habit and then slip up someday, pasting in private or sensitive details without thinking.
If I can eliminate that as a potential source of mistakes, so much the better.
I certainly not use Github for everything.
That being said, I've way too often encountered people for whom git oddly only exists in the form of Github, to the point of those being synonymous to them. The concept of a self-hosted repo seems totally foreign to them, and the only local git trees they have is what their IDE created for them in the directory of their checked out or created source project. It's weird, but such people exist… and lots of them.
It's like learning react first ,and html+css later, it's doable but a bit convoluted