GitHub Copilot exposes names in TODO comments
twitter.com
twitter.com
https://www.verywellfamily.com/top-1000-baby-boy-names-27576...
Soon we'll discover that Copilot is also exposing api keys!
What it shows is people often put brackets with their name after TODO and before the comment.
a MAJOR element of prompt engineering in stable diffusion for image generation is specifying the artist name(s) you want it to be like. so picking a defensive coder, terse coder, enterprise coder, etc, may reflect on the type of code you want.
And when the time comes to explain Copilot to laypeople (say, in court), this is an example they can understand.
I feel like the tweet is panic like there’s something wrong and data is being leaked which I don’t think is the case. However someone said it exposed keys. That could be bad for sure!
[1]: https://fossbytes.com/github-copilot-generating-functional-a...
AIUI, if you wrote a story about a character named "Harry Potter" and used nothing else from the originals, you'd still be violating copyright.
Not that these names are important. I just think a lawyer might use this example to show that Copilot is copying (without anyone's eyes glazing over), then show some complex code and argue that the code was both copied and important. The first makes the second easier to understand and accept.
GitHub watch for those and notify the provider of the API key so it can be revoked, as well as emailing the owner(s) of the repo which contained the key.
if it was in enough places that it got reproduced exactly via neural network, then there is zero chance it escaped their secret-scanners.
they even have some GHAS feature that lets you define your own regexes for secrets and will notify you if any are found.
none of this makes it safe to store secrets in code, though, these things just help tell you when you've done it.
[1]: https://fossbytes.com/github-copilot-generating-functional-a...
if I argue that "cars in the US must be equipped with seat belts," would you say "not before 1968!" and believe that we were both talking about the same time period?
I guess, as always, I should be absolutely specific about everything, lest someone fill in a gap incorrectly...
I’ve bookmarked this comment for the inevitable.
but ok. I'll have fun occupying a small portion of your brain for a while.
With the way machine learning networks work, they don't really learn one-off high-entropy strings like API keys.
The weights powering the models are only so big, there simply isn't enough memory to waste storing that type of data.
Instead, it's learned the general pattern of what an API key looks like and the fact that it's high-entropy random jibberish, with no predictable pattern. And when auto-completing code, it's just going to generate a random string of characters in approximately the correct format.
This can't be true. Copilot used to reproduce code to the inverse square root function from Quake source literally. They blocked it off by blacklisting the function name. Given that, why would learning API keys be impossible?
Second, it's like the complete opposite of "one-off". It's a block of code that has been copy/pasted into so many code bases, and then there have been hundreds of blog posts, articles and comments written about it. It literally has a Wikipedia article. Most of those places copy/paste the block of code verbatim. It shows up so many times in both the original gpt-3 datasets, and the extra datasets that GitHub did additional training on.
That block of code is famous. No wonder why the model though those sequence of tokens were important enough to store and regurgitate verbatim with the minimal amount of prompting.
But none of the same applies to an API key found in only one code base.
TODO: Fence the stolen diamonds! [GranPC]
They are, and they probably won’t be stopped, so each of us at the individual level have to decide if we are okay with it and what additional action to take to protect ourselves, if necessary.
Even in a court, it's really not a 'should' question. It's a "is this legal, and what is the penalty" question.
But a lot have barely any repos. There is even a "ferran" who doesn't have anything (public) on GitHub.
So, my guess is that Copilot stored a lot of GitHub accounts, and when we type "@", it autocompletes with any random ones from that list.
There is no relation with the code that it generates.
That’s not how that works and an mis-feature. No one would want a real username from a random list and there’s no reason to think it has a list of usernames somewhere. It for sure generates usernames in real-time the way you would if I told you to imagine a plausible one.
If you publish "TODO(@name)" in open source code in a public git repo, of course "name" is out there for anyone to read.
It might be more accurate to say "Publicly publishing code with your name in it on the internet exposes your name" than "github copilot exposes your name".
@joseph: I dress as a cow
@joebiden: I drink oil like a big Mack truck
@olivertwist: Yeah, I stole your wallet you can't have it
Haha, I got the President too!