GitHub watch for those and notify the provider of the API key so it can be revoked, as well as emailing the owner(s) of the repo which contained the key.
if it was in enough places that it got reproduced exactly via neural network, then there is zero chance it escaped their secret-scanners.
they even have some GHAS feature that lets you define your own regexes for secrets and will notify you if any are found.
none of this makes it safe to store secrets in code, though, these things just help tell you when you've done it.
[1]: https://fossbytes.com/github-copilot-generating-functional-a...
if I argue that "cars in the US must be equipped with seat belts," would you say "not before 1968!" and believe that we were both talking about the same time period?
I guess, as always, I should be absolutely specific about everything, lest someone fill in a gap incorrectly...
I’ve bookmarked this comment for the inevitable.
but ok. I'll have fun occupying a small portion of your brain for a while.
With the way machine learning networks work, they don't really learn one-off high-entropy strings like API keys.
The weights powering the models are only so big, there simply isn't enough memory to waste storing that type of data.
Instead, it's learned the general pattern of what an API key looks like and the fact that it's high-entropy random jibberish, with no predictable pattern. And when auto-completing code, it's just going to generate a random string of characters in approximately the correct format.
This can't be true. Copilot used to reproduce code to the inverse square root function from Quake source literally. They blocked it off by blacklisting the function name. Given that, why would learning API keys be impossible?
Second, it's like the complete opposite of "one-off". It's a block of code that has been copy/pasted into so many code bases, and then there have been hundreds of blog posts, articles and comments written about it. It literally has a Wikipedia article. Most of those places copy/paste the block of code verbatim. It shows up so many times in both the original gpt-3 datasets, and the extra datasets that GitHub did additional training on.
That block of code is famous. No wonder why the model though those sequence of tokens were important enough to store and regurgitate verbatim with the minimal amount of prompting.
But none of the same applies to an API key found in only one code base.