The Python Package Index is now a GitHub secret scanning integrator
github.blog
github.blog
- AWS IAM token for S3 upload access to a throwaway dev bucket. The bucket had already been deleted but still... Got an email about it informing me the IAM token had been revoked by AWS within 5 minutes
- A Slack webhook notification URL/secret. Committed as a example on a working branch and then git rm'ed but still active. Got an email about it and token revoked by Slack automatically within 5 minutes.
- A Mapbox API token. This one was funny. The token was indeed in there and functional but was in the docs/sample code for a dependency. Still, we got an email within the hour about it and were able to investigate.
Edit: In this case we intentionally kept the commit history. A safer alternative (and one we normally practice) is to start a fresh repo for the open source variant.
Commit histories can spill a lot of secrets that are easy to overlook.
Same problem here with inner source, that goes open source.
I feel sorry for all our internal committers, however I know of "secrets", that went into the commit history. We are still considering our option, but tend to opt for deleting our commit history entirely and build a wall of fame for the former committers.
I also like shhgit[2] for looking for secrets in repositories. (I don't think shhgit will look back in the git history for you though).
1. Get an old copy 2. run dictionary attack 3. prosper
https://github.com/zulip/zulip/blob/3.3/tools/zanitizer
https://github.com/zulip/zulip/blob/3.3/tools/zanitizer_conf...
The script was really fast (all ~10000 commits in a few minutes), which allowed us to iterate quickly on its configuration as we audited using gitk and other tools for remaining items to scrub.
Doing this work allowed us to release with an essentially complete history going back to the first commit in 2012, which has been a really valuable resource for understanding why various Zulip subsystems were written the way they were.
Nowadays there are other tools for scrubbing history that might be more polished, like BFG: https://rtyley.github.io/bfg-repo-cleaner/
Note that it's also possible to go back and rewrite history (e.g. if you know what the tokens are and where/when they were committed), to preserve Git history while cleaning out tokens. It can be mildly slow or complicated, but there are tools to automate it, such as BFG Repo Cleaner[0] which is relatively easy to use (once you learn it).
There are other awesome rewriting tools, like git filter-repo[1], but that operates solely on the structure of the repository (i.e. it can manipulate basically anything except file contents). Great for removing unwanted files or directories extremely fast, but not good for removing tokens (unless you want to remove the entire file the token was in).
[0] https://rtyley.github.io/bfg-repo-cleaner/
[1] https://github.com/newren/git-filter-repopsanford also mentioned truffleHog and others, lstamour mentioned https://github.com/cloud-gov/caulking which is built on gitleaks which looks good. caulking's customized list of patterns for gitleaks is here https://github.com/cloud-gov/caulking/blob/master/local.toml Looks like it would have found the keys in my example case no problem.
The integration means that GitHub knows to recognize this format, and calls some API of pypi.org when it finds one so PyPI can revoke it.
As always, please allow me to lament that we don't have a standard for this, such as secret-token:pypi.org/9NX39cdNn0AH1cCl1bMT48eKzf4Rhvw1mipk1FZTPrpR9, which would let any system know that this string is a secret and that pypi.org should be notified (for example via POST pypi.org/.well-know/compromised-secret). See also https://news.ycombinator.com/item?id=25978185
This page also mentions that they "strongly recommend you implement signature validation in your secret alert service", but I'm not sure why. Isn't the fact that they send valid tokens proof that they have really found a leak?
Something similar for tokens would be really useful.
They're actually just macaroons[1] internally, which means that they could easily be upgraded at some point to include a reporting URL like you mention.
Just as a tidbit: they were originally prefixed with "pypi:" rather than "pypi-", but that colon caused problems for a few packaging utilities. Any sort of in-band signaling like that is unlikely to gain widespread adoption for exactly that reason :-)
[1]: https://en.wikipedia.org/wiki/Macaroons_(computer_science)
Your reporting endpoint seems protected by a secret key that GitHub holds. Any reason PyPI can't accept anonymous submission of compromised tokens? If I find a PyPI token on my own server, can I not post it to https://pypi.org/_/github/disclose-token without getting a key from you first?
I don't believe it's something standardized or considered by the original whitepaper. Macaroons have the ability to contain arbitrary data, however, so it wouldn't be difficult to add revocation information to them.
> If I find a PyPI token on my own server, can I not post it to https://pypi.org/_/github/disclose-token without getting a key from you first?
I wasn't part of the design, but my first thought goes to preventing the endpoint's use as an oracle: after a compromise, a malicious agent might find it useful to have an unlimited endpoint to test their stolen credentials against. Restricting use to a limited set of trusted entities avoids that.
Also, the endpoint sends a 204 with no information about the validity of tokens, making it not much of an oracle. I think the payload is processed in the background too, preventing timing attacks.
Rate-limiting was just the easy example. Other endpoints are subject to additional constraints: tokens don't directly carry their user information (IIRC), so someone with a collection of stolen tokens may not know which projects they can control. Similarly, tokens are scoped, so "create a new project" isn't an ability that an arbitrary token can necessarily do to gain more information about its rightful owner.
Like I said, I don't know too much about the actual design decisions for that endpoint! That was an educated guess, based on what I might have done.
As for the endpoint being an oracle, the endpoint doesn't really need to respond to the reporting client other than the revocation request has been received.
Whether or not it's a security issue depends on how the token is being used. Allowing potentially arbitrary parties to revoke tokens right before, say, a critical security release feels like a potential issue to me. Then again, I suppose they could do that by proxy by just publishing it on GitHub and letting the secret scanner do the work.
Long story short: I'm idly speculating. For all I know, they did it because allowing arbitrary parties to report leaked secrets would result in unacceptably high FP rates. I wasn't privy to the decision.
If the third-party has the token, they can make releases *adding* critical security issues.
0) Cache invalidation
1) Naming things
5) Asynchronous callbacks
2) Off-by-one errors
3) Scope creep
6) Bounds checking
http://people.cs.uchicago.edu/~wiseman/humor/ai-koans.html
Moon instructs a student
One day a student came to Moon and said: “I understand how to make a better garbage collector. We must keep a reference count of the pointers to each cons.”
Moon patiently told the student the following story:
“One day a student came to Moon and said: ‘I understand how to make a better garbage collector...
[Ed. note: Pure reference-count garbage collectors have problems with circular structures that point to themselves.]
FTFY
Usually its junk, but occasionally you do get lucky and find tokens.
There's no winning these battles..
How else would you want to pronounce it?
They have integrations with a bunch of services to recognize the tokens, and disable them. This means malicious users can't copy/paste them, spin up servers and leave you with a big bill. (Ideally, of course it could still happen, but the aim is to prevent that kind of thing.)
I think it’s good because the risk of a package being taken over is low, but very damaging if it occurs in a widely used package.
Is there an argument (security by obscurity?) that that makes it easier to spot it and abuse it?
Or would it be better to encode it in the secret bits somehow, add 16 control bits that have known values?
I've seen devs share a snippet of code with an AWS Access Key Id/Secret in it using gists and we immediately got a notice from Amazon about that key being compromised.
I usually try and prefix e.g. fields in config files with "secret" to make it obvious they shouldn't be committed.
https://docs.github.com/en/code-security/secret-security/abo...
Glad they got our back.
> Fixes #6051
> See #7124 reverted in #8555 due to #8554 which is addressed in #8562 (pfew...)
> Should not be merged before #8562: EDIT:
>
> Re-revert of the code. The bug that caused revert was splitted into #8562
Software development in a nutshell, everyone.Unfortunately it's marked as "Future," so it's still a ways out.
Is it an easy mistake to make, for someone to inadvertently commit and push a "secret PyPI token"?
Most python devs will probably never publish to PyPi, but this can save some headaches for those who do, especially for the first time.
When I write, say, bash scripts which do work using ssh, I don't specify a password: The user running the script will provide their own manually, or use ssh-copy-id, or edit the authorized_keys file on the target machine if they want to save themselves some typing. That is - authentication is decoupled from my script's actual work. Why is that not how things work with PyPI?
It's only a rule because people have made the mistake enough to learn the lesson...
No, it's not “totally verboten” (forbidden by whom?), and people do it all the time. Mostly, perhaps, for stuff they aren't planning to share, but plans change.
It doesn't help that lots of example code embeds placeholders for secrets directly (with notes to replace with your actual credentials), so lots of stuff gets embedded in the course of copy-and-paste coding.
The elimination of a distinction between “safety” and “security” is unhealthy imo, as it leads to a failure to distinguish between unintentional harm caused by nature, and intentional harm caused by other people.
E.g. “safety first” is only intelligible if it doesn’t also prevent you from trusting anyone (which is what would be implied by “security first” as a general priority).
IMHO, Github should make it mandatory for integrated services to provide this feature.