Characterizing secret leakage in public GitHub repositories
blog.acolyer.org
blog.acolyer.org
[1]: https://www.dannyguo.com/blog/i-published-my-aws-secret-key-...
Specifically, I'm building a SaaS (https://www.locktower.com/) for organizations (or security teams) looking to have a managed solution for detecting leaked secrets in GitHub/BitBucket/etc. I'm in the process of building an on-prem version as well. Overall, I really hope to help drive down the number of unresolved leaks that the authors found.
Webhook > AWS API Gateway > Lambda
The Lambda uses the new(ish) Layers feature so it can use Git. I then use the truffleHog[0] library to scan for entropy/regexes inside the commit.
If something is detected, it posts to an SNS topic, which is currently subscribed to by another Lambda that posts an alert to my team and the Security team's Slack channel.
It then calls the GitHub API to make the repo private to limit the exposure.
Rephrasing:
What if you use a git server like GitHub that doesn't allow you to install server side hooks?
Cloud providers have proprietary solutions, but those don't work on other providers (or your local dev env).
Rolling your own secrets server seems like an expensive centralized disaster waiting to happen.
It seems like putting a secret into source code is one of the least risky options. Just make sure it's not in a public git repo.
Do repo cleaning tools such as https://rtyley.github.io/bfg-repo-cleaner/ leave the original commit's SHA-1 ID intact?
https://help.github.com/en/articles/about-token-scanning
https://github.blog/2018-10-17-behind-the-scenes-of-github-t...
Using a .env file is a bit of an anti-pattern, partly because many applications expect it and thus will be affected by it in ways you might not want. But also because passing configuration to applications via environment is not great, because then the values ar all static and the only way to change them is to restart the app. Better to have a function that can reload a real data format (json, yaml, ini) at run-time.
Each environment that runs your app will need to have its own 'env file' because every environment is slightly different, and they're coupled to deployments. So I'd keep your environment stuff wherever your deployment stuff is; with your terraform/ansible/puppet/chef configs, or etc/consul, or an S3 bucket, or SSM, etc. Create it at deploy time, pull it into the app at run time.
If you are running locally, you should be using your own secret keys that are configured in your user directory with
aws configure
If you are running on anything within AWS you should be using a role attached to your EC2 instance or lambda and the SDK can retrieve your keys automatically.Unfortunately, every single third party code sample on the internet has you including the secret keys in your code.
I have seen people doing this in their Dockerfile
ADD . /src
To add all the sources in the image, and inadvertently just made public all the secrets that were in the .env file.Personally, I like to keep all my secrets very far from my repository.