Finding secrets by decompiling Python bytecode in public repositories
blog.jse.li
blog.jse.li
While for the other secrets, I’ve got nothing, there is never a reason to have AWS secret keys in your code or in application specific configuration files.
Every AWS SDK will automatically read your keys from your config file in your home directory locally. Just run
aws configure
When you run your code on EC2, Lambda or ECS, the same SDK’s will automatically get the keys associated with the attached role.Repo should have a `README` that says the secrets are in the team 1password account, talk to <team member> to get a 1password account, and get added to the group that has access to the vault with the creds.
The repo should have a source script that will pull the credentials[1] and `export` them to your ENV. `direnv` can make that happen automatically[2], or you can run that script from your `.bashrc` or similar.
You can do something similar with your favorite secrets manager. I've used a similar approach before, with good results.
[1] https://support.1password.com/command-line-getting-started/
Have your code fetch any secrets it needs from AWS Secrets manager, using the name of the secret (which does not have to be kept secret, so it can be in your source repo)
That way, you don’t have to put secrets in your environment, with the risk of leaking them.
I'm just pointing out that AWS Secrets Manager is not at automatic, no-brainer win
The other solutions don’t integrate with AWS IAM. Something has to grant access to the password vault. In the case of Secrets Manager/Parameter Store you just grant access to the role attached to your EC2 instance/ECS cluster/Lambda.
Question: what does python do if it doesn't have write permission in the current working directory? Not write the cache?
You can also set PYTHONPYCACHEPREFIX if you want to use 'a mirror directory tree at this path, instead of in __pycache__ directories within the source tree' - https://docs.python.org/3/using/cmdline.html#envvar-PYTHONPY...
The PEP mentions your case at https://www.python.org/dev/peps/pep-3147/#case-5-read-only-f... . But honestly, I don't really follow the answer. I think it means "ignore creation and write failures."
E.g., if you want to back up your home dir, and omit caches, since they can be regenerated. It's a lot easier if programs write their cache data to ~/.cache / $XDG_CACHE_HOME than if they intermix it / scatter it about.
At least, that's what I interpret "OS-specific" to mean.
I have seen a surprising number of people — some who are engineers by profession too, and ought to know better — just git add everything, and then commit it all without looking. One should review the diff one has staged to see if it is correct, but alas…
Perhaps it's worthwhile for someone to blog about this more/promote this as a best practice? Though what's missing is the hook to connected it as appropriate for the given platform.
I see now that "OS-specific" was meant to be interpreted as "the OS-defined mechanism to find a cache directory", not "a cache directory which differs for each operating system".
I would not have been confused by the term "platform dependent", which is what Python's tmpdir documentation uses, as in: "The default directory is chosen from a platform-dependent list" at https://docs.python.org/3/library/tempfile.html?highlight=tm... .
In the case where the interpreter can't find any place it has write access to to write the cache, it will not write it, yes. That means it will have to re-parse the source file into bytecode (and fail to write the bytecode to cache) every time it is loaded.
I'd argue that it's more appropriate to follow the 'explicit is better than implicit' from the 'guidelines' of https://www.python.org/dev/peps/pep-0020/ .
The bulk of your pycs are generated during package install. What tends to remain in the usual case is a handful of files representing app code or similar.
I have lost count of the number of times I've seen someone lose an hour due to it. I can also count many instances of QA environments becoming inexplicably bricked by it. The correct fix for this requires opening the .py and hashing its content, at least doubling the amount of IO required to start a program. They were a great feature when parsing small files was noticeably slow, but this hasn't been true for almost 20 years.
It's therefore worth turning the question around: why do you think pyc files are useful?
https://docs.python.org/3/reference/import.html#pyc-invalida...
Verifying a hash is a bit slower than checking the timestamp but far faster than parsing and byte compiling the source file, so I don't think this option is "significantly inert".
> The current Python pyc format is the marshaled code object of the module prefixed by a magic number [7], the source timestamp, and the source file size. The presence of a source timestamp means that a pyc is not a deterministic function of the input file’s contents—it also depends on volatile metadata, the mtime of the source. Thus, pycs are a barrier to proper reproducibility.
That is, they were made for a quite different use case than you or I were talking about.
I looked at the PEP to see if it gave timing numbers. No luck - would be a good blog post if I were still blogging. It does say:
> The hash-based pyc format can impose the cost of reading and hashing every source file, which is more expensive than simply checking timestamps. Thus, for now, we expect it to be used mainly by distributors and power use cases.
If you’re seeing this regularly, it suggests there may be something unique or uncommon in your set-up. You may wish to isolate and change whatever that is.
So if performances allow you to do so, disabling them when you dev is not a bad idea, as long as you keep them enabled when you run CI.
IMO, when you dev, you should set PYTHONHASHSEED (random is predictible), PYTHONDEVMODE (verbose warnings + tooling + sys.flags.dev_mode=True) and PYTHONDONTWRITEBYTECODE (remove the need for cleaning them).
In CI and prod, you should make sure those are NOT set, and use PYTHONOPTIMIZE=2 (remove assert, __debug__ lines and docstring) if you trust your dependancies to be well written (which I would check by running tests in a pre-push hook).
I guess the author did find cases where secrets.pyc is committed but secrets.py is not? It's hard to fathom how that could have happened (especially inside "organization" settings). Sounds like the result of absolute rookies in both Python and git following a tutorial with a step "add secrets.py to .gitignore" but unfortunately takes ignoring __pycache__ and ﹡.pyc for granted, which is too much to ask for some people.
> it is very easy for an experienced programmer to accidentally commit their secrets
No, it doesn't take an experienced programmer to put __pycache__ and ﹡.pyc to global ignore, or use a gitignore boilerplate at project creation, or notice random unwanted files during code review.
> never get to remove them from history.
Scrubbing specific files from git history isn't hard.
s/__pycache__ and ﹡pyc/secrets.py/g and people will also commit it in. PEBCAK.
Nonetheless, this is a common mistake, whether you believe it or not. And if it is common, then it will be exploited.
To be clear, they’re not even text. You don’t need to know Python at all to realize something’s not right when you’re committing unknown binaries to source control.
It's very easy to verify:
secrets.py:
import os
SECRET = os.getenv('SECRET')
Then $ python -m compileall secrets.py
$ uncompyle6 __pycache__/secrets.cpython-38.pyc
# uncompyle6 version 3.7.0
# Python bytecode 3.8 (3413)
# Decompiled from: Python 3.8.2 (default, Mar 10 2020, 12:58:02)
# [Clang 11.0.0 (clang-1100.0.33.17)]
# Embedded file name: secrets.py
# Compiled at: ...
# Size of source mod 2**32: 40 bytes
import os
SECRET = os.getenv('SECRET')
# okay decompiling __pycache__/secrets.cpython-38.pycSo a snippet like “os.environ[‘my_super_secret’]” won’t contain anything else than the bytecode to fetch that environment variable.
[1] https://github.com/radareorg/radare2/tree/master/libr/asm/ar...