Environment Variables Considered Harmful for Your Secrets
movingfast.io
movingfast.io
To address my second claim first: file permissions work at the user or group level. ACLs / MAC likewise. SELinux can be configured to assist in this case but it's not as trivial as it appears at first glance, it would be easier to use environment variables.
In the example case of spawning imagemagik, it's running as the same user and therefore has the same level of access to the properties file. That is, it can access the secrets without negotiating any authorisation to do so.
Depending on how imagemagik is launched and how the parent process handles the config file, it's possible that imagemagik could inherit a file handle, already open to the file.
Now to address my first claim, if the parent process is following best practice then it will sanitise the environment before exec'ing imagemagik, that should mean launching the imagemagik process with only the environment it needs.
To give a concrete example, the postfix mail transfer agent is extraordinarily high quality software, its spawn process owns the responsibility of launching external processes, potentially sysadmin supplied / external to postfix. This case would be very comparable to the web app invoking imagemagik.
We can see that it explicitly handles this case as I've suggested is best practice: https://github.com/vdukhovni/postfix/blob/master/postfix/src...
EDIT: accidentally posted before finishing.
If the parent sanitised the child's environment, then the only way for the child to access the data would be to read the parents memory. In practice this is quite easy - try the "ps auxe" command for a sample, however this access can much more easily be controlled by SELinux policy than can file access.
Any obfuscation technique applicable to a config file can similarly be applied to an environment variable.
I applaud to postfix for sanitising the ENV, and it's very good practice to do so. But are all the frameworks doing it correctly? Maybe some code is then also just spawning new processes without sanitising? You could argue that's a bug then (which I completely agree), but not all projects are run like postfix...
... mail delivery is done by daemon processes that have no parental
relationship with user processes. This eliminates a large variety of
potential security exploits with environment variables, signal handlers,
and with other process attributes that UNIX passes on from parent to child.
[edit] sorry for the spam, there was some problem with the submission form.This is an extension of the "don't put secrets in the query string" advice for HTTP: it's not that it's more secure: anything along the way that reads the query string can also read the post data just as easily, but that more bits of software may inadvertently leak the query string but not the post data (in logs, for instance).
Sure you can sanitize the environment, and make sure to wipe out sensitive data after reading it, but if that step fails for some reason, the only way you will notice is when you have a data leak. That's not a way to fail safe. You may the one to never make mistakes, but your coworkers aren't that perfect. That can only lead to security problems in the end.
When given the choice to store your secrets in regular config files or make up something with environment variables, choose the former. Doing it the expected way will buy you lots of goodness later on: You can use ACLs or policies to restrict access, you can allow just a limited number of reads, or pretty much anything else the VFS allows you to.
When you are not alone in your programming, do not inflict onto others what you wouldn't want inflicted on yourself. Avoid surprises.
The other kink is that the secret server itself reads the secrets from a symmetrically-encrypted file and when it boots, it doesn't actually know how to decrypt it. There's a master key for this that's stored GPG encrypted so that a small number of people can retrieve it and use a client tool that sends the secret server an "unlock" command containing the master key. So any time a secret server reboots, someone with access needs to gpg --decrypt mastersecret | secret_server_unlock_command someserver
There are some obvious drawbacks to this whole system (constraining pushes to require an SSH agent connection is a biggie and wouldn't fly some places, and agent forwarding is not without its security implications) and some obvious problems it doesn't solve (secrets are obviously still in RAM), but on the whole it works very well for distributing secrets to a large number of apps, and we have written tools that have basically completely eliminated any individual's need to ever actually lay eyes on a secret (e.g. if you want to run any tool in the mysql family, there's a tool that fetches the secret for you and spawns the tool you want with MYSQL_PWD temporarily set in the env, so you need not copy/paste it or be tempted to stick it in a .my.cnf).
> When an app on another server is started up, it must be done from a shell
That's a no-go for many setups. It doesn't integrate well with how Linux distros usually start services (systemd, upstart, sysv init, ...), and means you have to have another way to manage dependencies between your services.
> When an app on another server is started up, it must be done from a shell (we use cap) which has an SSH agent forwarded to it. In order for the app to get its database passwords and various other secrets, it makes a request to the secret server (over a TLS-encrypted socket), which checks your SSH identity against an ACL
At this point you could have used ssh right away, no? Any reason you used TLS + checking SSH agent instead?
> That's a no-go for many setups. It doesn't integrate well
> with how Linux distros usually start services (systemd,
> upstart, sysv init, ...)
Change the daemon config file to use a small wrapper script, which initializes the SSH environment and then execs the target binary. Assuming a reasonable setup, this should be trivial. > At this point you could have used ssh right away, no?
> Any reason you used TLS + checking SSH agent instead?
It sounds like they take an SSH identity certificate from the agent, send it via TLS, and then the remote process verifies it. This would have fewer potential security issues than trying to lock down a user's SSH login shell.Well, the point is that the ssh needs to have forwarded agent from somewhere else. If the host on which the service is run can initiate it, the whole security aspect is gone.
> This would have fewer potential security issues than trying to lock down a user's SSH login shell.
Locking down a login shell (usually be not running a shell in the first place) is a solved problem, and for example gitolite uses it has the base of its architecture. Yes, you have to be careful, but you must also be careful when manually validating certificates.
Yeah, using the SSH login method is actually quite slow for something you want to call at app startup on N instances during a push (at a minimum, your process responsible for whatever gatekeeping you do has to be respawned for every request, which necessarily puts a lower bound on the latency). I'm sure this could have been tracked down and optimized, but as jmillikin points out, another downside is that the additional per-user config can get kind of messy and error prone. Implementing logic like this at the .ssh/config level is (in my opinion) kind of easy to goof up and hard to test.
One of the interesting (and optional) things is does, is provide a agent to run on your instances that require the secrets, the agent implements a FUSE filesystem, and access to this filesytem is controlled by policy. For example - A policy can say "Allow exactly 1 read of /secrets/AWS.json within 120 seconds of boot". Any out of policy access attempts can cause the instance to be blacklisted, preventing any future secret access etc..
[1]: https://www.openstack.org/summit/portland-2013/session-video...
We've been using it to store our application secrets for some time and had no complaints.
It looks like process-private storage is one of its features.
Thanks for pointing this out.
Most (all?) of the current brand of UNIX variants have locked this down quite a while ago, which is a good thing. There are still a few old boxes kicking around though so if you're writing code that is meant to be widely deployed please don't put stuff there. For example: https://github.com/keithw/mosh/issues/156
Even if you are sure that your code will only be running on modern machines I think this article gives good advise. Unless you purge the secret environment variables when you launch they'll get sent to all of your subprocesses and it's quite possible that one of them won't consider their environment secret.
"Classic UNIX behavior was that environment variables were public (any user could see them with the right flags to "ps") so it was well-known not to put anything secret there."
The same logic applies to Windows environments and modern *nix too...
Storing private data in a public location is obviously a bad idea.
I don't get the reason for all that disagreement. Both methos have clear practival advantages and problems, and both are nearly as secure as the other. Use whatever fits better your problem.
But why should we jump through hoops for something that was broken over 15 years ago?
Do you still filter ping packets on your router because back in the '90s large pings would crash[1] many operating systems?
I also don't really buy the arguments for ENV storage of even non-sensitive data. There's just not really any good reason to do so; your config has to be put in place by some tool, even if it is in ENV; why not make your tool produce a file, with reasonable permissions in a well-defined default location? The 12 Factor App article seems to believe that config files live in your source tree, and are thus easily accidentally checked into revision control. That's not where my config files live. My config files live in /etc; or, if I want it to isolate a service to a user, I make a /home/user/etc directory.
One could say, "Don't store passwords in the revision control system alongside your source." And that would be reasonable. But, there's no reason to throw the baby out with the bath water.
The issue is that ultimately, one has to trust some infrastructure as being of an ultimately trusted network (whether it's Puppet/Chef/cfe2+ or offline authoritative CA's)
Env makes things really flexible.For instance,you can call a program with env variables directly thus overriding the default ones,which simplify configuring applications. You dont have to have a test set-up,a production-setup or a staging set-up,just start a server or an app with different env variables in the command line.
You don't want your config or keys to depend on an OS,or a language. Finally Env variables can be restricted to a set of users,so third party process started with a different one cant access them.
I believe env variables are better than other solutions.
Particularly this allows you to package a general configuration file, a user can provide an external configuration file and can for a specific instance even overwrite that with env (or command line arguments).
[1] http://docs.spring.io/spring-boot/docs/current/reference/htm...
ENV variables are not restricted by user though, your process can spawn another process under a different user and give it the same environment. It's the nature of the environment that it is usually inherited from the parent which causes the issues when we're talking about secrets.
Just reference a different config file on the command line.
>You don't want your config or keys to depend on an OS,or a language.
A path to a config file isn't an OS or language. The ini format is pretty widely supported.
>Finally Env variables can be restricted to a set of users,so third party process started with a different one cant access them.
This is exactly what filesystem permissions do for files already.
BEGIN {
$API = new Backend($ENV{credentials});
delete $ENV{credentials};
};
Filesystem permissions do not make it possible for a program to internally partition access to those credentials unless you (a) start it as root, or (b) delete the credentials file after reading it.Either of that seems to be able to be read straight from argv[0] of the process, so don't forget about cleaning that up.
Once an application has been written to get its secrets from the environment, there is a question of how the secrets are obtained. They can be sourced from a file in an init script, but today we are seeing a lot of momentum towards containerized architecture, and the use of service discovery and configuration systems like etcd, zookeeper and consul.
However, secrets require much more attention to security concerns than the data that these tools are designed to handle. Least privilege, separation of duties, encryption, and comprehensive audit are all necessary when dealing with secrets material. To this end, we have written a dedicated product which provides management and access for secrets and other infrastructure resources (e.g. SSH, LDAP, HTTP web services). The deployment model is similar to the HA systems provided by etcd, consul, SaltStack, etc. It's called Conjur (conjur.net).
On the distribution as a virtual appliance: is it a black box? How do I back it up? How do I upgrade it? How do I check the integrity of the secrets database?
In addition, the security team can make decisions about the management of the API key, such as:
* It can be kept off physical media (e.g.in /dev/shm)
* The process(es) by which the API keys are created and placed on the machines can be carefully managed
* A dedicated reaper/deprovisioner process can be used to retire (de-authorize) API keys once they go out of service.
Regarding the virtual appliance:
* Is it a black box? Essentially yes, although there are specific maintenance scripts on the machine (e.g. backup / restore) that you may occasionally run. For normal operation, only ports 443 and 636 (LDAPS) are open.
* How do I back it up? A standby master and read-only followers can be used to keep live copies of the database. A backup script can be used to capture full GPG-encrypted backups. A restore script will restore a backup onto a new appliance.
* How do I upgrade it? You create a standby master and followers of the new upgraded version of the appliance, and connect them to your current master. Once the standby master and followers are current with the data stream, they are promoted via DNS and serve the subsequent requests.
* How do I check the integrity of the secrets database? We have written an open-source Ruby interface to OpenSSL called Slosilo (https://github.com/conjurinc/slosilo) and subjected this library to a professional crypto audit. On the advice of the auditors, Slosilo encryption employs message authentication (cipher mode AES-256-GCM) which ensures the integrity of the secrets. Nevertheless, the Conjur appliances must obviously be afforded the highest level of protection. Unlike a multi-master system, Conjur's read-only "followers" contain only read-only copies of the secrets. So compromise of a follower cannot be used to modify secrets or authorization policies, or to confuse a master election scheme.
If you're shelling out to commands you think might snarf credentials, the environment is easy for them to pick it out of, but if they're running as the same user then they could probably read the secrets from the config file. If they aren't running as the same user, you need a way of passing in the secrets - and we tend to come back to environment variables..
The good practice here is just to reset the environment when calling shell commands, as he notes. It's not hard to do.
But there is one main difference: that tool would need to do so explicitly, with the intent of reading (and possibly exposing) your secrets. For me, that's a huge difference from having the secrets being implicitly available to the process through the processes environment.
Disadvantage, of course, is that you will have to update the list of trusted callers whenever they change. Mac OS X automates that by automatically trusting binaries signed with the same key (trades some security for convenience)
Another disadvantage is that this doesn't work well with scripts (the kernel service would have to know how to find the running scripts in order to checksum them, for every possible scripting language on the system). Also, any form of extension support in a trusted application is problematic.
I don't know of full open source equivalents. Parts of Apple's code are open source, though, for example http://opensource.apple.com/source/security_systemkeychain/s... (may not even compile; not all Apple's open source releases do)
If one absolutely needs to centralize secrets (TLS/SSL private keys, entropy sources, etc.) (at risk of SPOF or some HA setup), use some PSK style setup that delivers them directly, out-of-band (via separate NICs) or prioritized ahead of regular traffic. Keep it simple. Otherwise, prefer something like zookeeper with encrypted secrets (again PSK keying per box). Try to not deploy the same secret on every box, if possible. Also, try to avoid magic secrets if you can too (remove all local password hashes, use only auth keys).
If you're uncomfortable with plaintext secrets, encrypt them (as end-to-end as possible) and require an out-of-band decryption key at the last possible moment.
It's like having a secure document viewing system... ultimately, someone will need to browse just enough of the plaintext version, or it's not a document viewing system.
Neither files nor env variables.
Most chipsets have a rather unused TPM function, and it should be possible to have developers and processes hook into that.
Perhaps using tmptool ? On master process startup ask user for passphrase, and use that to query the TPM stored values ?
Cloud servers
=============
Xen supports virtual TPM.I'm no Amazon EC2 expert, but a quick google exposed a few keen souls who tried to use vTPM and failed. This would suggest that Amazon does not yet support vTPM.
Re-entering passphrases
========================
Well, unless the machine is permissioned by default you will need to give a fresh instance new authorization. Permissioning by default is the same security problem you're trying to avoid though... just shifted. Your overall goal is to have the credentials inaccessible to sniffing, right ?I guess you could set up some form of ssh-agent handshake to make the process less manual.
I would say that you should consider malicious sniffing too
I would propose to use just one folder like /secret and put your config files in there. Exclude this folder from backup on all relevant hosts.
Then spend your time on security of your hosts, applications (OWASP) and monitoring / alerting. Something that you have to do anyway.
why not just make use of the OS's secrets store? for example, like how https://pypi.python.org/pypi/keyring operates
Same thing while debugging : only the encrypted key is printed.
And instead of a secret key which is easily searchable, your method could just do some substitutions, something a bit more complicated than a Caesar cypher. Yes it's really weak but it beats an unencrypted secret key.
I know security minded people are not gonna like it, but until we have a real battle tested solution it's better than nothing.
A determined attacker will almost always win against our best defenses. I think we have to do our best to make their job hard, but at one point we have to accept that offense is really easier than defense.
It's been too long since I used it to remember the details, but I believe process-private keys are one of this API's features.
With pecl/hidef, I can hook a bunch of text .ini files into the PHP engine and defines constants as each requests comes in.
Originally, it was written for a nearly static website, which was spending a ridiculous amount of time defining constants, which rolled over every 30 minutes.
Plus those .ini files were only readable to root, which reads it and then forks off the unprivileged processes.
But with the hidef.per_request_ini, I could hook that into the Apache2 vhost settings, so that the exact same code would operate with different constants across vhosts without changing any code between them.
Used two different RPMs to push those two bits so that I could push code & critical config changes as a two-step process.
And with a bit of yum trickery (the yum-multiverse plugin), I could install two versions of code and one version of config, and other madness like that with the help of update-alternatives.
That served as an awesome A/B testing setup, to innoculate a new release with a fraction of users hitting a different vhost for the micro-service.
I'm rambling now, but the whole point is that you need per-request/per-user config overlay layers, for which env vars are horrible, slow and possibly printed everytime someone throws some diagnostics up.
We used Michael's trick: environment variables pointing to config files works unbelievably well if you ever need to implement a multiple tenant cloud offering.
So apart from the security aspect, there's the aspect that it is a more versatile design.
If you have software in your deployment that will send "error reports" to untrusted third parties then you have bigger problems than your shell environment.
2. The whole environment is passed down to child processes
If you don't trust your child processes then you have bigger problems than your shell environment.
3. External developers are not necessarily aware that your environment contains secret keys.
And?
I'm not sure what you mean by "external developer" and what you expect them to do with your environment. E-Mail it out when an error occurs?
If you tolerate that kind of developer on your project then you.. oh well, see above.
This has nothing to do with your choice of configuration method.
I think this article is a response to people's practice of keeping API keys as an environmental variable so as to keep them off of the filesystem (or at least what git sees and checks in) so that they don't accidentally publish them, as happened in that article article where some gem he was using to respect .gitignore didn't work for some reason.
would this work as a solution?
Our most common crash report scenario is the airbrake gem sending crash reports from our rails app to errbit. I can confirm airbrake gem does not post any sensible env data.
Of course, this is only a good news for that specific case, others apps may transfer environment, and we can't just wonder for each app installed "what will it ever send ?".
• Configuration files can be trivially forged (rsa without signing)
• Many configuration files can be trivially decrypted without the key (rsa use on potentially big files)
• Keys can be trivially recovered in multiple ways (global variable, ptrace)
Users might be tricked into using your library despite the fact it offers them no real security except false security.
I recommend placing a warning that it is not intended to be used by anyone at any time.
* Preventing forgery/tampering is an orthogonal issue and can be done by intrusion-detection software.
* "trivially decrypted without the key" -- How? The size of the file is not going to make any difference here.
* Of course, the key can be recovered if you have the ability to ptrace. What's your point? That's going to be true of any solution out there.
Why don't you make some constructive criticism instead of being obnoxious.
This is a job for people who take programming seriously, not amateurs who can't be bothered to read the documentation of the modules they are using, e.g.
http://stuvel.eu/files/python-rsa-doc/usage.html
Note especially the bits about signing RSA, and how to use it to encrypt files. Doing raw RSA on strings longer than 245 bytes is unwise, and inventing new protocols for RSA is unwise. That the author does these things is strong evidence that the author is not demonstrating "great care".
If the author believes forgery/tampering is orthogonal and/or protection from in-process attacks, then the author should state that. That's part of the analysis step.
> Of course, the key can be recovered if you have the ability to ptrace. What's your point? That's going to be true of any solution out there.
Nonsense. If you delete the key from memory, it is obvious that someone cannot use ptrace to recover it afterwards. Detailing what is at risk, and what is not at risk is part of the analysis step.
The first section of his documentation states what the purpose of the module is and what's it is trying to protect against. The use of RSA there is pretty reasonable, so I don't know why you're being so hard on him. It also talks about how this project is a variation on another project, so it's not like he is going off half-cocked and implementing something crazy.
> This is a job for people who take programming seriously, not amateurs
Everyone starts out as an amateur. You can either be an internet know-it-all that wants to show off how little you know or you can provide constructive feedback.
> who can't be bothered to read the documentation of the modules they are using, e.g. http://stuvel.eu/files/python-rsa-doc/usage.html > Note especially the bits about signing RSA, and how to use it to encrypt files. Doing raw RSA on strings longer than 245 bytes is unwise, and inventing new protocols for RSA is unwise. That the author does these things is strong evidence that the author is not demonstrating "great care".
Did you bother to read the documentation yourself? It directly talks about encrypting large files. Yes, it's true that trying to encrypt more data than is supported for a key size will leak information, which is why implementations throw an error if you try to do that. The usage documentation you pointed to discusses that and provides options to cope with it. But, in this case, he is not even encrypting the whole file, just the values of certain properties, which are probably going to be small enough to not require any extra measures.
And, again, what does signing have to do with how RSA is being used in greybox?
> If the author believes forgery/tampering is orthogonal and/or protection from in-process attacks, then the author should state that. That's part of the analysis step.
They did state the scope of the project and there is really no need for them to go into deep detail about deployment. Are you going to criticize him for not pointing out that the whole system needs to be secured from tampering? Of course not, so stop acting like a douchebag.
>> Of course, the key can be recovered if you have the ability to ptrace. What's your point? That's going to be true of any solution out there. > Nonsense. If you delete the key from memory, it is obvious that someone cannot use ptrace to recover it afterwards. Detailing what is at risk, and what is not at risk is part of the analysis step.
You have to ask, where did the key come from in the first place? The private key file has to exist somewhere that is accessible on the box so that the process can read it in the first place. If someone has ptrace, they can probably read a file as well.
Ultimately, it's about the level of risk people are willing to live with. For what greybox seems targeted at, it's encrypting properties in files that are committed to a repo and then using a private key that is only available on certain machines to read that property. For some use cases, that is probably a perfectly fine level of security and better than what is being done in some cases anyways.
It is sensible to type it in while the system is still in single-user mode.
If you left the private key in a file that is right next to your encrypted configuration file, then it remains as I stated previously: No security except false security.
> Everyone starts out as an amateur.
This is not the forum for an education on programming. A library that states it was written as a learning exercise with a request for criticism will be treated that way. A library proposed to solve problems recognised in the linked article will not.
> Ultimately, it's about the level of risk people are willing to live with.
People are notoriously bad at recognising and evaluating risk.
Shut up and read the goal of the project already. It is trying to protect secrets that are stored in a repo and it achieves that goal.
Secrets on a deployed host will always be in the clear on that host, it's just a fact of life. Many barriers can be put in place, but at the end of the day, a program will always need access to the plaintext version of the secret at some point.
> It is sensible to type it in while the system is still in single-user mode.
No it isn't. Services have to restart all the time, saying that a human has to be ready to type in a password at any moment is not practical.
> This is not the forum for an education on programming.
But it is a forum for you to post invalid criticisms of a project and act like a dick? That's bullshit. This is "hacker news", it's a perfectly fine place for a technical discussion. If you thought there were problems with the project and this wasn't the right place to discuss it, then file issues on the github project.
> A library that states it was written as a learning exercise with a request for criticism will be treated that way. A library proposed to solve problems recognised in the linked article will not.
There is no difference, you're just trying to justify your bad behavior.
wouldnt that, in turn, also be stored in a code repository, likely accessible in the same way as the main coe repo? then, this feels like a non-solution to me.
Alternatively, if you're running on AWS you could also fetch the secrets config file from an S3 bucket which is only accessible by your production servers.
Mreinsch's s3 reco is a good example. I use this method for storing extra role secrets for AWS.
It's good to keep in mind the words Morpheus when dealing with Chef(and all this stuff really). Free your mind.
Now, of course no one would install and run this... but I could imagine someone accidentally typing the name of a Gem wrong, someone accepting a bad PR (a sub-dependancy perhaps even doing so?), etc and somehow something untrusted getting in there. Yes, that means you have other problems, but it isn't outside the realm of possibility that accidental access like this is had.
Just because it shouldn't happen, doesn't mean it will never happen.