Real-world stories of how we’ve compromised CI/CD pipelines
research.nccgroup.com
research.nccgroup.com
Some thoughts:
1. Hardcoded credentials are a plague. You should consider tagging all of your secrets so that they're easier to scan for. Github automatically scans for secrets, which is great.
2. Jenkins is particularly bad for security. I've seen it owned a million and one times.
3. Containers are overused as a security boundary and footguns like `--privileged` completely eliminate any boundary.
4. Environment variables are a dangerous place to store secrets - they're global to the process and therefor easy to leak. I've thought about this a lot lately, especially after log4j. I think one pattern that may help is clearing the variables after you've loaded them into memory.
Another I've considered is encrypting the variables. A lot of the time what you have is something like this:
Secret Store -> Control Plane Agent -> Container -> Process
Where secrets flow from left to right. The control plane agent and container have full access to the credentials and they're "plaintext" in the Process's environment.
In theory you should be able to pin the secrets to that process with a key. During your CD phase you would embed a private key into the process's binary (or a file on the container) and then tell your Secret Manager to use the associated public key to transmit the secrets. The process could decrypt those secrets with its private key but they're E2E encrypted across any hops between the Secret Store and Process and they can't be leaked without explicitly decrypting them first.
The two real problems with environment variables are:
1. Environment variables are traditionally readable by any other process in the system. There are settings you can do on modern kernels to turn this off, but how do you know that you will always run on such a system?
2. Environment variables are inherited to all subprocesses by default, unless you either unset them after you fork() (but before you exec()), or if you take special care to use execve() (or similar) function to provide your own custom-made environment for the new process.
I think that this would require being the same user as the process you're trying to read. Access to proc/pid/environ should require that iirc. You can very easily go further by restricting procfs using hidepid.
And ptrace restrictions are pretty commonplace now I think? So the attacker has to be a parent process or root.
> 2. Environment variables are inherited to all subprocesses by default, unless you either unset them after you fork() (but before you exec()), or if you take special care to use execve() (or similar) function to provide your own custom-made environment for the new process.
Yeah, this goes to my "easy to leak" point.
Either way though you're talking about "attacker has remote code execution", which is definitely worth considering, but I don't think it matters with regards to env vs anything else.
Files suffer from (1), except generally worse. File handles suffer from (2) afaik.
Embedding the private key into the binary doesn't help too much if the attacker is executing with the ability to ptrace you, but it does make leaking much harder ie: you can't trick a process into dumping cleartext credentials from the env just by crashing it.
IIRC, this was not always the case. But fair enough, this might not be a relevant issue for any modern system.
There is a really good article that explains a different way of securing these systems though sets of attestations.
https://grepory.substack.com/p/der-softwareherkunft-software...
The article you linked to is about signing. It doesn't solve "I need to put an AWS key into the environment of a process".
The answer is very much, 'it depends'. For oen thing, developers can run whatever code in CI before it's benn reviewed. I could just nab the env vars and post them wherever. If there are no sensitive env vars for me to nab and you have enforced code review, then I need a co-conspirator, and my change is probably going to leave a lot more of a paper trail.
Another risk is accidental disclosure - I have on at least two occasions accidentally logged sensitive environment variables in our CI environment. Now your threat model is not just a malicious developer pushing code - it's a developer making a mistake, plus anyone with read access to the CI system.
I don't know about your org, but at my job, the set of people who have read access to CI is a lot larger than the set who can push code, which is again a lot larger than the set of people who can merge code without a reviewer signing off.
> but drawing these boundaries seems like a nightmare for everyone...
As someone currently struggling with how to draw them, yup.
Yeah I don't think this gets talked about enough.
If you're talking about private repos in an organization then CI often runs on any pull request. That means a developer is able to make CI run in an unreviewed PR. Of course for it to make its way into a protected branch (main, etc.) it'll likely need a code review but nothing is stopping that developer who opened the unreviewed PR to modify the CI yaml file in a commit to make that PR's pipeline do something different.
Requiring a team lead or someone to allow every individual PR's pipeline to run (what GitHub does by default in public repos) would add too much friction and not all major git hosts support the idea of locking down the pipelines file by decoupling it from the code repo.
Edit: Depending on which CI provider you use, this situation is mostly preventable -- "mostly" in the sense that you can control how much damage can be done. Check out this comment later in this thread: https://news.ycombinator.com/item?id=29967077
If a developer changes the CI pipeline file to make their PR's code run in `deployment: "production"` instead of `deployment: "test"` doesn't that bypass this?
Edit:
I'll leave my original question here because I think it's an important one but I answered this myself. It depends on which CI provider you're using but some of them do let you restrict specific deployments from being run only on specific branches or by specific folks (such as repo admins).
In the above case if the production deployment was only allowed to run on the main branch and the only way code makes its way into the main branch is after at least 1 person reviewed + merged it (or whatever policy your company wants) then a rogue developer can't edit the pipeline in an unreviewed PR to make something run in production.
Also with deployment specific environment variables then a rogue developer is also not able to edit a pipeline file to try and run commands that may affect production such as doing a `terraform apply` or pushing an unreviewed Docker image to your production registry.
CI should never ever have access to anything related to production; not just for security but also to prevent potentially bad code being run in tests from trashing production data.
As a concrete example, GitLab has the concept of protected branches and code owners, both of which allow you to restrict access to the corresponding environments’ credentials to a smaller group of people who have permission to touch the sensitive branches. That allows you to say things like “anyone can run in development but only our release engineers can merge to staging/production” or “changes to the CI configuration must be approved by the DevOps team”, respectively.
That does, of course, not prevent someone from running a Bitcoin miner in whatever environment you use to run untrusted merge requests but that’s better than access to your production data.
For private repositories, that means access to credentials. Probably read-only credentials, but it requires network access.
Would you be suggesting that everyone should commit all dependencies?
I'm a proponent of lock-file approaches, which gain 99% of the benefits with far less pain. It requires network access, though.
However, it's problematic. Only use it if you're certain it will solve a specific problem you have.
Consider using Docker to build images that include a snapshot of node_modules.
I agree that there are tokens and variables that are dangerous to expose via CI, but throwing the baby out with the bathwater confused me.
1. stand up a private package mirror that you control, that uses a whitelist for what packages it is willing to mirror;
2. configure your project's dependency-fetching logic to fetch from said mirror;
3. configure CI to only allow outbound network access to your package mirror's IP.
The disadvantage — but also the point — of this, is that it is then a release manager's responsibility, not a developer's responsibility, to give the final say-so for adding a dependency to the project (because only the release manager has the permissions to add packages to the mirror.)
I guess that's beside the point if your goal is only to reduce risk of compromised CI/CD.
Every CI system I've ever seem has pulled dependencies in from the network.
git clone --recursive --branch "$commitid" "$repourl" "$repodir"
img="$(docker build --network=none -f "$dockerfile" "$repodir")"
docker run --rm -ti --network=none "$img"
Sure, CI pulls in from the network... but execution occurs without network.First:
> git clone --recursive --branch "$commitid" "$repourl" "$repodir"
The `git clone` will take your git repository URL as the $repourl variable. It will also take your commit id (commit hash or tagged version which a pull request points to) as the $commitid variable (`--branch $commitid`). It will also take a $repodir variable which points to the directory that will contain the contents of the cloned git repository and already checked out at the commit id specified. It will do so recursively (`--recursive`): if there are submodules then they will also automatically be cloned.
This of course assumes that you're cloning a public repository and/or that any credentials required have already been set up (see `man git-config`).
Then:
> img="$(docker build --network=none -f "$dockerfile" "$repodir")"
Okay so this is sort've broken: you'd need a few more parameters to `docker build` to get it to work "right". But as-is, `docker build` usually has network access so `--network=none` will specify that the build process will not have access to the network. I hope your build system doesn't automatically download dependencies because that will fail (and also suggests that the build system may be susceptible to attack). You specify a dockerfile to build using `-f "$dockerfile"`. Finally, you specify the build context using "$repodir" -- and that assumes that your whole git repository should be available to the dockerfile.
However, `docker build` will write a lot more than just the image name to standard output and so this is where some customization would need to occur. Suffice to say that you can use `--quiet` if that's all you want; I do prefer to see the output because it normally contains intermediate image names useful for debugging the dockerfile.
Finally:
> docker run --rm -ti --network=none "$img"
Finally, it runs the built image in a new container with an auto-generated name. `-ti` here is wrong: it will attach a standard input/output terminal and so if it drops you into an interactive program (such as bash) then it could hang the CI process. But you can remove that. It also assumes that your dockerfile correctly specifies ENTRYPOINT and/or CMD. When the container has exited then the container will automatically be removed (--rm) -- usually they linger around and pollute your docker host. Finally, the --network=none also ensures that your container does not have network access so your unit tests should also be capable of running without the network or else they will fail. You could use `--volume` to specify a volume with data files if you need them. You might also want to look at `--user` if you don't want your container to have root privileges...
And of course if you want integration tests with other containers then you should create a dedicated docker network and specify its alias with `--network`: see `man docker-network-create`; you can use `docker network create -d internal` to create a network which shouldn't let containers out.
Does that answer your question?
It's been my experience that usually pull requests which make a library capable of being built offline are welcome. Usually. So get off your duff and get to fixing the problems in the free packages you use.
Luxury! When I were a youngun our devnet was airgapped and we had to sneakernet dependencies to it. Using vhs tape on a 9track spool. Because real tape was too expensive.
More seriously I've seen jenkins and gitlab pipelines with only access to sneakernet maintained mirrors.
However, it'd probably be better if you could have the CI framework collect and inject that information into the build using some hard-coded deterministic logic, rather than giving the build itself (developer-driven Arbitrary Code Execution) access to that capability.
Same idea as e.g. injecting Kubernetes Secrets into Pods as env-vars at the controller level, rather than giving the Pod itself the permission to query Secrets out of the controller through its API.
I'm failing to understand how that procedure even works. How do you run the tests?
It's sort of like how I tell my mom, "Even you don't want to know your passwords", when explaining that she should use a password manager. The reason in this case is because she may get fished and redirected to a fake website. If there's a password manager, and she doesn't even know her own password, then it won't recognize the certificate and she's safe. However, if she does know it, then she has the ability to leak it.
This is the same with environment variables, and a reason why secret storage is preferred over them. Think of environment variable as your mom knowing her password, and secret storage to her using a password manager. More or less haha.
Maybe. Back in the old days if you had the commit bit your badge didn’t get you into the server room. I get the impression a lot of shops are effectively giving their devs root but in the cloud this time, which isn’t necessary.
Queues map pipelines to agents. Agents can be assigned IAM roles. If you want a certain build to run as an IAM role, you give it a queue where the agents have that role. For AWS, Buildkite has as a Cloud Formation stack that sets up auto scaling groups and some other resources for your agents to run.
Basically it adds a signed web identity file into the container which can be used to assume roles.
— They claim that their solution has the same isolation level ("4 stars") than gVisor, unlike "standard containers", which are "2 stars" only (with Firecracker and Kubevirt being "5 stars). This is very wrong - as far as I can tell, they use regular Linux namespaces with some light eBPF-based filesystem emulation, while the vast majority of syscalls is still handled by the host kernel. Sorry, but this is still "2 stars" and far away from the isolation guarantees provided by gVisor (fully emulating the kernel in userspace, which is at the same level or even better than Firecracker) and nowhere close to a VM.
— Somehow, regular VMs (Kubevirt) get a "speed" rating of only "2 stars" - worse than gVisor ("3 stars") and Firecracker ("4 stars"), even though they both rely on virtually the same virtualization technology. If anything, gVisor is the slowest but most efficient solution while QEMU maintains some performance advantage over Firecracker[2]. These are basically random scores, it's not a good first impression–if you do a detailed comparison like that, at least do a proper evaluation before giving your own product the best score!
— They claim that "standard containers" cannot run a full OS. This isn't true - while it's typically a bad idea, this works just fine with rootless podman and, more recently, rootless docker. Allowing this is the whole point of user namespaces, after all! Maybe their custom procfs does a better job of pretending to be a VM - but it's simply false that you can't do these things without. You can certainly run a full OS inside Kata/Firecracker, too, I've actually done that.
Nitpicking over rating scales aside, the claim that their solution offers large security improvements over any other solution with user namespaces isn't true and the whole thing seems very marketing-driven. The isolation offered by user namespaces is still very weak and not comparable to gVisor or Firecracker (both in production use by Google/AWS for untrusted workloads!). False marketing is a big red flag, especially for something as critical as a container runtime.
Anyone who wants unprivileged system containers might want to look into rootless docker or podman rather than this.
[2]: https://www.usenix.org/system/files/nsdi20-paper-agache.pdf
For the CI/CD usecase on AWS, sysbox presented the right balance of trade-offs between something like Firecracker (which would require bare metal hosts on AWS) and the docker containers that already existed. We specifically need to run privileged containers so that we could run docker-in-docker for CI workloads, so rootless docker or podman wouldn't have helped. Sysbox lets us do that with a significant improvement in security to just running privileged docker containers as most CI environments end up doing.
Just switching their docker-in-docker CI job containers to sysbox would have mitigated 4 of the compromises from the article with nearly zero other configuration changes.
rootless docker works inside an unprivileged container (that's how our CI works).
https://docs.docker.com/engine/security/rootless/#rootless-d...
> Anyone who wants unprivileged system containers might want to look into rootless docker or podman rather than this.
Perhaps I'm missing something, but I have been running full OS userlands using "standard containers" in production for years, via LXD[1].
> By default containers are unprivileged […]
https://linuxcontainers.org/lxd/docs/master/security/#contai...
As for LXC:
> LXC containers can be of two kinds:
> - Privileged containers
> - Unprivileged containers
> […]
> The latter has been introduced back in LXC 1.0 (February 2014) […]
- Regarding the container isolation, Sysbox uses a combination of Linux user-namespace + partial procfs & sysfs emulation + intercepting some sensitive syscalls in the container (using seccomp-bpf). It's fair to say that gVisor performs better isolation on syscalls, but it's also fair to say that by adding Linux user-ns and procfs & sysfs emulation, Sysbox isolates the container in ways that gVisor does not. This is why we felt it was fair to put Sysbox at a similar isolation rating as gVisor, although if you view it from purely a syscall isolation perspective it's fair to say that gVisor offers better isolation. Also, note that Sysbox is not meant to isolate workloads in multi-tenant environments (for that we think VM-based approaches are better). But in single-tenant environments, Sysbox does void the need for privileged containers in many scenarios because it allows well isolated containers/pods to run system workloads such as Docker and even K8s (which is why it's often used in CI infra).
- Regarding the speed rating, we gave Firecracker a higher speed rating than KubeVirt because while they both use hardware virtualization, the latter run microVMs that are highly optimized and have much less overhead that full VMs that typically run on KubeVirt. While QEMU may be faster than Firecracker in some metrics in a one-instance comparison, when you start running dozens of instances per host, the overhead of the full VM (particularly memory overhead) hurts its performance (which is the reason Firecracker was designed).
- Regarding gVisor performance, we didn't do a full performance comparison vs. KubeVirt, so we may stand corrected if gVisor is in fact slower than KubeVirt when running multiple instances on the same host (would appreciate any more info you may have on such a comparison, we could not find one).
- Regarding the claim that standard containers cannot run a full OS, what the table in the GH repo is indicating is that Sysbox allows you to create unprivileged containers (or pods) that can run system software such as Docker, Kubernetes, k3s, etc. with good isolation and seamlessly (no privileged container, no changes in the software inside the container, and no tricky container entrypoints). To the best of our knowledge, it's not possible to run say Kubernetes inside a regular container unless it's a privileged container with a custom entrypoint. Or inside a Firecracker VM. If you know otherwise, please let us know.
- Regarding "The claim that their solution offers large security improvements over any other solution with user namespaces isn't true". Where do you see that claim? The table explicitly states that there are solutions that provide stronger isolation.
- Regarding "The isolation offered by user namespaces is still very weak and not comparable to gVisor or Firecracker". User namespaces by itself mitigates several recent CVEs for containers, so it's a valuable feature. It may not offer VM-level isolation, but that's not what we are claiming. Furthermore, Sysbox uses the user-ns as a baseline, but adds syscall interception and procfs & sysfs emulation to further harden the isolation.
- "False marketing is a big red flag, especially for something as critical as a container runtime." That's not what we are doing.
- Rootless Docker/Podman are great, but they work at a different level than Sysbox. Sysbox is an enhanced "runc", and while Sysbox itself runs as true root on the host (i.e., Sysbox is not rootless), the containers or pods it creates are well isolated and void the need for privileged containers in many scenarios. This is why several companies use it in production too.
> It's fair to say that gVisor performs better isolation on syscalls, but it's also fair to say that by adding Linux user-ns and procfs & sysfs emulation, Sysbox isolates the container in ways that gVisor does not.
Have a look at what gVisor actually does: https://gvisor.dev/docs/architecture_guide/security
It fully implements a subset of the Linux kernel ABI in userspace, including procfs and sysfs and even memory and process management. No untrusted code ever interacts with the host kernel. Filesystem and network access goes through an IPC protocol and is handled by the gVisor processes on the host, which in turns runs inside a user namespace and a seccomp sandbox for defense in depth.
This is a much, much stronger level of isolation than your approach or, arguably, even VMs (the trade-off is performance). "Sysbox isolates the container in ways that gVisor does not" just isn't true.
The sysbox approach is one kernel bug away from host system compromise, same as using regular containers. Emulating procfs and sysfs and using user namespaces takes away some of the attack surface and is great defense in depth, but does not provide isolation from the host kernel.
> Also, note that Sysbox is not meant to isolate workloads in multi-tenant environments (for that we think VM-based approaches are better)
I've read numerous claims that sysbox is suitable for untrusted workloads, for instance in [1] and [2].
It's a nice product and certainly much, much better than running docker-in-docker using privileged containers, but given the significant remaining attack surface, this claim could put your customers at risk and should come with a big disclaimer.
> While QEMU may be faster than Firecracker in some metrics in a one-instance comparison, when you start running dozens of instances per host, the overhead of the full VM (particularly memory overhead) hurts its performance (which is the reason Firecracker was designed)
Firecracker was designed for memory efficiency, faster cold start times and security (by virtue of being written in a memory-safe language). It means you can run more containers per host, but the actual workload performance overhead is identical to "normal" VMs and, in some cases, even slightly higher since Firecracker lacks some of the optimization that has gone into QEMU.
> Regarding gVisor performance, we didn't do a full performance comparison vs. KubeVirt, so we may stand corrected if gVisor is in fact slower than KubeVirt when running multiple instances on the same host (would appreciate any more info you may have on such a comparison, we could not find one).
KubeVirt is just plain QEMU VMs using libvirt, which have been compared to gVisor quite extensively[3][4]. There's almost no overhead for memory/CPU and quite a lot of overhead for syscalls (but with big improvements recently with the introduction of VFS2 and soon LisaFS[5]). It's a classic trade-off - gVisor is more secure and efficient than QEMU, allowing a much larger number of instances to run on a host by virtue of better cooperation with the host kernel scheduler and memory management, but for raw performance, a QEMU VM always wins.
> Regarding the claim that standard containers cannot run a full OS, what the table in the GH repo is indicating is that Sysbox allows you to create unprivileged containers (or pods) that can run system software such as Docker, Kubernetes, k3s, etc. with good isolation and seamlessly (no privileged container, no changes in the software inside the container, and no tricky container entrypoints). To the best of our knowledge, it's not possible to run say Kubernetes inside a regular container unless it's a privileged container with a custom entrypoint. Or inside a Firecracker VM. If you know otherwise, please let us know.
Firecracker runs a full Linux kernel inside the VM, so it could always run regular Docker, Kubernetes or anything else. See [6] for a practical example.
For containers, this used to be the case, but the situation improved in recent kernel releases.
For podman, almost every combination works - running systemd unprivileged, running podman inside podman, or even running rootless-podman-in-rootless-podman[7] and so does Kubernetes-in-rootless-{podman,docker}[8] (requiring very recent kernel features, though - notably cgroupsv2 and unprivileged overlayfs).
Running docker:dind-rootless inside unprivileged Docker containers also works, however, it requires "--security-opt seccomp=unconfined".
Sysbox definitely got to that point earlier and has better usability.
> - Regarding "The claim that their solution offers large security improvements over any other solution with user namespaces isn't true". Where do you see that claim? The table explicitly states that there are solutions that provide stronger isolation.
Apologies, then, for misinterpreting that.
[1]: https://blog.nestybox.com/2020/10/06/related-tech-comparison...
[2]: https://github.com/nestybox/sysbox/issues/120#issuecomment-9...
[3]: https://object-storage-ca-ymq-1.vexxhost.net/swift/v1/6e4619...
[4]: https://www.scitepress.org/Papers/2021/104405/104405.pdf
[5]: https://gvisor.dev/blog/2021/12/02/running-gvisor-in-product...
[6]: https://github.com/innobead/kubefire
[7]: https://www.redhat.com/sysadmin/podman-inside-container
> Have a look at what gVisor actually does
I am aware of what it does, though I had missed the fact that the Sentry and/or Gopher run within a user-ns (could not find this in the docs). Had also missed the fact that it does perform procfs/sysfs emulation (makes sense), so I stand corrected on that. In light of this, I'll modify the Sysbox GH table to show gVisor as having a stronger isolation rating (in fact, our Sysbox blog comparing technologies [1] did give gVisor a stronger isolation rating).
> the sysbox approach is one kernel bug away from host system compromise
All approaches are one bug away from host system compromise (gVisor, VMs, etc.), though I agree that approaches like gVisor and VMs have a reduced attack surface.
> I've read numerous claims that sysbox is suitable for untrusted workloads
It's not a black or white determination in my view. Users choose based on their environments & needs. We always make it clear to our users that VM-based approaches provide stronger isolation, per the Sysbox GH repo:
"Isolation wise, it's fair to say that Sysbox containers provide stronger isolation than regular Docker containers (by virtue of using the Linux user-namespace and light-weight OS shim), but weaker isolation than VMs (by sharing the Linux kernel among containers)."
> Firecracker runs a full Linux kernel inside the VM, so it could always run regular Docker, Kubernetes or anything else
That's good to know (thanks), though the table in the Sysbox GH repo meant to compare Sysbox against Kata + Firecracker (since Kata is a container runtime). To the best of my knowledge running Docker, K8s, k3s, etc. inside a Kata container is not easy (see [1] and [2]).
> For containers, this used to be the case, but the situation improved in recent kernel releases.
It's correct that rootless docker/podman approaches are improving as far as what workloads they can run inside containers, although they still have several limitations [3], [4].
With Sysbox, most of these limitations don't apply because the solution works at the more basic "runc" level, Sysbox itself is rootful, and it uses some of the techniques I mentioned before (user-ns, procfs & sysfs virtualization, syscall trapping, UID-shifting, etc.) to make the container resemble a "real host" while providing good isolation.
Good discussion, please let me know of any more feedback.
[1] https://github.com/kata-containers/kata-containers/issues/20... [2] https://github.com/daniel-noland/docker-in-kata [3] https://docs.docker.com/engine/security/rootless/#known-limi... [4] https://github.com/containers/podman/blob/main/rootless.md
After that podman and buildah have gotten a lot of great reviews from people so I think they're awesome.
For an old time Unix sysadmin it just doesn't make sense to run something as root unless you absolutely have to.
Which also makes the client excuse in the article so strange, they had to run the container privileged to run static code analysis. wtf. Doesn't that just mean they run a tool against a binary artefact from a previous job? I fail to see how that requires privileges.
That also eliminates the risk of accessing the host docker.
We had this "deploy" Jenkins box set up with limited access for devs, because it had assume-role privs to an IAM role to manage AWS infra with Terraform. The devs run their tests on a different Jenkins box, and when they pass, they upload artifacts to a repo and trigger this "deploy" Jenkins box to promote the new build to prod. The devs can do their own CI, but CD is on a box they don't have access to, hence less chance for accidental credential leakage. Me being Mr. Devops-play-nice-with-the-devs, I let them issue PRs against the CD box's repo. Commits to PRs get run on the deploy Jenkins in a stage environment to validate the changes.
This one dev wanted to change something in AWS. But for whatever reason, they didn't ask me (maybe because they knew I'd say no, or at least ask them about it?). So instead the dev opens a PR against the CD jobs, proposing some syntax change. Then the dev modifies a script which was being included as part of the CD jobs, and makes the script download some binaries and make AWS API calls (I found out via CloudTrail). Once they've made the calls, they rewrite Git history to remove the AWS API commits and force-push to the PR branch, erasing evidence that the code was ever issued. Then close the PR with "need to refactor".
In the morning I'm looking through my e-mail, and see all these GitHub commits with code that looks like it's doing something in AWS... and I go look at the PR, and the code in my e-mails isn't anyware in any of the commits. He actually tried to cover it up. And I would never have known about any of this if I hadn't enabled 'watching' on all commits to the repo.
Who'd have thought e-mail would be the best append-only security log?
I was actually fired early in my career as a contractor when an over-zealous security big-wig decided to go over my boss's boss's head. I had punched a hole in the firewall to look at Reddit, and because I also had a lot of access, this meant I wasn't trustworthy and had to go. People (like me) make stupid mistakes; we should give them a second chance.
I've maybe managed to explain this process to one other extant employee, so pretty much everybody bugs me or one of the operations people any time there's an issue. That could be a liability in an outage situation, but I don't have a concrete suggestion how to avoid this sort of thing.
I've seen few things get engineer pushback quite like trying to tell engineers that they need to rework how they build and deploy because someone outside their team said so. It's just dev, not production, so why should they be so paranoid about it? Sheesh, stop screwing up their perfectly good workflows...
GitHub because it’s better UX. It’s even quite simple to setup good automation around a codebase.
Platform teams are using Argo, dev teams not really doing too much ci/cd which I like.
& to be honest CI/CD requires continuous investment as things continuously change. Not that it isn’t necessary… but in an enterprise environment you I’ve seen teams become more successful on their own rather than trying to fulfill any “reciprocity” bs.
I reckon this has to do with how the CI tools are configured.
Everyone knows you shouldn't commit a secret to Git, so tools like GitLab CI which require all their config be in git naturally will see less of this specific issue.
Jenkins is a batteries excluded pattern in one of its worst possible incarnations.
Jenkins is basically a CI framework for trusted users only. Untrusted workloads must not have access to anything Jenkins.
Global find on some terms like "key", "password" etc were great fun. It really showed most people, our team included, struggled with getting the pipeline to work at all. Let alone doing it in a secure manner.
This is a 50k+ employee financial institute. I am honestly surprised these kind of attacks are not much more widespread.
To fix this - almost anywhere - stop using shared secrets. Every time you visit a (HTTPS) web site, you are provided with the credentials to verify its identity. But, you don't gain the ability to impersonate the site because they're not secret credentials, they're public. You can and should use this in a few places in typical CI / CD type infrastructure today, and we should be encouraging other services to enable it too ASAP.
In a few places they mention MFA. Again, most MFA involves secrets, for example TOTP Relying Parties need to know what code you should be typing in, so, they need the seed from which to generate that code, and attackers can steal that seed. WebAuthn doesn't involve secrets, so, attackers who steal WebAuthn credentials don't achieve anything. Unfortunately chances are you enabled one or more vulnerable credential types "just in case"...
A secret value ought to be very carefully guarded even from the host machine itself.
.NET for example has SecureString, which is a good start — it can’t be accidentally printed or serialised insecurely. If it is serialised, then it is automatically encrypted by the host OS data protection API.
Windows even has TPM-hosted certificates! They’re essentially a smart card plugged into the motherboard.
A running app can use a TPM credential to sign requests but it can’t read or copy it.
These advancements are just completely ignored in the UNIX world, where everything is blindly copied into easily accessible locations in plain text…
What planet are you from and can I go there?
SecureString provides one layer in the "defence in depth". If someone accidentally logs it, then it won't leak. If it is used as a script parameter, then the prompt will use password input characters automatically.
Seriously, I want you to try this PowerShell snippet right now. (You can even run it on Linux too with pwsh v7, so no excuses!):
PARAM(
[parameter(Mandatory=$true)]
[securestring]$Secret
)
Write-Warning "Oops I didn't mean to log this: $Secret"
$Secret | ConvertTo-Json -Compress
$Secret | Export-Clixml 'accidental serialisation.xml'
Get-Content 'accidental serialisation.xml'
Run the the script and see what happens.There's even low-level protection built-in, such zeroing out the memory when it is garbage collected (unlike normal strings).
Why is this a bad thing!?
Do you like secrets leaking everywhere unless everyone is always hyper vigilant? Or do you prefer to roll your own half-baked secret storage type that nothing else is compatible with?
PS: The page with that advice is based on an "archived" read-only repo with a bunch of open issues of people befuddled as to why this bad, BAD, BAD advice is being published there.
Secrets are awesome part of pwsh. You also have `Get-Secret` and friends for full blown pass-like solution.
One of the many reasons I detest stringly-typed programming is that many small things become difficult, including (but not limited to) the handling of secrets.
These kinds of statements are giving major "draw the rest of the owl" vibes.
https://i.kym-cdn.com/photos/images/newsfeed/000/572/078/d6d...
Ultimately most CI/CD setups are basically systems administrators with privileged access to everything, network connected and running 24/7. It's pretty dangerous stuff.
I don't have an answer though, expect maybe to keep the CI and CD in separate, isolated instances that require manual intervention to bridge the gap on a case by case basis. That doesn't scale very well though.
There is an argument to be made for a minimalist CI/CD implementation that can handle task scheduling and dependencies, understands how to fetch and tag version control, count version numbers and not much else. Even extracting test result summaries, while handy, maybe should be handled another way.
For many of us, if CI is down you can't deploy anything to production, not even roll back to a previous build. Everything but the credentials should be under version control, and the right people should be able to fire off a one-liner from a runbook that has two to four sanity checked arguments in order to trigger a deployment.
These could just as easily be run from a developers machine with the right credentials.
With carefully chosen creds, Devs can run the CICD things directly, thus making feedback loops much faster in certain circumstances. Maybe don't let them delete databases, or run expensive VMs, but other than that go for it.
In my experience if the pipeline is even slightly different Devs will treat it as a foreign object and will whine about it not working the way they want... they don't feel like they can change the pipeline, which is odd. Dude, it's code, just change it, I'm happy to approve your PR and/or help you make a change.
I disagree with this 1,000%. This is a massive security risk.
From an InfoSec perspective it's a massive security risk because it implies the developers have access to AWS API's directly, otherwise how could they modify the pipeline or even engage it, when it manipulates infrastructure?
And that's not even to mention all the other security implications that come with allowing developers to edit a pipeline... holy s-... I don't even want to imagine the chaos brought about by a quick edit to the `.gitlab-ci.yml` file that runs `terraform destroy -auto-approve` an hour before handing in one's notice.
And no, code reviews aren't a sufficient barrier on their own.
This is just such a bad idea on so many levels.
> In my experience if the pipeline is even slightly different Devs will treat it as a foreign object and will whine about it not working the way they want...
Who cares? That's not how DevOps works. It's not how a business works. Everyone has a part to play and it's not always going to be comfortable. Operations specialists are just that: specialists at handling the operations. Let them do it, just as they let developers use the tools they want or handle things the way they see best.
Comfort doesn't apply.
IMO, the worst CI/CD tool on the market is Bamboo, and I've been using CI since you had to edit an XML file and restart.
And the reason for that is information hiding. The Bamboo UI is completely undiscoverable. All of the things you can do with it are buried in the docs. Anything you don't have permission to do is eliminated from the UI, so unless you RTFM you don't even know what tools you have available to solve problems.
>>“Pretend you have compromised a developer’s laptop.”
Most companies will fail right here. Especially outside of the tech world security hygiene with developer's laptops is very bad from what I have seen.
cries in security
This problem is solvable without hard coding env variables into your docker-compose.yml file.
You can commit an .env.example file to version control which has non-secret defaults set so that all a developer has to do is run `cp .env.example .env` before `docker-compose up --build` and they're good to go.
There's examples of this in all of my Docker example apps for Flask, Rails, Django, Phoenix, Node and Play at: https://github.com/nickjj?tab=repositories&q=docker-*-exampl...
It's nice because it also means the same docker-compose.yml file can be used in dev vs prod. The only thing that changes are a few environment variables.
I put it `insecure`. I think it makes it clear that the password, and file, aren't secure by default and should be treated as such.