The answer is very much, 'it depends'. For oen thing, developers can run whatever code in CI before it's benn reviewed. I could just nab the env vars and post them wherever. If there are no sensitive env vars for me to nab and you have enforced code review, then I need a co-conspirator, and my change is probably going to leave a lot more of a paper trail.
Another risk is accidental disclosure - I have on at least two occasions accidentally logged sensitive environment variables in our CI environment. Now your threat model is not just a malicious developer pushing code - it's a developer making a mistake, plus anyone with read access to the CI system.
I don't know about your org, but at my job, the set of people who have read access to CI is a lot larger than the set who can push code, which is again a lot larger than the set of people who can merge code without a reviewer signing off.
> but drawing these boundaries seems like a nightmare for everyone...
As someone currently struggling with how to draw them, yup.
Yeah I don't think this gets talked about enough.
If you're talking about private repos in an organization then CI often runs on any pull request. That means a developer is able to make CI run in an unreviewed PR. Of course for it to make its way into a protected branch (main, etc.) it'll likely need a code review but nothing is stopping that developer who opened the unreviewed PR to modify the CI yaml file in a commit to make that PR's pipeline do something different.
Requiring a team lead or someone to allow every individual PR's pipeline to run (what GitHub does by default in public repos) would add too much friction and not all major git hosts support the idea of locking down the pipelines file by decoupling it from the code repo.
Edit: Depending on which CI provider you use, this situation is mostly preventable -- "mostly" in the sense that you can control how much damage can be done. Check out this comment later in this thread: https://news.ycombinator.com/item?id=29967077
If a developer changes the CI pipeline file to make their PR's code run in `deployment: "production"` instead of `deployment: "test"` doesn't that bypass this?
Edit:
I'll leave my original question here because I think it's an important one but I answered this myself. It depends on which CI provider you're using but some of them do let you restrict specific deployments from being run only on specific branches or by specific folks (such as repo admins).
In the above case if the production deployment was only allowed to run on the main branch and the only way code makes its way into the main branch is after at least 1 person reviewed + merged it (or whatever policy your company wants) then a rogue developer can't edit the pipeline in an unreviewed PR to make something run in production.
Also with deployment specific environment variables then a rogue developer is also not able to edit a pipeline file to try and run commands that may affect production such as doing a `terraform apply` or pushing an unreviewed Docker image to your production registry.
CI should never ever have access to anything related to production; not just for security but also to prevent potentially bad code being run in tests from trashing production data.
As a concrete example, GitLab has the concept of protected branches and code owners, both of which allow you to restrict access to the corresponding environments’ credentials to a smaller group of people who have permission to touch the sensitive branches. That allows you to say things like “anyone can run in development but only our release engineers can merge to staging/production” or “changes to the CI configuration must be approved by the DevOps team”, respectively.
That does, of course, not prevent someone from running a Bitcoin miner in whatever environment you use to run untrusted merge requests but that’s better than access to your production data.
For private repositories, that means access to credentials. Probably read-only credentials, but it requires network access.
Would you be suggesting that everyone should commit all dependencies?
I'm a proponent of lock-file approaches, which gain 99% of the benefits with far less pain. It requires network access, though.
However, it's problematic. Only use it if you're certain it will solve a specific problem you have.
Consider using Docker to build images that include a snapshot of node_modules.
I agree that there are tokens and variables that are dangerous to expose via CI, but throwing the baby out with the bathwater confused me.
1. stand up a private package mirror that you control, that uses a whitelist for what packages it is willing to mirror;
2. configure your project's dependency-fetching logic to fetch from said mirror;
3. configure CI to only allow outbound network access to your package mirror's IP.
The disadvantage — but also the point — of this, is that it is then a release manager's responsibility, not a developer's responsibility, to give the final say-so for adding a dependency to the project (because only the release manager has the permissions to add packages to the mirror.)
I guess that's beside the point if your goal is only to reduce risk of compromised CI/CD.
Every CI system I've ever seem has pulled dependencies in from the network.
git clone --recursive --branch "$commitid" "$repourl" "$repodir"
img="$(docker build --network=none -f "$dockerfile" "$repodir")"
docker run --rm -ti --network=none "$img"
Sure, CI pulls in from the network... but execution occurs without network.First:
> git clone --recursive --branch "$commitid" "$repourl" "$repodir"
The `git clone` will take your git repository URL as the $repourl variable. It will also take your commit id (commit hash or tagged version which a pull request points to) as the $commitid variable (`--branch $commitid`). It will also take a $repodir variable which points to the directory that will contain the contents of the cloned git repository and already checked out at the commit id specified. It will do so recursively (`--recursive`): if there are submodules then they will also automatically be cloned.
This of course assumes that you're cloning a public repository and/or that any credentials required have already been set up (see `man git-config`).
Then:
> img="$(docker build --network=none -f "$dockerfile" "$repodir")"
Okay so this is sort've broken: you'd need a few more parameters to `docker build` to get it to work "right". But as-is, `docker build` usually has network access so `--network=none` will specify that the build process will not have access to the network. I hope your build system doesn't automatically download dependencies because that will fail (and also suggests that the build system may be susceptible to attack). You specify a dockerfile to build using `-f "$dockerfile"`. Finally, you specify the build context using "$repodir" -- and that assumes that your whole git repository should be available to the dockerfile.
However, `docker build` will write a lot more than just the image name to standard output and so this is where some customization would need to occur. Suffice to say that you can use `--quiet` if that's all you want; I do prefer to see the output because it normally contains intermediate image names useful for debugging the dockerfile.
Finally:
> docker run --rm -ti --network=none "$img"
Finally, it runs the built image in a new container with an auto-generated name. `-ti` here is wrong: it will attach a standard input/output terminal and so if it drops you into an interactive program (such as bash) then it could hang the CI process. But you can remove that. It also assumes that your dockerfile correctly specifies ENTRYPOINT and/or CMD. When the container has exited then the container will automatically be removed (--rm) -- usually they linger around and pollute your docker host. Finally, the --network=none also ensures that your container does not have network access so your unit tests should also be capable of running without the network or else they will fail. You could use `--volume` to specify a volume with data files if you need them. You might also want to look at `--user` if you don't want your container to have root privileges...
And of course if you want integration tests with other containers then you should create a dedicated docker network and specify its alias with `--network`: see `man docker-network-create`; you can use `docker network create -d internal` to create a network which shouldn't let containers out.
Does that answer your question?
It's been my experience that usually pull requests which make a library capable of being built offline are welcome. Usually. So get off your duff and get to fixing the problems in the free packages you use.
Luxury! When I were a youngun our devnet was airgapped and we had to sneakernet dependencies to it. Using vhs tape on a 9track spool. Because real tape was too expensive.
More seriously I've seen jenkins and gitlab pipelines with only access to sneakernet maintained mirrors.
However, it'd probably be better if you could have the CI framework collect and inject that information into the build using some hard-coded deterministic logic, rather than giving the build itself (developer-driven Arbitrary Code Execution) access to that capability.
Same idea as e.g. injecting Kubernetes Secrets into Pods as env-vars at the controller level, rather than giving the Pod itself the permission to query Secrets out of the controller through its API.
I'm failing to understand how that procedure even works. How do you run the tests?
It's sort of like how I tell my mom, "Even you don't want to know your passwords", when explaining that she should use a password manager. The reason in this case is because she may get fished and redirected to a fake website. If there's a password manager, and she doesn't even know her own password, then it won't recognize the certificate and she's safe. However, if she does know it, then she has the ability to leak it.
This is the same with environment variables, and a reason why secret storage is preferred over them. Think of environment variable as your mom knowing her password, and secret storage to her using a password manager. More or less haha.
Maybe. Back in the old days if you had the commit bit your badge didn’t get you into the server room. I get the impression a lot of shops are effectively giving their devs root but in the cloud this time, which isn’t necessary.