Terraform module for scalable GitHub action runners on AWS
github.com
github.com
My experience is that it works until it does not, and then you are down on a rabbit hole trying to figure out why.
The only reason we keep using it is because we have other priorities, but once we have more spare time, this module is going to hell, it's not reliable at all.
One of the machines in the fleet failed to sync its clock via NTP. Once a job X got scheduled to it, the runner pod failed authentication due to incorrect clock time, and then the whole ARC system started to behave incorrectly: job X was stuck without runners, until another workflow job Y was created, and then X got run but Y became stuck. There were also other wierd behaviors like this so I eventually rebuilt everything based on VMs and stopped using ARC.
Using VMs also allowed me to support the use of the official runner images [0], which is good for compatibility.
I feel more people would benefit from managed "self-hosted" runners, so I started DimeRun [1] to provide cheaper GHA runners for people who don't have the time/willingness to troubleshoot low-level infra issues.
[0]: https://github.com/actions/runner-images [1]: https://dime.run
If something fails and you don't have idle runners (hence wasting unnecessary resources), things start to snowball.
Instead of being a service, I'm also open to sell the software+hardware solution behind it, so you can have it on-prem. Do you think that's something you would consider given the constraints on supply chain security?
That said, do many see value is in this? The current stance is you either self-host or accept paying github's runner rates out of laziness. In the end it's all very mild when it comes to extra costs.
When things go wrong, you need people going through logs working out that there's nothing actually wrong with CI and in the end finding "oh it was a re-provisioning of compute", so...let's run it again and hope that goes well, actually no...let's schedule it for 11pm Eastern which probably isn't that optimal.
I can see it working for some, but horrendous for others. Maybe if you have burstable CI loads it's worth doing and re-rolling on failure.
Still, a marvellous readme. If only others in open-source would take note that this is how you do it.
There’s also people who use GitHub Enterprise who still need to host a runner fleet, so they could benefit from this too.
* Dependabot and AI PRs might be exceptions.
We actually looked at using this project from OP at the beginning of our ci journey but eventually decided on actions-runner-controller due to its higher level of robustness and configurabiltiy
There are several other different architectures that range from simpler to more complex. The architecture I recommend people start out with is a single long lived beefy buildkit instance that a bunch of runners share, since that is much much simpler to implement. It of course has the downside that you have to refresh/rebuild the cache if the instance ever goes down. For runs that need read/write locks on volumes (eg Gradle build cache) my recommendation after trial and error to rsync those to the runners and then rsync them back after the run completes so you don’t have a bunch of locks fighting each other for the same folder.
You even need to self-host if you want to test code that uses AVX512 outside of an emulator since the default runners do not support that. Same if you want to test aarch64-specific code paths on Linux, Windows, or macOS.
GitHub don't host runners for Linux × arm64, so if you need this, you need to self-host. You can also run custom AMIs with pre-installed packages, which can speed up workflows that depend on those packages.
> When things go wrong, you need people going through logs working out that there's nothing actually wrong with CI…
I'm on a small team who've been running the Philips Lab self-hosted runners for the past year. It hasn't been difficult to operate. Once deployed, it pretty much "just works".
In my experience, the things that go wrong originate from the GitHub workflows themselves. We usually have to review workflow logs regardless of whether the workflow uses a self-hosted runner or not.
Is this not possible via QEMU like everyone has been doing for a very long time now?
The real issue here seems to be large companies wanting their own syntactical turf everyone abides by that they can later profit from.
We have extremely bursty CI loads. One push can kick off up to 60 different CI jobs. If a few people push at the same time, we easily have 300 jobs running in parallel. This happens the whole day long.
It absolutely makes sense to self host this, since the cost to run this on Github runners would be prohibitive (Github runners are 8x the cost of the equivalent AWS instance I think). All our runners are ephemeral, so we only pay for them when they are actually running jobs. After all the jobs are over, the runners are immediately scaled to zero.
I build this whole thing myself, so it's a bit sad to find someone had already built the whole thing before.
https://github.com/actions/actions-runner-controller
I think Kubernetes is a better platform than EC2 for runners, it's faster and more integrated with your tooling (if you're using Kubernetes).
Eventually I stumbled on the idea of running the VPC-requiring commands from an AWS CodeBuild script, and invoking it from a workflow executed on a GitHub-owned runner. Works beautifully and I was able to remove a ton of complexity from my infra that this Terraform module adds.
It all comes down to the instances. GitHub runners are insanely complex and also the best part of actions, along with community driven reusable workflows. They not only pack so much stuff without conflict, they also spin up really fast (forgot which virtualization infrastructure they use). So when you move to self-hosted you now have instances that take longer to be ready, increasing cold start times from pipelines, that will require that you install everything you need in it to run your jobs, and that, depending on how you set them, will require more thorough cleanup after each job.
The scalability is really good though, and all the benefits already cited, like having access to private resources on your VPC, and delegating permissions to instance profile through IAM roles, make this project a godsend.
I am in the process of setting it up on a cheap Hetzner box. If it works, would be a great deal! You can get a 64 GB RAM box for 35 EUR/mo at server auctions with unlimited traffic. I don't mention CPU or GPU, as typically this isn't a bottleneck for my projects.
Plus, I can configure cache sharing via host-mounted dir. E.g. pnpm cache can be all in one place, and be locally available to pods via a mounted dir. Same for the Docker image cache. This would speed up CI runs and also reduce network traffic by a huge margin.
GitHub Actions effectively has no local caching. There's an action for caching, but it uses a blob storage for cache artifacts. Which then gets network fetched, gzip'ed and gunzip'ed each time, and from my experience this has never been a gain for medium to large npm projects, as they have thousands of small .js files in node_modules, and thus takes a long time to compress and decompress. I think npm edge cache servers are already so optimized and fast, that in my experience almost always it's faster to install from npm directly. I even tested this on AWS, where the cache was stored in S3, in the same region as CodeBuild (CI), and direct installs from npm were still faster by about 30%.
So other than adding more hardware resources, local caching is the only way to significantly speed up GH Actions, from my experience, and thus you must have your runner.
Hey there, we offer Ubicloud Runners that are 10x cheaper than GitHub and bill them by the minute. You get a fresh VM with each job; and we use Hetzner as our underlying provider.
If you're already setting up a Hetzner box, I'd love to get your input. Any thoughts or feedback for us?
https://www.ubicloud.com/use-cases/github-actions
https://github.com/ubicloud/ubicloud/blob/main/routes/web/we... (our github actions integration is also openly available)
We use BuildJet now and it's $0.008/min for 4 vCPU and 16 GB, while your offer is 0.16¢/min for the same (20x more expensive).
We didn't quote our prices in $ because the number of trailing zeros confused people. Maybe still go ahead and switch back to that?
Right now I’m using Karpenter, ARC, EFS and buildkit and it’s great, but it was also like a month of setup and is nontrivially complex.
We're using this Philips Labs module at $dayjob. It's a great piece of work!
1. A codebuild job running a container image which starts a runner with custom labels
2. A lambda to receive the webhook from github and run the codebuild job on demand.
What would implementing all the infra give you in terms of benefits over codebuild?
The readme doesn't mention anything about codebuild that I can see.
Shameless plug: I wrote a CDK construct that lets you create runners on demand in response to a GitHub webhook (so basically what was suggested here only with handling of more corner cases that came up). It lets you start runners in CodeBuild, Fargate, ECS, EC2, or Lambda so you can pick whichever works best for you. There is a table in the readme that shows the difference between them and why you might want to choose one over the other. https://github.com/CloudSnorkel/cdk-github-runners
It’s funny because just the other day I thought about implementing the same for GitLab in order to learn more about terraform.
(Note: I work for GitLab, but as a Frontend Engineer and usually not on CI topics)
It will be in production with a lot of usage soon so I can evaluate how it scales.