3,698 karma · joined June 3, 2011
Previously at Era Software, Thorn, Webflow, Playlist.
jacobwgillespie@gmail.com
https://jacobwgillespie.com
https://github.com/jacobwgillespie
https://twitter.com/jacobwgillespie
[ my public key: https://keybase.io/jacobwgillespie; my proof: https://keybase.io/jacobwgillespie/sigs/F-KFv_EcnLeEUyxHBMSuyDH8bw4t5-4sfhrxE-zbdMs ]
I'd like to have a more automatic integration at some point - the challenge is that a lot of BuildKit's architecture performs best when many different build requests all arrive at a single build host, it is then able to efficiently deduplicate and cache work across all those build requests. So you really want the many different Actions jobs all communicating with the same BuildKit host.
We have some ideas for reducing the amount of change to Actions workflows to adopt ^ - longer term we're also working on our own build engine, to free those workloads from being confined to single hosts (be that single CI runners or single container builders).
Our original version of that system used vanilla BuildKit + EBS volumes + orchestration, nowadays we've replaced EBS with a distributed ceph storage cluster for significantly faster IOPS and throughput and have modified BuildKit to be better suited for high-performance distributed builds.
Both the container build service and the Actions runners are in the same AWS VPC, so they get good network performance between the two and don't need to egress over the internet.
- For instances with >= 32 vCPUs, traffic to an internet gateway can use 50% of the throughput
- For instances with < 32 vCPUs, traffic to an internet gateway can use 5 Gbps
- Traffic inside the VPC can use the full throughput
So for us, that means traffic outbound to the public internet can use up to 5 Gbps, but for things like our distributed cache or pulling Docker images from our container builders, we can get the full 12.5 Gbps.
[0] https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/ec2-inst...
With Depot, we're moving towards deeper performance optimizations and observability than vanilla GitHub runners - we've integrated the runners with a cache storage cluster for instance, and we're working on deeper integration with the compute platform that we built for distributed container image builds - as well as expanding the types of builds we can process beyond Actions and Docker, for instance.
But different options will be better for different folks, and the `philips-labs` project is good at what it does.
1. Launch, prepare basic software, shut down
2. A GitHub job request arrives at Depot
3. The job is assigned to the stopped VM, which is then started
4. The job runs on the VM completes
5. The VM is terminated
So the pool exists to speed up the EC2 instance launch time, but the VMs themselves are both single-tenant and single-use.
The challenge with Arm is actually just that GitHub doesn't have a runner image defined for Arm. For the Intel runners, we build our image directly from GitHub's source[0], and we're doing the same for the Arm runners by patching those same Packer scripts for arm64. It also looks like some popular actions, like `actions/setup-*`, don't always have arm support either.
So the disclaimers for launching Depot `-arm` instances at the moment is basically just (1) we have no idea if our image is compatible with your workflows, and (2) those instances take a bit longer to start.
On achieving fast startup times, it's a challenge. :) The main slowdown that prevents a <5s kernel boot is actually EBS lazy-loading the AMI from S3 on launch.
To address that at the moment, we do keep a pool of instances that boot once, load their volume contents, then shutdown until they're needed for a job. That works, at the cost of extra complexity and extra money - we're experimenting some with more exotic solutions now though like netbooting the AMI. That'll be a nice blog post someday I think.
[0] https://github.com/actions/runner-images/tree/main/images/ub...
- Depot runners are hosted in AWS us-east-1, which has implications for network speed, cache speed, access to internet services, etc. (BuildJet is hosted in Europe - maybe Hetzner?)
- Also thanks to AWS: each runner has a dedicated public IP address, so you're not sharing any third-party rate limits (e.g. Docker Hub) with other users
- We have an option to deploy the runners in your own AWS account or VPC-peer with your VPC
- We're integrating Actions runners with the acceleration tech we've built for container builds, starting with distributed caching
GitHub's incentives and design constraints are different than ours. GitHub needs to offer something that covers a very large user-base, to cover the widest possible number of workflows, and they've done this by offering basic ephemeral VMs on-demand. CI and builds are also not GitHub's primary focus as an org.
We're trying to be the absolute fastest place to build software, with a deep focus on achieving maximum performance and reducing build time as much as possible (even to 0 with caching). Software builds today are often wildly inefficient, and I personally believe there's an opportunity to do for build compute what has been done for application compute over the last 10 years.
GitHub Actions workflows are more of an "input" for us then (similar to how container image builds have been), with the goal of adding more input types over time and applying the same core tech to all of them.
We have some docs on this for our container builder product - still need to write the docs for Actions runners too, though they use the same underlying system: https://depot.dev/docs/self-hosted/overview.
But even with them both using Wireguard, there are choices involved that affect performance, for instance whether to use the Wireguard kernel model or userspace implementation, how to configure routing, packet filtering, firewalls, etc.
The downside is durability and operations - we have to keep Ceph alive and are responsible for making sure the data is persistent. That said, we're storing cache from container builds, so in the worst-case where we lose the storage cluster, we can run builds without cache while we restore.
No idea if this holds if/when the email crawler bots start executing JS on crawl.
Happy to answer any questions / provide any additional details!
But we needed a way to authenticate our CLI within those public workflows. This OIDC issuer is the result of that need, and works like so:
1. The pull_request workflow makes a "claim request" to the OIDC issuer, claiming certain details about the workflow like the ID, run ID, repository, etc.
2. The OIDC issuer responds with a "challenge code" that the workflow must periodically print to its logs
3. The OIDC issuer connects to the GitHub Actions websocket endpoint for log streaming, validates that the challenge code is being printed, then returns a new OIDC token to the workflow
This is working well for us, and lets us acquire an OIDC token similar to the GitHub Actions native OIDC token. The issuer itself runs as a Cloudflare Worker.
Happy to answer questions and I'd love any feedback you may have!
We're now using Ceph, running a storage cluster in top of i3en instances with local storage. This lets us thin-provision the cache storage volumes, elastically grow them as needed, and deliver much better IOPS and throughput compared to the gp3 volumes we used to use. Some very large builds that used to take as high as 8 minutes to construct image layers now take less than a minute!
For the most part, Ceph has been very seamless to deploy / operate, though we've been learning many fun facts about low-level Linux filesystems. :)
That preference has had an impact on the Node ecosystem at large, given how prolific an OSS contributor Sindre is, but IMO that influence has been earned by the large body of work he's contributed.
> I think we solved it with buildkit cache
One big thing we're doing here, if you're familiar with BuildKit cache, is providing builds a stable cache SSD that's reused between builds. This means we support all of BuildKit's caching features, including things like cache mounts that aren't directly supported in ephemeral CI environments. Plus Depot doesn't need to save or load the cache to a remote store like S3 or GitHub Actions cache, instead the previous cache is immediately available on build start.
This may not be any better or different than what you're doing, I just wanted to mention the detail for anyone familiar with trying to make BuiltKit more performant.
To make this performant we keep a certain number of spare "warm" machines ready for build requests so that you don't have to pay the instance launch time penalty yourself.
Both Kaniko and BuildKit can be run in rootless mode - we are not doing this, instead we give every builder access to an isolated VM, so builds are a bit quicker as well by avoiding some of the security tricks that rootless needs to work.
The most important bit though is that we have a `depot/build-push-action` that implements the same inputs as Docker's `docker/build-push-action`, so just swapping that line and adding a project ID and access token are the majority of what you'd need to do:
- uses: depot/setup-action@v1
- uses: depot/build-push-action@v1
with:
project: <your-depot-project-id>
token: ${{ secrets.DEPOT_TOKEN }}
# Whatever other inputs:
context: .
push: true
tags: |
...
I think that's along the lines of what you're describing as a Depot GitHub Action: https://github.com/depot/build-push-action.Just to note, you can totally use Depot within your GitHub Actions runs, even if those runs are happening inside self-hosted or BuildJet-hosted runners. You might get the best of both worlds that way, having your builds and tests outside of Docker run on BuildJet, and Docker builds accelerated with Depot.