BuildKit in depth: Docker's build engine explained
depot.dev
depot.dev
Also, buildx won't commit the intermediate layers during the build. So if something fails, you can't just grab the previous intermediate layer and do `docker run --entrypoint /bin/bash` on that layer to poke around.
I've had cases where BuildKit will get stuck and `docker buildx prune -a` will leave things around. The way to fix it is with `systemctl restart containerd.io`. This seems to happen a lot when I interrupt the build (eg, I know it will fail and need to restart it with a fix), but it's also happened on successful builds.
All of that is on top of docker itself having extremely poor error messages.
I really do not like docker or buildx.
- There's no docker.sock API socket by default, which makes sense because Podman doesn't have Docker's daemon-based architecture. You can run one with `podman system service [...]` if you have non-CLI clients expecting to connect to some $DOCKER_HOST
- Mounts/volumes behave subtly differently IIRC, I had a bunch of scripts that didn't work with podman-as-docker out of the box, I think the issue was that Docker will create some directories automatically that Podman won't?
Aside from that, rootless Podman is great and I've had far fewer issues with it than I did for rootless Docker, only real caveat is that -p <some port> doesn't actually connect to the host, it's still namespaced. You have to use --net=host for actual host networking.
The mount/volume issues are probably caused by SELinux / other security decisions
That said, my pipelines are fairly simple, as are my Dockerfiles, so...
It shaved significant time off my pipelines, though, not having to wait for a docker-in-docker service to spin up (I use Gitlab with Kubernetes runners).
There's actually a whole debugging client available.
Docs: https://docs.docker.com/engine/reference/commandline/buildx_...
We ended up switching to using multiarch/qemu-user-static:7.2.0-1 <https://hub.docker.com/layers/multiarch/qemu-user-static/7.2...> which also mutates binfmt_misc in buildx's context in order to exec the static copy of qemu in it but I find their shell script much more legible about what's going on: <https://github.com/multiarch/qemu-user-static/blob/v7.2.0-1/...> not to mention the fact that it more clearly represents the recent qemu version versus secreting it away inside the moby/buildkit:heheh-good-luck-friend
They have an entire GH label for that: https://github.com/docker/buildx/labels/area%2Fqemu and my journey into the buildx sewers started from https://github.com/docker/buildx/issues/1170
Delete the entire docker stack. It's utter shite. Then install Podman and Buildah, and rejoice! For all the evil buildx magic can now be done explicitly.
This solution was born when trying to do native builds on CircleCI - where you do not have a buildx cluster with multiple architectures. You have to build the archs as independent build steps, then pull them together afterwards. This is something that Docker never considered, and Docker only functions on the happy path. Buildah has multiarch manipulation commands: https://danmanners.com/posts/2022-01-buildah-multi-arch/ (not a plug, first blog on search)
We migrated to GHA, but still use separate builders to avoid the Docker trite. We also ditched DCT (Docker signing) for Cosign because - again - DCT is home-grown garbage.
`buildg debug` (Dockerfile debugger based on BuildKit) to rescue: https://github.com/ktock/buildg
I did some digging and wiresharking, and I'm pretty sure it's undocumented because the API is _insane_. It starts as HTTP/1 from the Docker client to the Docker engine, but a key BuildKit feature is that the engine pulls files & data from the client on-demand, which is hard in a normal REST API, so how does it do that? By renegotiating the connection to flip the direction after the client connects.
That means: the client sends an HTTP/1 request, the server offers to upgrade to HTTP/2 by in reverse, and then the server becomes an HTTP client and the client becomes an HTTP server, still on the same existing connection. All actual communication then happens as gRPC, but backwards.
Absolute madness, and very difficult to document or support in 3rd party SDKs (and so they haven't) but it's very clever. Some more context here: https://twitter.com/i/web/status/1423353288129396740
We're working on another project in this realm that might interest you. Happy to send over more details via email if your interested in better build APIs. My contact info is in my profile.
One example: https://ravichaganti.com/blog/2022-11-28-building-container-...
You can also use docker save to get a tarball and ship that file to another machine, which can be run through docker load and then run as if it was built locally.
If you have an oci bundle, you might look at runc instead: https://github.com/opencontainers/runc
> I got familiar with container intervals via, basically, "building images without docker"
That's been my approach, as well. However, right now I'm stuck at getting overlayfs to work without privileges (easy: use `unshare`), while not breaking rootless runc (apparently not so easy).
https://github.com/psanford/runc-examples/blob/master/rootle...
It's not exactly what you're asking for as it relates to running containers via runc. But this walks through how OCI layers are actually built up behind the scenes. In case it's helpful: https://depot.dev/blog/building-container-layers-from-scratc...
Anyway, fortunately I seem to have found a solution for now (running runc with an overlay rootfs without root), see the link in the other sibling/nephew comment I posted.
You could likely get very similar performance by provisioning a single host with good hardware and simply leverage the on-host cache.
We've optimized BuildKit for remote container builds with Depot. So we've added things like a different `--load` for pulling the image back that is optimized to only pull the layers back that have actually changed between the build and what is on the client. We've also done things like automatically supporting eStargz, adding the ability to `--push` and `--load` at the same time, and the ability to push to multiple registries in parallel.
We've removed saving/loading layer cache over the network. Instead, the BuildKit builder is ephemeral, and we orchestrate the cache across builds by persisting the layer cache to Ceph and reattaching it on the next build.
The largest speedup with Depot is that we build on native CPUs so that you can avoid emulation. We run native Intel and Arm builders with 16 CPUs and 32GB of memory inside of AWS. We also have the ability to run these builders in your own cloud account with a self-hosted data plane.
So the bulk of the speed comes from persisting layer cache across builds with Ceph and native CPUs. The optimized portions of BuildKit really help post-build currently. That said, we are working on some things in the middle of the build related to the DAG structure of BuildKit that will also optimize up in front of the build.
Seeing that reminded me of some healthy discussion in https://news.ycombinator.com/item?id=39235593 (SeaweedFS fast distributed storage system for blobs, objects, files and datalake) that may interest you. control-f for "Ceph" to see why it caught my eye
Writing your own front end to BuildKit can be pretty simple. At my work, this is sort of a starter task for everyone who joins the team.
My frontend was based on intercal and I don't recommend anyone use it but it was fun to play around with[1].
[1]: https://github.com/adamgordonbell/compiling-containers/tree/...
We are building on top of it, and sometimes adding to it.
For instance, building up the DAG of all the build steps and then scheduling things so that various parts can be built in parallel. Buildkit does a lot that we use beyond just being a way do build things inside a runc container.
That said, we have a fork of buildkit for the various things we add that don't fit well in the upstream.
Already our auto-skip feature and our branching are implemented on top of Buildkit rather than using it. Probably as we grow and add more build centric features we will continue to diverge.
I'm just a DevRel person though, so that's just my 2 cents. The core team may disagree with me.
This is one of my hobbies as a programmer and I'm extremely interested in this area. Can you share more or is it a trade secret?
Furthermore, do you have links to tools or papers that deal with this?
I don't have any great references to be honest. It's not that its a trade-secret, its just that its an area I don't work on.
The AWK book describes how to build you own Make using AWK. You might like that. And an explanation of it and python implementation is found here:
https://benhoyt.com/writings/awk-make/
Hopefully that helps.
I am considering writing a small library / framework to optimize scheduled tasks by making a DAG and inferring which ones are safe to be parallel and which aren't. So this can be useful.
Dagger is by the same folks that brought us Docker. This is their fresh take on solving the problem of container building and much more. BuildKit can more than build images and Dagger unlocks it for you.
Is that not the case?
1. You can get the developer UX improvements without shared caching
2. You can run it in k8s and get shared caches there. I believe it should support any BuildKit caching solutions, and they will help you figure that out in Discord
What problems does it solve that cannot be easily achieved with separate build and runtime files?
I am running a (relatively new) NX monorepo, which deploys ~6 microservices to various K8s environments using Skaffold. All services are Typescript, Node 20.11, with pnpm as a package manager.
I have separate build files for certain common base images, which mostly are Node-($v)-alpine with a few CI/common dependencies installed. I use Skaffold to compose the image builds.
When I Dockerize each service, I need to install dependencies for that specific service only, and in that case the install/build specification must be the same as the runtime specification. Rather than creating separate build image files and composing them with Skaffold, I simply use a build stage to install the microservice dependencies (using one of the base images with NX and PNPM), then copy them to the final image (which does not have those deps). This is a pretty common pattern, and I like working with it. The advantages to build stages are readability and simplicity (no extra orchestration needed), and the whole setup makes it very difficult to make mistakes.