CNCF to Host Cloud Native Buildpacks in the Sandbox
cncf.io
cncf.io
Presentation to CNCF TOC: https://docs.google.com/presentation/d/1RkygwZw7ILVgGhBpKnFN...
Formal specification: https://github.com/buildpack/spec
Sample buildpacks: https://github.com/buildpack/samples
Here's the node one: https://github.com/buildpack/samples/blob/master/nodejs-buil...
What you're looking at here is informally the "v3" effort, which extends and consolidates the "v2a" and "v2b" designs evolved by Heroku and Cloud Foundry respectively from the original Heroku design.
In v2 both Heroku and Cloud Foundry provide supported PHP buildpacks, as well as Java, Ruby, Python, .NET Core and I forget the rest right now. There are hundreds of community buildpacks.
What is the rationale behind having the build step tied to the packaging format?
Typically this means that buildpacks look for files that correspond to the relevant ecosystem. Maven buildpacks look for pom.xml. PHP buildpacks look for composer.json. Etc.
Nothing in this creates a hard binding. Detect steps may use whatever logic they need to decide on whether to signal they can work on a codebase.
Edit: in the v3 design the detect script can also provide dependency information that later steps can pick up. So for example, a JDK buildpack can say "yes, I can interpret this codebase, and I can contribute a JDK". A later buildpack can then look for this contribution as a condition, eg. the Maven buildpack can say "I will proceed if I see a pom.xml and if there is a JDK available".
All jokes aside, this looks great. Super-early of course - seems like there are quite a few issues in the `pack` repository to be implemented - but I'm excited to see where this lands. Buildpack detection and application is not a straightforward problem.
(ps. I know what these surprises can feel like -- for what it's worth, I am sorry)
1. Developers.
What you need to know is: it just works. You don't even need to write a Dockerfile any more. The buildpack turns your code into a well-structured, efficient runnable image with no additional effort.
2. Operators.
What you need to know is: oh thank god no more mystery meat in production. Patching the OS is a ho-hum affair instead of a stone grinding nightmare that turns your developers into a white hot bucket of rage because you have to nag them or block them from deploying or both.
3. Platform vendors, buildpack authors and curious passers-by
What you need to know is: All the other stuff about detect, analyse, build or export.
Unless you are in group 3, the basic thing is that Buildpacks require less effort than Dockerfiles with more safety and faster builds.
The samples take advantage of this (as well as a separate transparent cache) in order to demonstrate the different aspects of the formal spec. A simple buildpack is not necessarily much more complicated than a simple Dockerfile.
My issue with Buildpacks is that it looks like a glorified bash script (which is a skill I am not bashing — pun not intended) whereas a dockerfile is much more human readable and the idea of layers, for a guy coming from a systems background, is much more intuitive for me. The analogy of a very lightweight VM makes perfect sense to me which means I’m much more productive with it.
Disclosure: I work for a Said Vendor.
With Docker it is easy to make this work using multi-stage builds, I'm not sure if buildpacks account for this at all.
This isn't strictly true of v2 designs (both Heroku and Cloud Foundry have schemes for multi-buildpack support) and in the v3 design explicit consideration is given to making mix-and-match a triviality. Buildpacks can cooperate quite easily and in multiple groups.
Where this shines is in updates. If I have an OS base layer, a JDK buildpack and a Maven buildpack, the layers they generate can be independently updated without needing a complete rebuild. So far as I am aware, this is not currently possible with a Dockerfile, multibuild or not. If you invalidate the OS layer, everything else gets rebuilt whether it needed to be or not.
I think you are right about the rebuilds, but maybe not in all cases.. If you had.
FROM nodejs:whatever as js
RUN npm build
FROM golang:whatever as go
RUN go build
FROM base
COPY --from js app.js
COPY --from go app
If only 'base' was updated rebuilding the image would just need to re-run the COPY commands. The only way everything would get rebuilt is if you also updated the go and js images.I think for more complicated applications https://grahamc.com/blog/nix-and-layered-docker-images is ultimately the best way to do things... ensuring that every step is isolated with explicit dependency information.
Google folks worked on "FTL", which is a technique for determining layering by reasoning about packaging system information. Jib[1] is one such system. There is a view that it will be possible to use FTL implementations or derivatives as buildpack components in the future.
[0] https://docs.google.com/document/d/1M2PJ_h6GzviUNHMPt7x-5POU...
We've also met to talk about buildpacks and have pre-existing working relationships with all the relevant folks in Google Cloud Builder and Google Container Tools teams through our work on Knative.
Though I understand that there are likely only dozens of us, my question is: what would the use of Buildpacks buy us for this specific use case?
Unfortunately the resulting image would be larger than just a single binary, but building it would be a lot easier and repeatable for other people working on the project. The base layers would be cached though so in practice it might not be that much larger on disk.
I looked into using bazel for doing similar things, and the biggest stumbling block is that bazel itself is a PITA to install on all platforms. May end up trying https://please.build/ at some point.
One thing I really like about bazel is the pkg rules.. I currently use fpm/goreleaser(nfpm) to build rpms for things, and it's nice having a single build tool that can build the app and spit out an rpm.
I will make mistakes, but ... AMA.
Here is my source code.
Run it on the cloud for me.
I do not care how.
The gist is that Docker containers are awesome for the Day 1 experience. I write a Dockerfile and I'm off to the races.But then Day 2 rolls around and I have a production system with 12,000 containers[0].
1. What the hell is in those containers, anyhow?
2. A new CVE landed and I want to upgrade all of them in a few minutes without anyone being interrupted (or even having to know). How?
3. I have a distributed system with many moving parts. I build a giant fragile hierarchy of Dockerfiles to efficiently contain the right dependencies, making development slower. Then I snap and turn it into a giant kitchen-sink Dockerfile with the union of all the dependencies in it. Now production is slow as hell.
4. Operations become upset about points 1-3. Now I can only use curated Dockerfiles, can only come through our elaborate Jenkins farm, rules rules rules. Wasn't the purpose of Dockerfiles to make this all just ... go away?
Buildpacks solve all of these. I know what's in the container because buildpacks control the build. I can update CVE flaws in potentially seconds. Each container can have what it needs - no more, no less.
And most important: the buildpack runs locally, or in the cluster, exactly the same. It's all the developer benefits of Dockerfiles/docker build, minus most of the suck.
[0] Yahoo! Japan has 12,000 prod and 8,000 dev containers using buildpacks: https://www.slideshare.net/iranainanimosuteteshimaou/yahoo-j...
If your use case fits into a scenario which is handled by an existing buildpack, then the claim is that you’ll be better off using the buildpack because the infrastructure can make optimizations that can’t be made with arbitrary containers.
If your use case isn’t covered by a buildpack, then you can either (1) make a buildpack or (2) revert to raw containers.
(Although your platform may not allow (2)).
Clound Native Buildpacks are based on a very well-proved model. Heroku do this at massive scale. So does Pivotal, so do many of our various customers and customers of other buildpack-using systems like Deis.
Step 2: Now that your Dockerfiles no longer contain any real information, retool your orchestration system to use source tarballs rather than docker images.
The cool part is that when the patch arrives, devs don't have to do anything to be patched. Cloud Foundry already does this with buildpacks and so does Heroku.
A big part of what's new is that layer rebasing could make this really really fast.
Aside from that; I think its a bad idea to have things update automatically. What if the upstream fix breaks things? It reduces trust in the build system.
> For companies with compliance requirements, you're required to change the systems explicitly; this is one of the reason why Docker images work so well; you can target specific tags for deployment and nothing changes with the same tag.
This isn't really true, though. Tags are floating targets, only the digest is stable. Taking Kubernetes as an example, suppose I push an updated Pod definition where I've changed an image tag from v1 to v2. If the tag is not properly locked, then I can be running multiple versions of the software without even realising it.
Speaking of regulation, we find a lot of people like buildpacks for that exact reason. Operators know exactly what OS is running in every container, exactly what JDK is running in every container, everything up to the runtime (and as FTL matures, up to the package dependencies as well). The platform doesn't have to accept any old container, they can all enter through a trusted pathway.
You can do this with docker builds, of course. You build CI/CD, you have centrally-controlled images, prevent non-conforming images from reaching production and so forth. But then you've pretty much recreated all the stuff buildpacks gave you, except you're the one having to maintain it.
> Aside from that; I think its a bad idea to have things update automatically. What if the upstream fix breaks things? It reduces trust in the build system.
Rebasing layers is close to instant. You can rollback the change as soon as it looks bad. More to the point, if the OS vendor or runtime have broken ABI compatibility, rebuilding a docker image won't necessarily help you to notice that before runtime.
Here's how I would do it:
rm Dockerfile
cf push
aaand I'm done.Someone using Heroku might do it wildly differently. It's super complicated:
git rm Dockerfile
git commit -m "Switch to Cloud Native Buildpacks"
git pushYou could use the resulting container from a buildpack in an image built by LinuxKit.