Kustomize – Templating in Kubernetes
blog.stack-labs.com
blog.stack-labs.com
> kustomize lets you customize raw, template-free YAML files for multiple purposes, leaving the original YAML untouched and usable as is.
That's the very first line of Kustomize's README[1].
Kustomize is not a templating DSL. The conversations here mocking templating DSLs are not relevant.
The next sentence[1]:
> kustomize targets kubernetes; it understands and can patch kubernetes style API objects. It's like make, in that what it does is declared in a file, and it's like sed, in that it emits editted text.
Kustomize is a patching framework. It takes valid k8s resources, and allows you patch them, with partial k8s resources. The entire point of Kustomize to not invent something new.
Is that different from a templating engine?
Some newer documentation can be found at https://kubectl.docs.kubernetes.io/ which is still a little bit barren, but I expect it to improve. Some of the older documentation can be found at https://github.com/kubernetes-sigs/kustomize/tree/a5bb5479fb... (before the tool was integrated directly as a subcommand)
Full Disclosure: I work with Joe @ VMware
EDIT: Although, after properly reading that linked issue, I see you mean applying the same patch to different resources. I interpreted "target" to mean build target (e.g. deployment/app), you can use the same patch across different deployments (apps) and overlays (environment). However, I see your use case is a bit different.
> I wish it folowwed a similar pattern as `kubernetes` resources where each resource had an `apiVersion`, and the `kustomization` itself had an `apiVersion`.
As of 2.0, this is now the case[1]:
> apiVersion: kustomize.config.k8s.io/v1beta1
> kind: Kustomization
[1] https://github.com/kubernetes-sigs/kustomize/blob/a5bb5479fb...
Simple yaml is just fine, even repetitive simple yaml, and you're often better off building tooling to make repetitive updates rather than condense your config. Jsonnet is for building when you've got dozens rather than handfuls of objects.
Well defined Kubernetes objects don't actually take that much configuration, don't generally need templating. Typically, my production and staging environments are precisely the same deployment definition, except image and replicas. Replicas isn't defined in the document - it's part of the working state rather than something that gets "configured" (i.e. that field is owned by an HPA, even if it's min and max values are the same)
All configuration is stored in Configmaps and secrets. Those are more likely to be templated, but still probably not. If two values are the same in production and testing (and they should be - service names are the same, with namespace defining where they point), why configure it? Use a sane default.
That said its design is quite opinionated and committed to a declarative model, and there are some things you just can't do without falling back on either generating a patch or performing unstructured edits. A good example is something as simple as tagging an image with a string that isn't known until build time (such as the commit sha). Another is shared configuration. You can't glob object names and apply the same environment patch, for example, to more than one resource. These constraints can be considered features, but they're nonetheless constraints.
Good times.
Just use Python (or Lua if you want something simpler/lighter). Or, like Brigade, even JS. Hell, use BASIC. But these bizarre template-language-DSL bastardizations are horrific when it comes to maintainability. Incremental evolution from a config file to a full language is a path wrought with peril, and only ends in damnation!
* saltstack Jinja templating [1]
* GitLab CI `include` directive for including and merging external CI files together [2]
A few tools that I have seen take a decently pragmatic approach:
* Kubernetes resources that use ConfigMap and `envFrom` that declaratively say where to resolve a value from [3]
* Circle CI commands which offer some reusability with its "commands" and "executors" type features [4]. To me, Circle CI has both good and bad aspects with some templating and some clever patterns
On the other side of things, there is essentially fully programmable type configurations like Jenkins Pipeline Groovy `Jenkinsfile`, which can be a nightmare, too.
I think it is tough to find a sweet spot between making it configurable and expressive for users while retaining a low barrier of entry and not turning the configuration into a complete program itself. Tools like Terraform are trying to find that sweet spot as they slowly introduce more programmatic ways of configuration while still being declarative, like the fairly recent introduction of if statements and soon (I think) to be released for loops. As soon as users of a tool and service have more complex use cases, there needs to be some way to solve that. The most common way (it seems) is taking the easy and familiar of introducing templating.
[1]: https://docs.saltstack.com/en/latest/topics/jinja/index.html
[2]: https://docs.gitlab.com/ee/ci/yaml/#include
[3]: https://kubernetes.io/docs/reference/generated/kubernetes-ap...
[4]: https://circleci.com/docs/2.0/configuration-reference/#comma...
Nowadays, Jenkins pipelines can be configured in either the Jenkins-provided "Declarative Syntax" [1] or the Apache Groovy-based "Scripted Syntax", with the Declarative Syntax used as the default for examples on the Jenkins website. I guess they've found the best way to not have users turn the configuration into a complete program is to provide declarative syntax only in the default option. It's good to see Kustomize is built with this in mind, too.
Those of you who use Kubernetes: What is a functionality it brings to the table you would miss if you automated your docker handling without it?
- namespace separations between resources
- optimized resource utilization and container scheduling
These are the biggest 3 for the smaller orgs I’ve been a part of
Not to be flippant, but any case where things need to connect to other things. At scale they usually need to connect to other things through some sort of load balancer. You get that out of the box with kubernetes for services hosted inside the cluster, and there are straightforward solutions to ingress for clients outside it. Another important feature is pod scheduling. Yes you could wire up a few machines using docker compose and any of a few different networking approaches, but if one of your VMs dies are those workloads going to move to a healthy instance by themselves?
And it just works. Simply and intuitively. With everything. Because it's been an IETF standard since 1983.
But you don't need containers either.
Service discovery just makes it easier to link up services if your architecture is truly microservice. We used to have an incredible amount of config that tightly coupled our API's.. now we use a combination of service discovery and an API gateway (Ambassador) to decouple the services, cut down on the number of random endpoints in our config, and we also get the added benefit of load balancing, rate limiting, and additional logging.
There's always a tradeoff with scale. If you have four servers than obviously all of this stuff is overkill.
I disagree, actually. I have four servers at home, and have some pods that have been running little tinkery things, and a bunch of open source software, with ridiculous uptime and little or no effort, even when I reboot one of those "servers" to do some gaming on Windows.
Now, do I need that uptime for all of those services? Not really. But for some of them, I want it, and it'd be annoying if I had to go figure out why they'd stopped running. The reality is, things just keep ticking without me worrying when they're on k8s.
These skills transfer into very in-demand job skills as well, and if I ever build anything that gains traction, I already have all the tools, configs, and knowledge to deploy that app across 500 generic cloud servers.
I wouldn't recommend Kubernetes if you only have one application you run in production. Just rsync your production image to production whenever you remember to do a release. But if you have more than 1 thing, it's time to start thinking about it, because the "do whatever" that works great with 1 thing starts to break down when 1 becomes 2. That is not Kubernetes's fault. That's just the nature of the beast.
- being able to force better design when engineers have to work with multiple services depending on each other (using start-up containers, readiness and liveliness, etc)
If you use a shell script to handle a cluster with a dozen docker images and nodes, with intermittent crashes, out-of-memory issues, running out of disk, or network failures, your shell script will be so complex that you will basically need to recreate kubernetes from scratch.
Load balancing and auto-scaling are some of the more basic and popular services that kubernetes provides. If your platform already provides them for you, it makes the case for kubernetes weaker. That said, kubernetes can be run on multiple cloud providers, and provides far more features.
Once the complexity of the cluster grows, and vendor lock-in issues increase, kubernetes makes more and more sense.
Even if we had a monolith instead, and no plans for an on-prem offering, ultimately I still think that a managed Kubernetes offering makes sense. Efficient resource utilization is all the more important for small companies, and once you have more than a handful of servers, Kubernetes makes it much easier to right-size your fleet by handling all of the scheduling for you.
If you have a couple of servers and that's it, then sure, Kubernetes isn't giving you much. But if you're building a professional offering then you're likely to outgrow that couple of servers pretty quickly.
Yeah, I'm pretty sure we can't simply run it all from a single server.
-Rolling deployments
-auto scaling pods and nodes
-health checks and self healing
In saying that, I've just performed a reasonably significant migration to k8s, and despite the significant time investment, I am quite happy with it.
However, prior to this we spent 5 years on Dokku[1], which is freaking fantastic project and was for the most part more than enough to meet our needs. It solves pretty much all the same issues start-ups use k8s for, with a lot less overhead.
The reason for the migration was simply that we've reached the point where our clients demand improved reliability (redundancy) and somewhat coincidentally we'd outgrown some other infrastructure; which on its own would require a large migration. So we moved to k8s at the same time as rejigging our infrastructure, for redundancy and future-proofing purposes.
* DNS. This lets all the various components talk to each other across nodes (Airflow, Spark, Zeppelin, etc.). I can also, via VPN, connect to things to look at what's going on.
* Load-balancing. I create pods and then they run on the node where there's capacity without me thinking about it.
* Auto-scaling. I spin up pods and nodes get created to handle the load.
* Helm charts for things so I don't have to figure out how to run them myself.
* Built-in support in the tools I use. Airflow will spin up k8s pods to run tasks. Spark will spin up pods to run a job. Etc.
Granted, I still do have a bunch of "opinionated" scripts, but they delegate the heavy lifting to Kustomize, and all the configuration itself is just stock-standard k8s resources.
I do like helm's interface around packages: install, upgrade and test. I like that packages have a tree structure that you can compose (if you opt to do so). It'd be nice to extend the interface further to have other common actions you could apply to a more advanced package (e.g. backup, restore, chaos test, add tenant, remove tenant, etc).
The aim of all this is to allow you to avoid tricky in-place cluster upgrades and to just spin up a new cluster, direct traffic to it and tear down the old one. An extra benefit is that it would allow you to give each dev their own k8s cluster in parity with your live envs, but they could select only a subset of charts.
Check out the sample project[2] for more of an idea about how to use it (but the actual sample is probably broken since it's under heavy development at the moment).
May I suggest a middle ground =) github.com/stripe/skycfg (checkout //_examples/k8s that I added)
Edit: also the fact that you have a 4 month old outstanding pull request isn’t giving me any warm fuzzies about the maintenance and velocity of skycfg.
Yeah not sure what is going on with that pr, tbh i forgot about it b/c we implement it in our own runtime. I got a bunch of others merged quickly since then.
This is auto generated from a swagger API spec. Of course the intent is that you define your service and launch it with the client, but if you want a yaml file you can dump the objects to yaml - it would be far better than templating.
https://github.com/kubernetes-client/python/tree/master/kube...
or the docs here: https://github.com/kubernetes-client/python/tree/master/kube...
These are python classes that serialize to dict which you can dump to yaml. In other words, you can define your kubernetes spec as python classes, with variables, inheritance, and all the other good things that come with a programming language.
Once you have your kubernetes spec, you can either deploy it with the endpoint helpers (https://github.com/kubernetes-client/python/tree/master/kube...)
Or you can serialize to yaml if what you want is to generate a bunch of yaml. There is no need for a yaml template language.
Although, it's good that you pointed it out since I think it will help people that need fairly basic resources without doing something more complicated.