Tanka: Our way of deploying to Kubernetes
grafana.com
grafana.com
There are many ways to classify workloads, but one big distinction that we find valuable is between stable infrastructure components and our own rapidly deployed services. The former have complicated configuration surfaces but change relatively rarely, while the latter are usually much simpler (because microservices) and change daily in many cases.
We find helm works very well for the infrastructure pieces. Yes it's complicated and easy to get wrong, but so are most other package management systems. Charts for complicated things can be quite dense and hard to comprehend. See the stable prometheus chart for an excellent example. But once the work is done and as long as there is commitment to maintain it the results can be pretty awesome. We use helm charts to install and upgrade prometheus, fluentd, elasticsearch, nginx, etcd and a ton of other tools. Yes we've had to fork some charts because the configuration surface wasn't fully exposed, but they are a minority.
For our own services charts are overkill. They're hard to read, and crufted up with control structures and substitution vars. Essentially all of our microservices are deployments with simple ingress requirements. We currently use kustomize to process base manifests that live in the service repos and combine them with environment-specific patches from config repos. Both are just straight yaml and very easy for backend developers to read and understand, and different groups (i.e. devops, sre, backend dev) can contribute patches to the config repos to control how the services are deployed.
Bottom line: if you're going all-in on kubernetes, which you really need to do to get the most benefit from it, then you're going to need more than one approach to deploying workloads.
https://github.com/helm/charts/blob/master/stable/prometheus...
https://github.com/helm/charts/tree/master/stable/prometheus...
In the end we got rid of it and replaced it with Terraform. If your infrastructure is 100% kubernetes then I think helm is great. Our infrastructure is not. We have databases, dns, buckets, service accounts and more so we were splitting our setup between terraform and helm. Passing data between the two tools was going to be a pain. We follow a layered approach of building up the infrastructure.
1) Networking: DNS 2) secrets, service accounts, buckets 3) DBs 4) Pre-application config (istio) 5) Services
Semi-related things are together and all of those cloud provider values we need are saved as secrets. We are on GCP so that means we need things like service accounts to access GCP resources (buckets, cloudsql) and all of those variables are available to our services to pick up.
And Terraform has STATE. This is unbelievably valuable when doing continuous delivery as you can tell what changed on every deploy and deploys are FAST. One thing that really bugged me about helm was that determining if a deploy failed was a post helm event. We were going to have to write monitoring for service health/uptime on deploy. This is not hard at all but you get it for free with terraform. If a service failed to start, terraform will throw an error...
I don't think people know that Terraform has a kubernetes provider. It does not support all the alpha objects but has decent support for 99% of the things you need. I wish someone made a provider for istio virtual services and service entries.
I found the documentation and it’s special syntax a pain. Just give me json with json schemas defines somewhere so vscode can give me great code completion and I can templatize as necessary
Is there something that does this in the wild?
Common, almost ubiquitous in my circles, and often reliable are things I would say about it, painless is not. I was a bigger fan until using it daily focused on a project for a month on it more than anything else.
I feel this strongly enough I'm looking at AWS CDK and troposphere, likely to learn I suck at writing my own imperative style terraform and to be more thankful. Lol
I really don't like how TF holds it state and it is errorpron. Allone how you have tto keep tf version in sync feels wrong.
The k8s support did not take us far. We had some kubectl script and it doesn't feel like a first class citizen. I wouldn't bet on tf for k8s.
I do love argocd. Great product.
I find state to be invaluable when a deploy of anything fails.
With ksonnet having gone quiet[0] this looks like a promising initiative.
I'd imagine that it'll need something like a package manager (or at least a curated list of common packages) in order to gain good adoption.
[0] - https://blogs.vmware.com/cloudnative/2019/02/05/welcoming-he...
In the end I decided that I'd collapse all environments down to behave identically, which is simpler, but does add a few constraints for development in particular.
Will take another browse through while considering options for upcoming infrastructure :)
[1]: https://github.com/grafana/jsonnet-libs
[2]: https://github.com/grafana/loki/tree/master/production/ksonn...
[1] https://www.terraform.io/docs/providers/kubernetes/guides/ge...
An example would be creating a dynamically names S3 bucket and passing it onto your service to use/manage. The same goes with anything needing credentials. Very powerful.
I'd be very tempted to just use terraform. But helm, with our forked charts, extensive values.yaml files, etc. has permiated the deployment.
So... rather than try to break up the helm beast; I leave it and keep the separation of concerns (infrastructure vs k8s) as described at the beginning. It works pretty well to be honest is rarely the source of problems.
Note: It uses the kubernetes terrform provider and only support those objects
For example, in our case we do `default_deployment.libsonnet` to which we add arbitrary overrides like `default_deployment + dev` or `default_dep + us_central1_overrides` depending on the cluster. And in these case we don't need to pre-specify what can be overriden in the `default_deployment` which makes it really powerful (also can come bite you in the arse though).
You might have to overload it somehow to make it work the way you're describing though.
But honestly, if you're going for the latter, might as well just use Kustomize and accept the chaos.
I really hope we'll get to a dominant standard soon. But this subject is much more complex than I thought https://github.com/kubernetes/community/blob/master/contribu...
Nobody can neglect the power of helm charts (because so many already exist), so I think Tanka will add support soon.
Even though we are focused on Jsonnet right now, this does not necessarily need to stay that way forever.
Grafana for example is popular because it supports multiple datasources, Tanka should probably do as well (e.g. Jsonnet, CUE, Helm, whatever results in JSON)
It's like the infrastructure version of fine woodworking - building dovetails and screwless joints by hand, using chisels and hand planes and card scrapers and shit, to build a box. It may be "fun", but it's also needlessly complicated and time-consuming. Give me the power tools, pocket hole jigs, torx screws, nail guns, square clamps. Yes, the dovetails will make a more sturdy box - but do you need a box with dovetails? Probably not.
What I actually want is something to create a cluster in AWS to run my app, or "create me a Fargate ECS cluster using R53, ACM, ALB, ECR, Lambda, CodeBuild, CloudWatch, SG, VPC, and IAM". If I feed that thing my AWS account's root credentials, it should first go through a list of default variables, explaining each to me and why I might want to change them. Then it should just create all of the above for me, and probably save it all as code. Next, I create a repository with my app and a Dockerfile. I then run a "hook up my git repo webhook and deploy my app right now" script, passing it my github creds. That should create a webhook into CodeBuild where my app is turned into a container, pushed to ECR, and then deployed into the Fargate container, as well as creating any CloudWatch alerts needed.
That's what I call a "product solution": an opinionated, complete solution that does everything I need for me, with a very light user interface and guidance on how to use it. Probably 95% of the above is two custom Terraform modules and some glue, and I should only have to answer like three questions by default.
If Terraform itself shipped with the above complete solution, that would be what I'm looking for. But I'm not aware of a catalog of solutions like that. Occasionally random people will publish partial modules on GitHub, but those usually need more modules and glue to actually work. So I'm basically looking for "no assembly required" solutions with the option to modify them later.
It’s designed by the BCL/GCL author as a replacement (Jsonnet is apparently a copy of BCL/GCL)
Currently we're mostly keeping a close look at CUE, but not really using it as of right now. However, during the holiday break I've been trying to get into CUE again and there are some things I need to figure out before being able to tell how to incorporate or replace some of our jsonnet projects with CUE, if we really want that.
Some parts of CUE seem like an obvious improvement to what jsonnet currently offers. So 2020 will be exciting in that regard.
We have chosen Jsonnet because it already had an ecosystem and served us well, but Tanka is open to other languages as well
I am also impressed by the clip at which CUE is improving, and how useful it already is, in spite of its relatively young age. It reminds me of Python or Go in its focus on tooling quality, and stability.
Their community slack is very welcoming, too. https://cuelang.slack.com
off(but actually on)-topic: I can't wait for the use of slack for open source communities to die in a fire, not only because the onboarding experience is horrible, and search is a mess, but about 75% of them appear to be the non-paid flavor so old messages are just held for ransom
1 = https://mattermost.com/nonprofit/ asks for a $390 "setup fee" and only for 1,000 users
I haven't seen another appealing solution in the space. I recommend using Kustomize if you can, but it is very limited in what you can do with it. When I looked at dhall-kubernetes it was lacking some crucial defaulting features that are getting integrated into the language now.
Further we use jsonnet a lot (including generating dashboards) and in general found it much more powerful and useful compared to plain YAML. From a customary glance, I can't find if kustomize supports jsonnet.
Helm felt okay for PnP, but I want to have an explicit understanding of what I’m deploying for infra, and it seemed to abstract too much away.
Kustomize seemed too rigid.
Ksonnet seemed too magical, although I didn’t deeply look.
I still don’t love using jsonnet, as I can’t seem to find full language documentation even on the website for it.
How might this compare to kubecfg to those who might be familiar?
1. Documentation is bad. We know that and work on improving that. Some resources that might help:
- https://tanka.dev/jsonnet/overview: Our own docs include some notes about Jsonnet in general for newcomers
- https://jsonnet.org/learning/tutorial.html: Taking the time to read this entire page opens eyes. Annoying and time consuming, I know but worth it.
Regarding Ksonnet and Kubecfg:
1. Ksonnet was magical. Tanka hopes to be Ksonnet without the magic. We got rid of all of those concepts, parameter merging and whatnot. You have Jsonnet and Tanka, a tool that pushes Jsonnet to Kubernetes. That's it. (ok, you also get a lot of handy features like CLI completion, diff and other things to make your dev experience better)
2. Kubecfg is similar, but has a smaller scope. It evaluates Jsonnet and pipes this to kubectl (basically). At the time we started Tanka, kubecfg was by the way part of the deprecated Ksonnet project, so we assumed it dead as well. Luckily it's not, as it is a very cool project, that inspired Tanka a lot.
Tanka after all aims to be like the `go` command: The one command you need to manage your entire complex kubernetes clusters. Also, Tanka is not strictly limited to a single language. For now we focus on Jsonnet, but more may come in the future
Hope this project get a lot of traction!
What I am doing for my env clusters is to have a versioned production yaml that acts as a source of truth, then if I need an env (regions, customer, dev, prod, feature, etc..) I take that source of truth, apply a transformation (usually a node script or bash... depending on the kubernetes entity) and then apply the resulting transformed yaml. Basically is: versioned production => transform => new env definitions
Do you have any recommendation/high level thoughts on how to integrate or substitute Tanka in this approach? Which are the downfalls that you see with this approach?
Thanks.
Integrating should be quite straigthforward. Install Tanka, create a new project (tk init), copy your source of truth YAML (without transformations) somewhere under lib/ (for example lib/foo).
Then go to lib/foo/foo, and import each of those yaml files:
{
foo: {
deployment: import "./deployment.yaml",
service: import "./service.yaml",
}
}
In environments/default/main.jsonnet: (import "foo/foo.jsonnet") + {
// patch environment specific things here
// https://tanka.dev/tutorial/environments#patching
}
Then use `tk show` to verify it works.Furthermore, follow the tutorial to get an in depth understanding of Tanka: https://tanka.dev/tutorial/overview
Do you know if there is an example or open source cluster using Tanka that I can try on minikube or a test cluster?
it would be really helpful to see how is the workflow to move from feature to feature, how are the envs when there is a bug, how do you replace the volume info from the cloud provider to minikube and all the considerations of the patched envs
Do you have a good pattern on how to use it with CI/CD for deployments? The biggest challenge we've had after writing deployments is getting it setup to work with something like Jenkins (right now we have a custom bash script that does a bunch kubectl things).
(PS any way this would help with static IPs on hosted Grafana.com Cloud to make access to firewalled datasources easier?)
We build an api and a cli util for these things.
The api takes care of providing sane deployment manifests based on data from a service registry, which is populated using the cli.
Cli can be run non-interactive for use in pipeline and other automations.
The cli and api basically do any config and deplyment related tasks in concert with service discovery (which at this time is consul).
Edit, rant: For automations json just owns yaml in all the ways!
This yaml templating business should stop! I think it’s silly tbh. :)
Our base k8s setup consisted of roughly 40k lines of yaml. Yaml is awesome for 10 lines of user input configuration, not 100, let alone 40k machine printed stuff. What happened here in this brave new DevOps world?!
1. A way to templatize YAML (though it could be used with Tanka too - the render command is split from the deploy command for this reason)
2. Monitors the rollout of resources - it's possible to detect successful and failed deployments more easily.
3. A way to run one-off commands during deployment.
4. Uses kubectl under the hood so easy to add to any workflow.
It has other features that I don't use including secrets management.Net net - I definitely think it could elegantly replace a bunch of bash scripts that do kubectl things.
One of the things we need to do is elastically scale the number of tasks (basically pods in tekton) that comprise our test suite run. This might be based on cluster utilization or whether it's the master branch. Since we have a single threaded test suite, we hoist parallelism up to the k8s level by breaking apart the tests into partitions each run by its own pod. For this we just render processed and parameterized erb to yaml. Eventually we'll dispense with yaml altogether and programmatically construct resources using a k8s REST api client.
We haven't moved into CD with any of this tooling yet.
I think this highly depends on your deployment process .. do you want full continuous deployment (CI deploying to the cluster)? In this case you could continue to use for example Jenkins to run `tk apply <environment>` on each merged PR.
Another option would be to use an in-cluster CD agent (for example https://fluxcd.io/), which uses Tanka to generate the yaml and applies it. Flux can be used with Tanka, needs some setup though: https://docs.fluxcd.io/en/1.17.0/references/fluxyaml-config-.... I guess we could simplify this in the future.
Feel free to reach out to me on Slack http://slack.raintank.io/ in the #tanka channel :D
I'm on dot net. And although I can deploy as microservices ( clean architecture with core, application, Infrastructure and api).
I seem to integrate the api into my app ( Eg. Add the api dll). So my app does the provisioning like a monolith.
It exposes all api controllers by default.
Messaging is internal always then ( domain vs Integration).
Overhead is practically none.
If I have a heavy component/api, I can split up an API and put nginx in front of it for routing and nats for Integration Events
So, basically I have a DDD app at the beginning with the strangler pattern already in place for scaling porpose. Although none of my apps need scaling right now.
I also can do every deployment myself and more easily. Since I don't have a deployment complexity currently.
--
What I don't have, is that my stack is language agnostic at the beginning. But it could be using the same method as scaling, with nginx.
It seems that I have the best of both worlds at the beginning.
- maintainability by forcing DDD
- minimal devops
- testability
- no service mesh overhead ( eg. Consult brings a 30-50ms average overhead, I finish most of my requests in 8-12ms)
- fast development ( slower than monolith, much faster then microservice)
While scaling could be refactored within the day, if an insane amount of request come in ( see: refactoring)
Most microservices are fixed within a single language though. So that's not a concern currently.
The added benefit is, is that I have insane custom implementation options.
I just need to change the Infrastructure in a deployment to use a clients database as a source if a component needs it.
( Eg. An order service for a webshop. I can easily integrate with an clients existing magento for a niche of their shop)
TLDR: I currently don't have a devops overhead. I'm too small for that, I'm glad though.
--
if anyone thinks that isn't a good solution for my use-case ( small dev shop) or have any better ideas. Please share ;)
Apparently the imperial features are there for good reasons.
And it's even more brittle...
A dozen microservices, several diff DBs, and some large stateful datasets, all supporting a basic Rest API in the end. What tools would you choose nowadays?
Jsonnet is a godsend. Don’t use a string templating language for structured data like yaml/json. Use an object templating language like jsonnet. You’ll start to love life again.
We had used mustache templates before and it was a PITA.
- they are bound to a single file. This won’t help you when trying to maintain multiple similar sets of Config
- anchors do not support patching. If you need to change a nested key, you can’t do so without it affecting all other nested keys as well
For instance, if we want to share steps in a YAML based CI config we can't if they are a list.
It seems to me like Terraform is good at describing desired deployment shapes and detecting drift between actual state and desired state.
Can someone clue me into why Terraform hasn't caught on as the abstraction above/that drives K8S?
The issue with Kubernetes is more expressing the state. Kube uses YAML, which quickly becomes verbose and hard to maintain. More on our blogpost: https://grafana.com/blog/2020/01/09/introducing-tanka-our-wa...
Tanka is trying to solve this issue by providing a more powerful language that overcomes these limitations hopefully.
There is a bit of a movement, however, behind using it to deploy software by pairing it with Packer. You use Packer to create an e.g. AMI whose sole job it is to run your software (like a Docker container) then use Terraform to launch a bunch of EC2 instances that have juuuust enough resources to effectively run your software. That'd allow you to eliminate k8s from your stack, though it remains to be seen which stack would be more cost-efficient to run on.
I have about 40 kubernetes services all as modules using the kubernetes terraform provider. I think I have 1000+ pods running on our one cluster all deployed through terraform.
It works very well because I can chain infrastructure resources into my service deployments. For example, I can create a dynamically named bucket and pass the name of that bucket as configmap/secret into my service to use.
We ended up using TF's Helm provider, sometimes with hacks like a helm chart which deploys an arbitrary YAML file (the so-called "raw" chart). At that point, Terraform is blind to what's actually happening inside K8S. You can still benefit from the ability of TF to pass data from your other infra automation into the Helm charts, of course, but it's really Helm actually managing the configuration of your K8S cluster. And that's the app we all love to hate.
The situation may have been improved, but my conclusion was that it would always be a somewhat incomplete interface.
For those things we use a direct kubectl yaml provider.
I wish there was an istio provider!
Actually, YAML has anchors and aliases, which helps a lot when the same thing needs to be reused in several places.
Tanka does generating, because it seems to be more robust, as the tool understands the output (instead of string substituting a fragile syntax)
See the docs on more details about generating: https://tanka.dev/tutorial/k-lib
Would that work for you?
I would really like first class support for templating on Grafana though. My desired workflow would be to use github for the templates source of truth. Then a dsl on Grafana's end would regenerate the actual config whenever the template was changed. The web editor could provide a fallback method of making changes if it could commit to source control. For instance if I saw a cool demo for graph functionality online I could implement it in the UI and then dig into the autogenerated commit in git. That userflow usually works really with teamcity and their kotlin dsl.
Edit: Teamcity gets around the style issues of autogenerated code by submitting changes from the UI as patches. First the kotlin dsl is run on the base config, then patches are applied. That way you are able to rewrite an autogenerated project config in the style and approach you want, and any UI generated changes will be limited to one folder.
General:
1. We run everything on Kubernetes
2. We configure it using Tanka, keep all Jsonnet in git
3. Changes are done using PullRequest
---
Grafana (the software) related:
1. We use provisioning: https://grafana.com/docs/grafana/latest/administration/provi.... This means dashboards are kept in .json files on the filesystem and loaded on startup
2. Those dashboard json files are ConfigMaps
3. Those ConfigMaps are created using Tanka
4. The content of those ConfigMaps is created using grafonnet-lib
This means, our dashboards are source-controlled! A change to a dashboard is reviewed, merged and automatically deployed (ConfigMap is changed, Grafana restarted, picks that up and done!)
---
Caveats:
- Edit dashboard in Grafana won't work anymore
- You need to mess with files - BUT this might change in the future, Jsonnet might become integrated to Grafana. Stay tuned :D
With Tanka you would usually use the functions provided by `k.libsonnet` to generate your manifests.
For CRD's, there are no helpers available, so you either just write the plain manifest as a Jsonnet object (JSON syntax), as YAML and `import` it (gives you an object as well) or even better write some helper functions yourself, publish them as a library on GitHub and make future users happy :D
Also what about state changes? I.e calculate the diff between your local definition and Cluster state and act appropriately (delete, apply, change)
Diff: Use `tk diff`. It shows the differences between the local Jsonnet and the cluster. `tk apply` makes them reality afterwards.
While Jsonnet is technically a superset, all you will see during use is most probably function calls and imports.
You won't have to actually write the Kubernetes objects anymore, they are generated using helper functions, just like real programming languages would do.
[0]: https://cuelang.org/
The site is hosted on Netlify and seems to be working for most other people.
Maybe reset your browser cache, check your network or try on your phone using cellular data instead?
If the issue persists, I'll take a closer look :D
Actually, YAML has anchors and aliases, which help a lot when the same thing needs to be reused in several places.