How we use HashiCorp Nomad
blog.cloudflare.com
blog.cloudflare.com
When I started with kubernetes I converted my small Docker compose files to kubernetes files. Later I rewrote everything in helm charts. Now it's almost more YAML and golang templates lines than business logic lines in my applications.
I'm considering to go back to Docker compose files. It's simple, readable, and easy to maintain.
Jsonnet in general is a pretty damn good configuration language, not only for Kubernetes. And much more powerful than HCL.
It's a 'framework' vs 'library' thing, but in the devops context.
Same use-case but uses Starlark dsl instead of jsonnet
Also you can actually write pure code libraries in starlark (basically a subset of Python)
But I hate the UX of kubernetes. Compose was beautiful.
I have hope - since compose files have become an independent specification. https://www.docker.com/blog/announcing-the-compose-specifica...
Kubernetes distro built around the compose file specification would be a unicorn.
Or people have to keep inventing half baked things.
It is simpler than other tools, because you can get started without even touching jsonnet or python or anything else, when using our generators. It does more than all the other tools combined, as it replaces helm+helmfile+gitcrypt or kustomize. It’s universal, so you can use it on non-kubernetes situations where other tools leave you high and dry.
With the new generator library, you can have 1 template and use it to configure many services. Check our examples at https://github.com/kapicorp/kapitan-reference
We have just released: * https://github.com/kapicorp/kapitan-reference a repo with examples for quick-start. It includes our generator as explained in https://medium.com/kapitan-blog/keep-your-ship-together-with...
* https://github.com/kapicorp/tesoro a secret “webhook controller” to seamlessly handle Kapitan secrets in your cluster. Better than sealed-secrets because there is no need to convert secrets and it supports KMS like google and aws together. Get started with our blog: https://medium.com/kapitan-blog or join our kubernetes slack on #kapitan
- a deployment - a service - an ingress - a config map (or several) - a secret (or several)
It's even worse for stateful applications.
And each of the resource definitions is 60% boilerplate, 35% application-specific and 5% release- or environment-specific.
Helm would probably be a nice and neat tool if it had stopped at maintaining a simple map of variable names to values. But since applications need things like "if the user said SQLite, add a pvc, a configmap and a secret and refer to them in the ss, if she said Postgres, go pull another chart, deploy it with these parameters, then add this configmap and this secret and refer to them in the ss", Helm is an overcomplicated mess.
If you want to have complete control of what you're pushing to the API, use Helm as an starting point instead, run `helm template` and save the YAML output to some file, publish it using `kubectl` or some other rollout tool. I recommend using `kapp` [1] for rollouts.
Kinda loses the benefit of Helm hooks, but if one is using Spinnaker, there are other ways to do the same thing hooks can do.
How do you make helm chart deployment declarative? `helm install` is not declarative (in my understanding `kubectl apply` is declarative and `kubectl create` is not. Let me know if my understanding of declarative is wrong). Thanks.
It is simpler than other tools, because you can get started without even touching jsonnet or python or anything else, when using our generators.
It does more than all the other tools combined, as it replaces helm+helmfile+gitcrypt or kustomize.
It's universal, so you can use it on non-kubernetes situations where other tools leave you high and dry.
With the new generator library, you can have 1 template and use it to configure many services. Check our examples at https://github.com/kapicorp/kapitan-reference
We have just released: * https://github.com/kapicorp/kapitan-reference a repo with examples for quick-start. It includes our generator as explained in https://medium.com/kapitan-blog/keep-your-ship-together-with.... * https://github.com/kapicorp/tesoro a secret "webhook controller" to seamlessly handle Kapitan secrets in your cluster. Better than sealed-secrets because there is no need to convert secrets and it supports KMS like google and aws together.
Get started with our blog: https://medium.com/kapitan-blog or join our kubernetes slack on #kapitan
Pay close attention to the section where they described why they went with Nomad (Simple, Single-Binary, can schedule different workloads and not just docker, Integration with Consul). Nomad is so simple to run that you can run a cluster in dev mode with one command on your laptop (even MacOS native where you can try scheduling non-Docker workloads). I'd go so far as saying that it would even pay off to use Nomad as a job scheduler when you only have one machine and might have used systemd instead. You can wget the binary, write a small systemd unit for Nomad, then deploy any additional workloads with Nomad itself. By the time you have to scale to multiple machines you just add them to the cluster and don't have to rewrite your job definitions from systemd.
In my opinion, HashiCorp should look at what Rancher did with K3s and offer something like that, integrated with the entire Hashi stack. The only reason most people choose nomad is the (initial) simplicity of it (which quickly goes away once you realize how "on an island" you are with solutions for ingress etc). Deliver kube with that simplicity and Integration and it's a much more compelling story than what Nomad delivers today.
Also with Consul 1.8 and Nomad 0.11 you’ll get Consul Connect with Ingress gateways which solves some of those problems you mentioned.
Nomad template stanzas are simply consul-template (https://github.com/hashicorp/consul-template#plugins), you could use `{{ with secret "" }}` and never touch Consul. You also have every function beyond service and the key* set of functions. You could build pretty static, or dynamic configurations using these blocks without ever touching Consul. On top of that, both Terraform templates and Levant work well for templating job specs out themselves, which contain template stanzas in them.
An example of something helpful would be if you wanted to drop a very small binary that changes with each deploy, you could use https://github.com/hashicorp/consul-template#base64decode and just change the contents each time you deploy the job.
If you wanted to use redis keys in your consul-template, simply drop a binary on each server as a plugin to consul-template, then for example: `{{ "my-key" | plugin "redis" "get" }}`.
Why not try running it yourself?
`nomad agent -dev` will get you a server running, then `nomad job init` will give you a full spec which you can run with `nomad job run example.job`.
"Nomad schedules workloads of various types across a cluster of generic hosts. Because of this, placement is not known in advance and you will need to use service discovery to connect tasks to other services deployed across your cluster. Nomad integrates with Consul to provide service discovery and monitoring."
So one of the very basic features of an orchestrator, which is service discovery, requires consul. Sure, I can use it without consul to just start jobs that don't communicate, but obviously that's that the normal use case of it, and you can see that by looking at their issues list.
It's rather disingenuous to compare to Kubernetes-based installations by pointing out that Nomad is a single Go binary that's easy to install, because they're also running Consul and unimog. For what it's worth, kubelet (which schedules jobs, like Nomad) is also a single Go binary that's easy to install, as is kube-proxy (which helps services running on a node to send traffic to the right node, so, roughly analogous to Consul).
If anything, the author practically makes the case against Nomad. Practically speaking, the only reason why CloudFlare adopted Nomad is because the stars aligned on their stack. Most companies will actually reduce complexity by adopting Kubernetes instead, compared to running two different schedulers for jobs and externally-managed service discovery.
But in really, a single binary should be at the bottom of your list of concerns.
I wasn't able to immediately track down the alleged new repo for hyperkube, though
From cloudflare’s perspective, K8s is a riskier choice, precisely because it’s large and multifunctional. If Nomad (or Consul) upstream went south, CF would probably be able to maintain it without seriously growing their headcount. And they’d probably be playing the lead role in that ensemble.
If kubernetes upstream went south, the above statements would almost certainly not hold.
For them, picking a stack means picking something they need to stick to, potentially for the next decade or two. When your decisions are made with that long term in focus, the biggest, shiniest, and most exciting community is not always the right choice.
Like you said, it seems like this is the right choice for CF. It’s probably not the right choice for $TECH_STARTUP, for whom pinning the future of business critical infrastructure to Google and RedHat’s stewardship of K8s is the least of their concerns, and that’s OK.
First - companies will want service discovery for all managed resources. Consul is just outstanding, so there’s that.
On-prem, nomad is simple to manage and it gives you a form of freedom that k8s locks up in a black box.
After all, it’s a process scheduler, go ahead and schedule what you want. A container, raw exec or batch job. Add parameters to the batch job and dispatch them at will.
If you start from scratch or can move absolutely everything to k8s? Well sure, but not a lot of places you’d call “enterprise” isn’t in that position.
Of the two Nomad seems much more sane because it does one thing only and is much simpler to manage and deploy.
That said, having have used it, we are mostly moving away from it. Consul + Docker/Docker-compose with systemd in "a service per vm" model has proved much easier to administrate to our scale (couple of datacenters, ~1k VM mark). It is actually so simple that is boring. Instead of fiddling with infrastructure the developers spend time solving business problems...
Our stack is really boring, developers can find their way around ansible and share galaxy roles to configure / provision infrastructure... Other cloud native projects like Prometheus and Fluentd give us all the visibility we need in a very straight forward and boring way.
p.s: Great article!
I see kubernetes as more useful when organizationally you have more heterogeneous services and separate dev/sre groups. This allows them to divide up responsibilities more easily.
I have a set of developers that can focus on their code and not have to worry about the infrastructure, TLS certs, DNS, storage or whatever. That all gets abstracted away.
For SRE, that gives one a common set of orchestration tooling one can develop against.
In the 300k person org it was for scale and reliability. In the 35 person org it's so that 1 guy can manage everything without trying.
I think the approach I mention ends up being a good compromise from both sides. If I were starting a company from scratch I would certainly consider it differently.
I was motivated by a move away from config management in favor of commands wrapping Nomad API calls. For a "devs on-call" model this was preferable to being gatekeepers of PRs against config management.
To glue the whole thing together I got Consul Connect going in the Nomad jobs, so service config complexity was comparable to docker-compose.
Saying this out loud makes me realize it was the organizational model I was pursuing that led me here (so called "production ownership"). And I'm not a big config management fan ;). I take it you have dedicated Ops teams or Devs willing to learn Ansible well?
It is all pretty standardized, is rare that someone needs to create a new galaxy role for example.
We expect our teams to know their way around Ansible ( which hasn’t been a problem ), we have internal docs and training
Terraform was a huge breath of fresh air after Cloudformation. If you ask me, deploying apps via k8s is even better.
Everyone wants a free lunch and for me, momentum + cloud vendor support + ultimately the Nomad Enterprise features that come free with K8s made the choice easy.
I'd rather learn Go than deal with JavaScript. Python would be fine. Elixir would be great (I'd much prefer Erlang but there are only so many miracles I'm allotted in this lifetime). Perl, please.
I never understand when people say operators don't like programming languages. Always seems weird to me that a person that constantly works with computers would prefer a static markup language instead of full programming language for managing the complexity of operating software in production environments.
My original comment was about this attitude. I personally think programming languages are much better than markup languages for almost everything. If you are presenting static content then markup languages are fine but if you're doing anything more than presenting static content then YAML just makes no sense.
A person who constantly works with computers is painfully aware of the people's ability to screw things up by an honest mistake when programming them. The more complex the thing, the easier it is to make a mistake, because human attention is finite.
Whenever you can limit the language to the task at hand, and build in guard rails that would prevent you from doing things you should not be doing, it's usually a productivity win, even if it precludes implementation of 3% of some most advanced and complex solutions. Reliability and simplicity go hand in hand, and reliability is often the thing devops people value very highly.
The solution to finite attention isn't less powerful tools. The solution is more powerful tools that augment operators and let them extend their finite attention spans into longer time horizons.
How is your description and viewpoint different from what I just described?
What I won't like is the "here are some classes and functions for accessing our API, go build the stuff somehow" approach.
And it makes sense. As we demand more flexibility to not repeat ourselves (DRY) we get clever and add little features.
Suddenly you can stick an entire cluster config (or several!!) inside a lambda that returns the declared resources... And you've gone straight to hell. You're more than Turing complete by then and engineers are reinventing conditionals and loops with lambdas all over your config codebase.
Might as well have just started and left it at Python/etc.
Disclaimer: I worked on Terraform at HashiCorp and am a contributor to Pulumi also.
Really dig using with it Cloud Run.
Pulumi has two different k8s modules, the standard one (pulumi-kubernetes) and "kubernetes-x" which has a much nicer API that reduces boilerplate:
https://github.com/pulumi/pulumi-kubernetesx
The second one doesn't seem to be used in the Pulumi tutorial content and has less stars/recognition.
Are these interchangeable? If so, why not use @pulumi/kubernetesx instead of the standard k8s module everywhere?
To my mind, the real power of having general purpose languages building the declarative model instead of a DSL is that libraries like k-x (or indeed AWS-x) can actually be built!
Their okta provider, for example, is based on some absolutely ancient, non-standard provider (https://github.com/pulumi/pulumi-okta/blob/v2.1.2/provider/g... ) instead of using the official one which receives regular updates: https://github.com/terraform-providers/terraform-provider-ok... ; and now I guess it's on me to do some git diff to find out if there are meaningful changes between the imported provider and the upstream one?
To say nothing of the more esoteric providers that one can cheaply build and place in the `.terraform.d/plugins` directory and off you go, in contrast to trying to find the dark magick required to use some combination of https://github.com/pulumi/pulumi-terraform-bridge and https://github.com/pulumi/pulumi-tf-provider-boilerplate but ... err, ... then one has to own that new generated pulumi provider code? So like a fork but worse? Maybe they should make a bot that tracks the terraform-provider topic on GH and uses their insider knowledge to generate pulumi providers for them.
Don't get me wrong: I anxiously await the death of TF, but until there is a good story to tell my colleagues about an alternative to it, they'll continue to use it.
Humans shouldn't write CloudFormation templates. If you're doing that here is a important free clue: use CDK.
One advantage of CDK over Terraform is that the 'state' is the CloudFormation stacks themselves as opposed to Terraform's state system. Another is that you get a real programming language (your choice of several) using CDK.
The disadvantage of CDK is that it's AWS only.
Also a massive warning to anyone wanting to use hard cpu limits and cgroups (they do work in nomad it’s just not trivial), they don’t work like anyone expects and need to be heavily tested.
I had the exact same problem with Kubernetes (and even just straight up containers), I just want to make sure folks are extremely aware of how big of a footgun it is, and to really, really, really test it well.
I have only tested it with Docker, again, that company Moby with this very little used software.
How dare their core systems have any systems that are not extremely simple to understand and have caveats.
The cost/maintenance trade-off works when you have more SPOF management hosts than Nomad servers (5). You decrease host images down to 2, Nomad server and client, versus N management images.
Though it does sound a bit like they're using config management rather than pre-build images.
Bonus, Nomad servers are more failure resistant using Raft consensus versus any N management hosts. And for discovery I found the optimal pattern is to put all of the Nomad servers in a "cluster" A record for clients to easily join (pattern works well for Consul too)
How are they making a profit when AFAIK all other CDNs charge you for the bandwidth (and I assume they have to pay for to their providers)?
There are people with free accounts moving GBs every month through their network and I imagine those free users must account for a very large percentage of their traffic.
Well, to a certain extent. I didn't mean to imply that their monthly spend for bandwidth is zero -- I'm sure they aren't anywhere close to peering 100% of their traffic, they aren't in the DFZ, and, of course, they've gotta pay somebody (or, more correctly, several somebodies) to connect all of those datacenters together!
---
Long answer:
About six years ago, they described their connectivity in "The Relative Cost of Bandwidth Around the World" [0]:
> "In CloudFlare's case, unlike Netflix, at this time, all our peering is currently "settlement free," meaning we don't pay for it. Therefore, the more we peer the less we pay for bandwidth."
> "Currently, we peer around 45% of our total traffic globally (depending on the time of day), across nearly 3,000 different peering sessions."
Remember, that was about six years ago. I wouldn't be surprised if both their peers and peering sessions have increased by an order of magnitude since that article was published -- just think of all the datacenters that they're in today that weren't back then, especially outside of North America!
Additionally, they've got an "open" policy when it comes to peering, as well as a presence damn near everywhere [1,2]. Since they're "mostly outbound", the eyeball networks will come to them, wanting to peer.
Running an anycast network and being "everywhere" also has some other benefits. They perform large-scale "traffic engineering" -- deciding which prefixes they advertise where, when, and to who, and the freedom to change that on the fly -- so they've got tremendous control over where traffic comes in to and, perhaps more importantly, exits from their network (bandwidth is ~15x more expensive in Africa, Australia and South America than North America, example).
So, yes, CloudFlare is still paying for transit but, at their level, it's relatively "dirt cheap". Plus, in addition to the increases mentioned above, bandwidth is likely an order of magnitude cheaper -- at least -- than it was six years ago!
---
EDIT:
Two years later, in August 2016, CloudFlare published an update [3] to the article linked above. A few highlights:
> "Since August 2014, we have tripled the number of our data centers from 28 to 86, with more to come."
> "CloudFlare has an “open peering” policy, and participates at nearly 150 internet exchanges, more than any other company."
> "... of the traffic that we are currently able to serve locally in Africa, we manage to peer about 90% ..."
> ".... we can peer 100% of our traffic in the Middle East ..."
> "Today, however, there are six expensive networks that are more than an order of magnitude more expensive than other bandwidth providers around the globe ... these six networks represent less than 6% of the traffic but nearly 50% of our bandwidth costs."
---
[0]: https://blog.cloudflare.com/the-relative-cost-of-bandwidth-a...
[1]: https://www.peeringdb.com/asn/13335
[2]: https://bgp.he.net/AS13335#_ix
[3]: https://blog.cloudflare.com/bandwidth-costs-around-the-world...
The Service is offered primarily as a platform to cache and serve web pages and websites. Unless explicitly included as a part of a Paid Service purchased by you, you agree to use the Service solely for the purpose of serving web pages as viewed through a web browser or other functionally equivalent applications and rendering Hypertext Markup Language (HTML) or other functional equivalents. Use of the Service for serving video (unless purchased separately as a Paid Service) or a disproportionate percentage of pictures, audio files, or other non-HTML content, is prohibited."
In other words, as soon as you start using significant CDN bandwidth for images, video, audio, etc. they will contact you and ask you to upgrade your account.
They aren't. In Q1 they made a loss of $33M on $91M in revenue.
In fact it seems they have been loosing more and more money over the years.
https://finance.yahoo.com/quote/NET/financials/
What are they counting on happening here? Obviously free customers won't switch to paid plans.
Colocation providers charge hundreds of dollars to have an (allegedly) unmetered 100 Mbps network link.
AWS charges network usage per GB. It's the same flat number if you use half the capacity half the time or so.
What do all providers have in common? They charge enough to cover the costs of providing the service and make a profit.
Still a mystery to me why "balancing" has SO MUCH mindshare. This is almost certainly not the optimal strategy for user experience. It is going to be much better to drain traffic away from older machines while newer machines stay fully loaded, rather than running every machine at equal utilization factor.
You are right that even balancing of utilization across servers with different hardware is not necessarily the optimal strategy. But keeping faster machines busy while slower machines are idle would not be better.
This is because the time to service a request is only partly determined by the time it takes while being processed on a CPU somewhere. It's also determined by the time that the request has to wait to get hold of a CPU (which can happen at many points in the processing of a request). As the utilization of a server gets higher, it becomes more likely that requests on that server will end up waiting in a queue at some point (queuing theory comes into play, so the effects are very non-linear).
Furthermore, most of the increase in server performance in the last 10 years has been due to adding more cores, and non-core improvements (e.g. cache sizes). Single thread performance has increased, but more modestly.
Putting those things together, if you have an old server that is almost idle, and a new server that is busy, then a connection to the old server will actually see better performance.
There are other factors to consider. The most important duty of Unimog is to ensure that when the demand on a data center approaches its capacity, no server becomes overloaded (i.e. its utilization goes above some threshold where response latency starts to degrade rapidly). Most of the time, our data centers have a good margin of spare capacity, and so it would be possible to avoid overloading servers without needing to balance the load evenly. But we still need to be confident that if there is a sudden burst of demand on one of our data centers, it will be balanced evenly. The easiest way to demonstrate that is to balance the load evenly long before it becomes strictly necessary. That way, if the ongoing evolution of our hardware and software stack introduces some new challenge to balancing the load evenly, it will be relatively easy to diagnose it and get it addressed.
So, even load balancing might not be the optimal strategy, but it is a good and simple one. It's the approach we use today, but we've discussed more sophisticated approaches, and at some point we might revisit this.
I'm sure your system has its benefits, I just get triggered by "load balancing" since it is so pervasive while also being a highly misleading and defective metaphor.
With nomad, you have too many options, you can run your service/job as a process or within containers, but with Kubernetes, there is only one way to do things -- container, which is really a simple choice to make.
Besides, Kubernetes got etcd builtin, so I don't have to use deploy and maintain consul.
Last but not least, I still see containers mysteriously gone, and have no idea how nomad did that. With kubernetes, such thing never happened.
[1] https://www.hashicorp.com/blog/hashicorp-nomad-autoscaling-t...
what can you achieve with nomad that you can’t achieve with terraform?
- have a restart policy outside of what the docker daemon offers
- automatically schedule across multiple machines based on constraints including affinity and anti-affinity, bin packing etc
- “dispatch” jobs
The Terraform Docker provider effectively provides a tiny subset of the functionality of Nomad, which provides similar functionality across many drivers, not just Docker.
I like the quality of products from HashiCorp, but k8s is far, far, ahead of where nomad is.
What I really want is better integration for Terraform and kubernetes. The current TF k8s support leaves much to be desired, too many things are missing or broken and I find there are several bugs that result in deployment flapping (i.e., constantly re-deploying the same thing when there are no changes).
https://www.nomadproject.io/docs/job-specification/lifecycle...
- no quotas for teams/projects/organizations
- no preemption (ie higher priority job preempts lower priority job)
- no namespacing
So generally it's somewhat useless in organizations where there are multiple different teams that should be able to coexist on a cluster without stepping on eachothers' toes, or even where you want a CI system to access the cluster in a safe manner.
Also, we run on bare metal, so there really isn't a way to request extra capacity within seconds.
We fully intend to migrate more features to OSS in the future -- especially as we build out more enterprise features. As you can imagine building a sustainable business is quite the balancing act, and there's constant internal discussion.
(I'm the Nomad Team Lead at HashiCorp)
It's also going to make the operators very unhappy because it'll be harder to monitor actual memory utilization (allocated memory vs memory really in use) in order to plan cluster extensions. Are there some tools around or work planned to make this kind of scaling and utilization easier?
Please file an issue or link me to an existing issue of you have time. This seems really compelling!
For my personal use I don't need them and for a business Enterprise won't break the bank. You'd be surprised how much your management might be interested in having an escalation path if everything goes south and you need a hot fix. I can vouch that Enterprise support is worth it if you're on a small team that can't spend all day on this stuff
* No autoscaling
* Can't reserve entire CPU core: https://github.com/hashicorp/nomad/issues/289
* No way to run jobs sequentially: https://github.com/hashicorp/nomad/issues/419
Disclaimer: I work at HashiCorp in Product Management.
I hope you'll also consider sequential tasks in a single jobfile. Without it it's kind of awkward because you have to be passing shell scripts around just to run a setup step before your actual workload.
https://learn.hashicorp.com/nomad/task-deps/interjob
There’s more work going on there to further improve the feature in upcoming releases.
Disclaimer: I’m an engineer on the Nomad team.
One of the best features is that you can bundle a CR to request a MySQL database and it will be satisfied with whatever config is in your cluster so that app only declares the need but not care how it’s done.
Disclaimer: I’m one of the maintainers of Crossplane.
The only reason they're pursuing this strategy is because they think they can get a piece of the kubernetes market. If that doesn't pan out, they will dump nomad like a bad habit.
For-profit companies aren't open source charities.
Isn't exactly declarative, nor really imperative, mostly but not entirely idempotent, and worst of all, it has faux-state that may or may not reflect the actual state of the services.
I used to love Terraform, but it broke my heart. Now we're done.