One small addition: Nomad recently implemented basic service discovery functionality into Nomad itself, so for simple use-cases you might get away with just a Nomad cluster without Consul running alongside: https://www.hashicorp.com/blog/nomad-service-discovery
There are many reasons to love Nomad, like its straight-forward architecture, how it plays well with other binaries to augment its functionality (Consul & Vault), super easy service mesh with Consul Connect + Traefik, and the ease of defining complex applications/architectures with the Job Spec [0].
I've used k8s a few times for work, but I just love Nomad.
The big thing for me is needing mTLS support for some of my services.
You can also go DIY with something like Traefik and just configure everything to require TLS with client TLS pass-through for enforced mutual TLS, that puts you into 'require a set of central load balancers' but you still get all the benefits.
HashiCorp has always tried to loosely couple our products so that while they work Better Together, they also work independently and allow users to choose the bits and pieces that fit their needs best.
My general suggestions for this question are that Terraform is what you should be creating your Nomad clusters with. In the cloud you can use it to define your autoscaling groups, load balancers, and use Packer to create golden images for Nomad servers and clients.
While Terraform is also useful for deploying "base" system jobs like log shippers or metrics aggregators, I generally don't recommend deploying all of your workloads with Terraform. You generally want more fine grained control over things like promoting-canaries-during-deployments than Terraform's architecture or interface could possibly support.
So now your left with how to run all of your workloads. Lots of folks hit the Nomad API or use the Nomad CLI straight from the CI runners, but Waypoint and Pack offer some useful tools.
Waypoint has more application developer oriented tooling than Pack. It's a developer experience oriented tool where Pack is more like Helm and oriented toward offering a package management solution for clusters. Since Waypoint manages building your project, it's a much more opinionated tool than Pack. You likely want to use Waypoint throughout your application development process.
Pack is much more useful if you just want to "package up" a bunch of existing code, whether your own or 3rd party, and run it in your cluster. It's intended to eventually have an extensive public registry/catalog of popular software ready to run in your cluster. This isn't an area Waypoint is concerned with at all.
Hope that helps!
This may be my favorite article on the subject, and I've read a lot! Others have suggested some minor corrections but overall you've done a fantastic job accurately and fairly portraying Nomad as far as I'm concerned.
I really appreciate the final paragraph as I constantly field questions about whether or not a k8s-only world is a foregone conclusion. I think there's room for more than one option, and I'm glad you noticed HashiCorp thinks so too!
I've been trying to get a MariaDB Galera cluster running on Docker Swarm for the past couple weeks and have been somewhat stymied a few things. Something like Longhorn would help a lot.
That being said they’re often glaring. So please report an issue to Nomad and the upstream of you hit an issue!
Nomad has 2 other storage features:
1. Host volumes where nodes can advertise a volume by name and jobs can request being placed on nodes with that volume available. I think a lot of folks run databases this way so they can statically define their database nodes and not worry about the numerous runtime points of failure CSI introduces.
2. Ephemeral disk stickiness and migration: during deployments an instance of a job’s intrinsic local storage can be set to either get reused on the same node (sticky) or migrated to wherever the new instance is placed.
Ephemeral disks is the oldest storage solution and is only a best effort (unlike CSI which may be able to reliably reattach and existing volume to a new node in the case of node failure). However I’ve heard of a lot of people happily using it for already distributed databases like Elasticsearch or Cassandra where the database can tolerate nomad’s best effort failing.
(Sorry for lack of links as I’m on mobile. Googling the key phrases you care about should get you to our docs quickly and easily)
Nomad also supports CNI. It's not uncommon for folks to run their network's control plane as a Nomad job and use a CNI plugin to integrate their other Nomad jobs with it. This sort of approach allows for running multiple logical networks within a single Nomad cluster (eg perhaps segments of your cluster use a service mesh while data intensive or legacy applications use host networking).
Personally our team found a match using docker swarm. The reality is that we don't have a DevOps role and we needed something that anyone could deploy on
Although the general sentiment is negative towards docker swarm, its simplicity is hard to beat
This guide was useful for setting up our stack https://dockerswarm.rocks/
Are you used to reading drunken writing? :)
I read this earlier too and was very impressed at how everything was explained in plain english. I wish I could send this article to myself 2-3 years ago, it would have saved me literally days of trying to figure out what parts of k8s do what and why.
Agreed. LWN is really good at this, partly because their editors have really good "tech BS" detectors (learned this when I was writing for them). It means you really have to understand what you're saying and can't just parrot buzzwords or lingo you've read on a tool's marketing website.
Clients (aka nodes-actually-running-your-work) can be run globally distributed without issue, and we've even added some more features to help tune their behavior:
https://www.hashicorp.com/blog/managing-applications-at-the-...
Disclaimer: I am the HashiCorp Nomad Engineering Team Lead
It would be great if you took some time to write about the question of climate change and energy waste.
In this regard it would be very interesting to better understand the impact tools like K8S have on global warming. Just from a pure practical POV I see many clusters burning a lot of CPU cycles with K8S processes that do nothing else than providing some kind of infrastructure for the real computing task.
I have seen K8S burning 50% and more CPU time - this with not one single payload operation, just doing its cluster thing!
This is a horrible waste of energy and we can not go ahead establishing such a bloated monstrosity as default data center OS of the future.
Not all data centers are 100% solar - I guess people in the field are LOL at this point knowing the realities. So I personally come to the conclusion that a thing like K8S is the opposite of what is needed in a world that is heading to a climate disaster. So many hours spent into optimizing the Linux kernel (successfully!) are nullified by a tool that wastes energy recklessly.
We need more than one article to help programmers of these tools understand why they should do their jobs with more love for future generations.
Also we need lots of articles about how to write efficient code that does not waste CPU cycles.
Wow. I really enjoyed this. Your writing was clear, simple and explained the landscape while navigating the marketing terms extremely well.
I wish to read other articles that you have written.
A very interesting series of articles would be "How to take down your K8S cluster" describing the many ways a cluster can be [mis]configured to easily be brought down by one single error / pod / application.
Thanks