The Future of Kubernetes
eficode.com
eficode.com
Everything gets SO FRICKIN' COMPLEX these days, even the very simple things. On the flip side, interesting things that used to be engineering problems is relegated to being a "cloud cost optimization" problem. You can just tell your HorizontalPodAutoscaler via some ungodly YAML incantation to deploy your clumsy server a thousand time in parallel across the rent-seekers' vast data centers. People write blogs on how you can host your blog using Serverless Edge-Cloud-Worker-Nodes with a WASM-transpiled SQLite-over-HTTP and whatnot. This article's author fantasizes and fetishizes interconnected lanscapes of OpenID-Connected workloads that are, and I QUOTE, "stateful applications that are stateless". Because "This means data should be handled externally to the application using other abstractions than filesystems, e.g., in databases or object stores." - yeah, fine and dandy, but which platforms of the future should these services be hosted on? And who will develop and shepherd and maintain and fix them?
Little parts of a system being over engineered may be the fault of an engineers ego somewhere, but the massive complexity of distributed systems isn’t some ego trip or frivolous fiddling with shiny new tech.
It’s just big and hard for mere mortals. It’s hard to say when k8s actually decreases complexity, but I think most will get there before ~100 software engineers do.
I lost track on how many hours (more like man-months, mythical or not) k8s has saved so far in my work, whether it was in team of 2 or team of 15 supporting multiple other teams.
On a highly functioning team with strong communication skills and no major egos it should go a lot smoother where everyone is leveling up on the knowledge at the same time.
Ten or fifteen years ago you'd be looking at a team of at least five, probably 7-10 engineers, mininmum, to design, maintain and keep up to date a similar system. The company gets a tremendous value out of their three engineers at $150-250k/year, ea
Managed Kubernetes services like EKS replace that with half a person or less + about $100/month overhead per cluster.
why not go further and call it "neo-feudalism"? The infrastructure providers are the territorial lords that control the arable land, the aristocracy of the digital world. Corporations that rent this infrastructure to build services on it are their vassals. They may claim to be freemen or nobility as they pay for their fief in money not in services and allegiance, but most of them are ministeriales that can't just move elsewhere without significant risk to business continuity (some can, like a oil mega-corp that just rents some IT infra). Lastly we have the cloud engineers, the serfs at the bottom of the pyramid, who work the land they don't own and have to give part of the profits of their work to the upper class in exchange for the opportunity to make a living on their infrastructure.
Next level for you: look at society and realize it works the same.
With that said, I'll echo your sentiment "Everything gets SO FRICKIN' COMPLEX these days, even the very simple things."
Nomad is now perhaps my only favored tool for container deployments/some orchestration. I also long for the days for the original container ... the folder! Put everything you need in and zip it up. instantly shareable and no spaghetti tooling needed to run it. The OS told you everything you needed to know about the process. I also oppose (my personal opinion) the trend to want to scale things crazy to numbers. I wholly believe on doing more with less. The company infrastructure stack I currently harbor admiration and put them on the pedestal is StackOverflow. Last I checked in 2021 it still ran on 4 MSSQL servers and 9 IIS servers with a few satellite services. Could easily be 4 MYSQL servers and 9 HTTPD servers.
It has a name - it is the Anti-Singularity, and it is coming for us all. As we blindly make ever more aspects of our lives completely dependent on increasingly complex and fragile systems, those very systems are becoming unmaintainable, despite ever-increasing numbers of engineer hours. One day, we will go too far, and will make a change that will break things. It won't be possible to go back, as there will be nothing to go back to. As the breaks ripple out, faster and faster, the Anti-Singularity will occur, freezing development in place permanently, and bringing productivity to a halt.
So I'm a bit skeptical about the long-term viability of the platform. Personally I think we'll see new paradigms emerge that will mostly obsolete containers and complex orchestration platforms, you can e.g. already see some pretty interesting ideas around WASM. Before Kubernetes OpenStack was the big thing, but now it's mostly used as a substrate to run Kubernetes on.
One of the reasons why we adopted Envoy at Ambassador for our API Gateway was precisely because there was no single _commercial_ entity who controlled Envoy.
As a note, technically Google & Lyft are VC-funded, so that’s true, although I wouldn’t put Google in the startup category, and Lyft’s business model is unrelated to Envoy.
1: https://www.businesswire.com/news/home/20210310005295/en/Tet...
Istio development is primarily driven by its steering committee [0][1], which is led by Google. If you look at the contributors [2], 6 of the top 10 are Google employees (perhaps more - not everyone lists their employer on GitHub). They have an open design process, posting design documents and similar to their Google Drive [3]. The team is likewise active at community events, where they solicit input that drives what they work on.
[0]: https://istio.io/latest/blog/2020/steering-changes/
[1]: https://docs.google.com/spreadsheets/d/1Dt-h9s8G7Wyt4r16ZVqc...
Here's the initial video where Matt Klein introduces Envoy: https://www.microservices.com/talks/lyfts-envoy-monolith-ser....
If you want to buy commercial support for Envoy & Istio, Tetrate certainly is an option.
(Note: I am/was the CEO at Ambassador Labs, which organized the Microservices Practitioner Summit where this talk occurred.)
I'm a little disappointed that you placed so much weight in that factor. I believe simplicity, reliability, and comprehensive documentation are far more important concerns than whether a commercial entity has control over a project. The flip side of community ownership is more politics, more complexity, bureaucratic decisionmaking, and slower evolution.
Connection/service proxies are a dime a dozen. I suspect you could have picked nginx, haproxy, or the Linkerd proxy and been just as well off.
This of course may be true for some communities, but certainly not the case for L7 proxies. In fact, at the time, the opposite was true: NGINX & HAProxy were content in slowly supporting their (captive) customer bases, with little incentive to innovate. This is why Lyft went off to build Envoy Proxy. At GA, it supported features such as HTTP/2, native observability, hitless reloads — none of which were easy / possible to do in NGINX or HAProxy at the time.
> Connection/service proxies are a dime a dozen. I suspect you could have picked nginx, haproxy, or the Linkerd proxy and been just as well off.
Linkerd proxy didn’t exist then (there was an early version of it, written in Scala IIRC; the Rust version didn’t come along to quite some time later).
I agree that there are many options for service proxies. I’ve found, though, that the Envoy community has been pushing innovation in the space quite aggressively.
I've been tinkering with a prototype that does this (and if anyone is interested in collaborating, my email is in my profile).
Also, in my experience, it turns out that you can quickly find yourself needing all those extra bits if you're not Google (which implements a lot of the same features in libraries linked to programs directly and depends on the fact that they run customized everything) or Netflix (similar).
I'm excited about WASM, but I don't see how it's fundamentally different when it comes to being built by venture-funded startups. Pretty much all of the interesting stuff in server-side WASM is being driven by venture-backed startups.
> WASM is 1% of what Kubernetes does, it's an alternative to runc nothing more
It'd be interesting if you could expand on this. I don't think WASM does _any_ of what Kubernetes does.
WASM has near instant startup, making a cold start for each request feasible. I think this could be a key property/advantage that makes it preferable and more economically efficient in the long run.
WASM I think has a stronger sandbox as well. Containers in a multi-tenant environment need to run inside a virtual machine. WASM can be run in multi-tenant without the need to isolate each tenant in a virtual machine, again making it more efficient in long run.
edit: also recalled a relevant tweet from creator of docker https://twitter.com/solomonstre/status/1111004913222324225
WASM just replaces JVM/BEAM/CLR/etc with its own virtual machine. It doesn't address all the stuff that container orchestration tools address such as lifecycle management, update management, artifact management, autoscaling, networking, management access control, logging & metrics, etc.
If you want to start faster than containers, you already can — by not using containers.
So given all this above, how is "starts faster than containers" an advantage?
And containers are not exactly slow to start. `time docker run -ti --rm ubuntu:20.04 true` => 0.5s. I suppose that's too slow for serverless cold start, but if you really care about that then the startup time of Ruby, Python, Node.js etc with all the packaged dependencies are also in that order.
Absolutely agree, it's just a core technology, which enables new model of operation, on which all that is necessary to run. Right now there is almost nothing.
> If you want to start faster than containers, you already can — by not using containers.
Only if don't care about security/isolation. WASM is strongly sandboxed. It's appropriate for multi-tenant systems, where containers/jvm/etc have to be run inside of virtual machines(e.g. firecracker) to get acceptable isolation.
> I suppose that's too slow for serverless cold start
Yep, half a second is way too slow for new instance per request. I think wasm startup time around 35 μs for lucet.
WASM is interesting client side, service side beside very simple use case like resize my image it's not very good imo
2. Even under interpretation, WASM doesn't require the interpreter to jump through hoops or optimization to get it to be performant. In this sense, it's more of a portable assembly format than something like LLVM IR.
Plenty of platforms would AOT compile it before execution or at installation time (like IBM/Unisys mainframes and microcomputers).
People are seeing WASM in the context of a distributed runtime. Obviously, this would require you to write your applications with something not POSIX-like, but it would allow you to scale your service on a cluster. Your unit of computation isn't just confined to a single machine.
How does kubernets do this ? I haven't hit that spot yet but I'm pretty sure I will.
I guess I'll also find an open source solution to avoid switching licenses.
I'm confused here, does WASM stand for something different besides WebAssembly because googling "kubernetes WASM' came up with nothing but WebAssembly. But I was under the impression that WebAssembly is just a virtual machine. I'm not much of a backend person but I don't see the overlap between container management (kubernetes) to a virtual machine (WASM).
Where have I seen this already?
Already been done:
Am I thinking of a different tech? Is “WebAssembly” == WASM?
WebAssembly is a new virtual machine that isn’t related to JavaScript. There’s no garbage collector or automatic memory management. It uses a binary compiled format for machine code instead of plain text. It’s kinda like Java if you replace the Java memory model with the C memory model, and the Java language with something that looks more like LLVM mashed up with Lisp.
i think the other key difference b/w nix and k8s is that a very large swath of k8s core contributors are employed by very large companies (mostly Google, Red Hat, Microsoft, VMware, and IBM) who are also deeply invested in k8s-derivatives sold in the marketplace. from what i've gathered, and this might not be true anymore, many contributions to Linux still come from the academy (universities, etc)
> 1998: Many major companies such as IBM, Compaq and Oracle announce their support for Linux
> 2000: Dell announces that it is now the No. 2 provider of Linux-based systems worldwide and the first major manufacturer to offer Linux across its full product line
> 2006: Oracle releases its own distribution of Red Hat Enterprise Linux.
No wonder that as IoT took off plenty of POSIX like clones with non-copyleft licenses sprung off.
Even on Android, the Linux kernel is the only GPL piece of the puzzle left standing.
- Corba, MRI -> http apis
- XML, Soap -> json
- Flash -> html5
- Perforce, svn, mercurial -> git
Etc. We can’t think of anything better than k8s right now, but that’s just how it works. The moment that new tech appears, we all will be asking ourselves “how could we possibly use k8s?”
In terms of the wire-format and the low-level concepts, HTTP2 and 3 are completely different from HTTP1.x. it's not an extension like HTTP1.1 was, it's a complete reinvention.
However, what they did was to reuse the high-level semantics and the "API" of HTTP1, so as a web developer you can pretty much pretend you're still using HTTP1 and let the browser and web server do the rest. This makes transition significantly easier.
I could imagine something similar with K9s. Maybe some other product will come along that supports (a subset of) the k9s API but does something different under the hood. I believe, k9s itself did so with Docker: You can still define k9s images using Dockerfiles, even though no piece of Docker is actually involved in the process.
However, it seems protobuffers (or alternatives like flatbuffers, cap'n proto) are still niches compared to JSON.
However I have yet to convince anyone at an org that isn't already using them that they will change things for the better.
Schema once, code everywhere. It is really nice.
Have you seen Buf? https://buf.build/
No, thanks for sharing! I will have a look!
The big thing is JSON has common basic building blocks (simple types) so you are unlikely to see the same class of interop problems like SOAP and friends had ... "oh look this defines a new type my Python soap / xml lib doesnt natively support" (happened to me a few times, most memorable was akamai soap api using a custom list-of-string type nothing in python would parse out of the box)
How was it before ? A bunch of unstandardized shell scripts and ansible playbooks that you could not reliably automate, let alone port to a different topology ? How could we possibly use that ?
No, the complexity of Kybernetes is mostly in its auto-tuning and auto-scaling features. That's the stuff that needs to be re-engineered from the ground up in an elegant way. If history is any guide, we'll probably end up with some variant of it in systemd sooner rather than later.
Also, I do wonder what "elegant way" would be, considering that k8s is a very elegant system (once you notice that YAML is just one of the serialization formats to make it easier for raw API access by humans).
Eh, nothing unstandardized about POSIX shell scripts, as opposed to a bunch of ever-changing YAML files with questionable semantics.
I think this generation of "cloud natives" have drunk too much of the cool aid served by "cloud providers" and their clever faux grassroots marketing/blogs.
As compared to the Linux and BSD "movement" (no such thing), which was about liberating ourselves from having to run apps on commercial OS's, based on a well-understood OS baseline (POSIX, a term coined by RMS).
Kubernetes is not a defecto and always better replacement of Ansible and shell scripts. In most small scale deployments, the value brought by Kubernetes is trumped by the costs of complexity. Bunch of scripts you say? Now you have bunch of yaml files laying around. Atleast before, you could find and fix problems manually quickly. Now, first find someone who has been in the weanies if kubernetes and it’s ecosystem to fix your problem.
In large scale infrastructure, kubernetes may bring the standardisation that is needed. Being able to hire people who knows kubernetes would be easier than hiring people who know how you wired your infra with ansible playbooks and scripts.
Portability you say? How easy do you think it is to move your infra(not just the apps, but - users, security rules, storage etc) to another kubernetes provider? EKS to AKS? Do you think AWS IAM and whatever equivalent Azure uses would be straight portable? Your S3 buckets and it’s policies to Azure block storage? Selling K8s with infra portability is a lie.
Applications - yes, but that is achievable by other means as well.
However, moving from 1 cluster in one region to another within another, is very simple. And that did not exists before kubernetes.
The reason k8s took off is because it provides the fief that devops engineers so desperately need to be motivated to stay on the job. (And that's mostly a good thing. Look at software testing to see what happens when you don't give developers career incentives.)
Portability between cloud providers is not imaginary. Same for portability of network configuration.
HTML5 is definitely not simpler than Flash was. I'd argue, it's far more complex, if you include all the sub-technologies and frameworks required to reach feature parity.
There was also no "natural" progression from Flash to HTML5. Oracle and the browser vendors had to run an involved depreciation campaign to get developers off flash and move them to HTML5. If you had left the web ecosystem on it's own, my guess is, we'd still be writing flash applications today.
> Perforce, svn, mercurial -> git
This progression was definitely natural, but I doubt again that git is less complex than svn or mercurial. However, one big advantage of git is that it is extremely easy to set up: You can just install it and use it locally without bothering any admins or setting up any servers. Once you are already using it, it's also easy to connect your repo to a server.
So, I'd argue, git has better UX and better availability than the alternatives, but it's not necessarily simpler.
I think most people don't need full feature parity. That's why a lot of these techs take off. They need something that is easier to understand.
Similar with kubernetes, there's like 5 permutations of software deployment people actually want to do. Deployments are one of those but there are still a few areas that you have to get hacky with StatefulSets to achieve. Once some PaaS-wrapper for kubernetes takes off we'll see the V2 of that idea as a stand alone cluster management thing. If someone nips some ideas from MaaS [0] we might even see on-prem bare metal take off again once people realize they're getting screwed on perf/bandwidth and many of the benefits of cloud providers are already provided to you by something like Kube.
[0] - https://maas.io/
As far as I remember Mercurial was easier to start with but Github came along and won.
Git is very simple. There are basically just four concepts: blobs, trees, commits, and pointers to commits.
However, git is far from easy! The UX is haphazard and confusing and many people are scared of venturing outside of their couple of memorized commands.
>Perforce, svn, mercurial -> git
Not sure I agree with those two. Edit: Someone beats me to it.
25 years later and still, worse is better.
The commentary around PersistedVolumes is true (don’t write user data as json pls) but simple. We had a few use-cases for big persisted volumes or even host volumes managed by a daemonset, where object store performance didn’t cut it. One was indexes for a search service. You can’t push these 64gb hunks of data into S3 without paying a ton, much better to use a file system. Same thing with translation string data, a PV with a local SQLite database or other indexed read-optimized format is 100x faster and 100x less expensive than “object storage”. We did enforce that the PV was always read-only for the “main” container in a pod, so you needed to add a sidecar or daemonset if you wanted to sync data there.
Where I work, we use Kubernetes very very heavily and some of the article reflects some of my experience, in the sense that any challenge we have, the community response is that we should add MORE complexity. Using secrets is complex? Rearchitect your whole app to use OIDC tokens is the answer (for an example from the article).
I think that Kubernetes was designed, released, and funded, to attack the AWS hegemony. As such, it was designed to provide a cloud abstraction layer for developers to target instead of the AWS API(s). As a result, it is designed for an organization to run a single cluster with many relatively trivial workloads (enter Helm for packaging "off the shelf" apps).
However, it turned out for some of us, the need is to run the same non-trivial application in multiple regions of multiple Cloud Providers. Here, Kubernetes really falls short, because Cloud Service Providers don't seem particularly motivated to make Kubernetes-based applications more portable, and Kubernetes remains a very leaky abstraction in terms of versions, storage, and networking.
To me, the future of Kubernetes looks like the community rallying around a single "distro" of Kubernetes that is reasonably consistent across all major cloud providers, but also makes basic assumptions such as how the CI/CD pipeline, observability and SLOs, self-managed packaging, etc... Here, I think the Linux analogy holds, in the sense that different Linux distros emerged for different sets of users, and communities were able to rallying around solution sets and provide support to each other.
With that experience in mind, the storage section of the article is a hand-wave. Kubernetes is a major platform for high performance databases themselves. Saying that Kubernetes developers should "use databases" is not a very plausible answer in this case.
Object storage is not a panacea either. It does not exist equally in all locations and building caches to make it work efficiently in low-latency apps is a non-trivial exercise. (Like years of implementation work by PhDs if you want something good.) Locally attached NVMe SSD is still the winner for a lot of applications. It needs to be first class in Kubernetes and it needs to be easy to allocate pods that offer it. Applications can deal with resilience through app-level replication. Databases have been doing it for years.
I think the tone of the article writes this as a positive, but I do not see this as a good thing. Looking at the ecosystem of k8s today, all that has happened is the same concepts that existed in OSes have been reinvented. This future abstraction sounds like yet another migration, but I wasn't able to tell what the net benefit was aside from addressing some intellectual or academic concerns (x should not y).
I'll definitely try to keep my mind open, but, to borrow a sentence pattern often used in the post, more abstractions should not be seen as a better thing.
We need something simpler than K8s. Something that does not need a million abstractions, yet accomplishes the basic goals that we want: efficiency, scalability, security, reliability, reproducibility, immutability, etc.
I think the future we should be driving for is a distributed operating system. A real operating system whose primitives run distributed applications. K8s already reproduces features that exist in the kernel, and the kernel already does a lot of the heavy lifting of networking and containerization. What's left is to add userland and kernel features so apps can make a syscall or library call to tell the operating system how they'd like to scale, or deploy, or be traced, etc. This would simplify without adding abstractions. It would also be controlled by the open source world, rather than by corporations, making an incentive to finally get rid of crappy design patterns and make the right design be top priority.
Those of you who have used Mosix clusters already have an idea how simple and effective this can be. SSI clusters didn't take off because they were focused on unmodified applications, and sharing memory and threads across systems is hard. But by requiring applications to design themselves for a distributed OS, we can instead allow apps to run independently yet communicate seamlessly across systems as if they were just on one.
Kubernetes, as designed, is a framework and a rather leaky one in my not so humble and subjective opinion.
it accomplishes a pretty lofty goal (bin-packing, deployment, eventual reconcilation of service and generalising the network/storage so that you don't care about underlying infra) but it does so at a rather heavy cost of complexity and because you can swap out underlying pieces: if you're not on a cloud provider then you're probably going to spend the majority of your life maintaining it.
But, you probably want Nomad, if you're thinking about kubernetes seriously.
and lord help you if you are but choose to try to use the same approach across environments and eskew the markup on something like EKS. I've been thumping my head against trying to have a rather trivial HA service working on both a Hyper-V and EC2 cluster equally for weeks and it seems near impossible for reasons no leading lights in the space quite articulates.
Say you wanted to write an application that inspects the current state of jobs running in Nomad and modify them. You need to design something specifically to do that within nomad. Essentially you need to integrate their proprietary interface into your application.
In a distributed operating system you would just do ls /proc/host/*/* and then issue syscalls or write to files in /proc/ and /sys/. No proprietary interface, it's just operating system primitives.
I think it's so simple that it's hard for people to grasp how insanely easier this would be than any current system. Look up Mosix/OpenMosix (but that's an SSI; I'm not saying we should do SSI exactly, it's just an example of the simplicity).
I truly cannot wait. I love the premise but hate the complexity. Nomad stole my mindshare in 2016 and I would still choose it over k8s today for small applications. Even choosing it for managing VMs in some cases. (Disclaimer: worked at hashi but not on nomad nor did i use it while there).
The irony of UNIX having won, only to be replaced by System 360 roadmap.
20 years ago I was deploying EAR files into application servers without caring which OS they would run on.
Nowadays I package managed runtimes into containers, that I don't care if they run directly on top of a type 1 hypervisor, VM or whatever.
Basically what mainframe OS virtualization looks like nowadays and traces its roots back to System 360.
The best comparison, IMHO, is that Kubernetes is continuation of ideas espoused perhaps most visibly by IBM JES3 (and before that, ASP)
Half the game engines on that list target Linux, not 5%, not 15%, half.
Thanks to rent seeking by Microsoft (the cloud providers got rekt by license cost for Windows) and Googles strong push for Stadia (forcing game engines to support linux, which usually shared some amount of code with the game server) things have changed.
Now; most things being made are on Linux. (the actual development is still mostly happening on Windows Desktops though)
I would rather focus on building a product and clicking a couple buttons that allow millions to use it (managed solutions). Oh no it costs 10k, let me spend 1m+ on hiring people to play with configuration files.
Idk why companies are pushing for k8s at points when they could rather hire 10 engineers/tech team to build better and more features than fiddling around with y’all files and adding so many feedback loops through new “processes”.
Even when things work it takes a long time to A) get there organizationally and B) it still doesn’t work lol
Kubernetes gives a standardized set of abstractions and concepts that means hiring and development is easier, not harder.
Try to hold out for the next generation! It will surely arrive soon?
Do you use any container orchestrators today?
I worked on Kubernetes/containers migration at Airbnb, and I wouldn’t recommend a Kubernetes investment for a company with <300 engineers, who can dedicate 10+ people to dealing with it. At Notion we are quite happy with Amazon ECS.
systemd may have a complex implementation but it's usage is rather simple and consistent across distros so in that way it definitely simplifies on what came before it, which was an init system simple in its implementation but finicky and inconsistent in usage.
It's kind of like I'd much rather have a programming language that has a complex compiler but it takes care of so many details that its usage (programming) is simple - because I am not the one who has to deal with the compiler complexity - rather than making the jobs of the compiler devs easier, (why, if that's their primary job?), but making my everyday job programming in such language harder?
This is one of the many things that's really wrong with Kubernetes. A big selling point of it is binpacking; not wasting resources, but it can't even achieve that. You are just overcomplicating your infrastructure without the big win of almost-perfect resource utilization.
> Interestingly, Linux was the platform upon which we built everything a decade or more ago. Linux is still ubiquitous, and part of our stack, but few developers care much about it
You definitely care about it when something breaks in a container and you have to figure out what's going on. Abstractions are fine until something doesn't work.
We're operating in a niche so we'll never have millions of concurrent users. We have ~1000 customers, and number of users per customer follows a power law. Our largest customer by far currently has a ~100GB database and ~100 concurrent users.
Since we're relatively small (currently ~25 employees) we don't have the resources to "constantly" be changing our stack, and so I fear vendor lock-in.
At least Kubernetes seems to be something that'll be around for a while, and it's available from several cloud providers. So while it seems overkill it would appear it has that going for it. However the complexity does scare me a bit. It's a giant leap from our current "dumb" setup of a database and core service on a plain server.
Reading recommendations or helpful suggestions would be appreciated. Are there better options out there for small fish like us?
K8s is worth all the work cus then you can scale magically. You don’t need to scale magically. You just need to be able to easily deploy.
In fact if you think scale will never be an issue you could just run it on hardware, seems like everything would easily fit in a single server.
Personally k8s is least interesting in scaling, because you could easily scale with other options if you had money to burn
With K8s, the basics that "just work" are pretty wide, the bits that are specific to vendor that are necessary are usually small and contained, and for everything else I have extensible API that allows abstractions like crossplane.io
it does seem completely plausible that within the last year or so hosted environments are stable enough and standardized enough that k8s can be a target directly even for small projects. cool. i will try for my next one
Biggest possible issues were if you wanted to pass things like cloud credentials to your apps etc., but standard k8s resources worked pretty well with at most some extra annotations to for example link a Service of type LoadBalancer with specific pre-allocated IP.
Thank you, I've looked into Terraform before, but had half-way forgotten it. Thanks for the reminder, will have to check out again.
> seems like everything would easily fit in a single server
We probably need a few. Our client application and integrations do quite a lot of crunching at times. The largest client mentioned has a 16-core server with 64GB RAM for DB and integrations alone, and that is not over-provisioned by any means. They have two or three Citrix servers for the client in addition.
Obviously things would be a bit different with something web-based, but more crunching would have to move to server-side due to DB latency.
For smaller customers we do fit several of them on a single beefy server currently, and could very likely move forward with that.
That's why I'd prefer to avoid getting to married to the solutions from a single cloud vendor.
that said, that is a very very small task to change cloud providers and shouldn't deter you. also as you said, you can easily just start with one.
I assume that for whatever reasons going very close with one PaaS from any of the cloud vendors is out (I am not a fan of them, but that's specific to me being in Poland and not Silicon Valley, and I can't pay Silicon Valley prices)
You have ~25 employees, so you need to reduce toil - you can't just throw a person full time to massage and understand each custom thing.
1) Take care to run your components <https://12factor.net/> style. Even if you run them 'classic way' started from a shell script, it will make things easier to move around if necessary.
2) Don't use raw manifests - this gets tiring over time, and is prime example of "toil". Your time will be better served writing anything (whether fancy or not) that can manipulate some basic data structures and spit kubernetes resources in json format.
Once you have that, you can template your resources easily and enforce common standards and behaviours. This helps a lot especially when you don't have many employees - I'm not advocating for "employees as cogs", but if you all your applications follow a standard scheme for labelling, where credentials and data is stored, etc. - each of those small improvements reduce a bit of toil - and critical to maintaining that is removing the toil of manually keeping those standards implemented. In one job I might have spent a week tailoring a templated deployment for an application, but the next person (who was dealing with the setup for the first time ever!) configured another, similar application in one day.
It's a huge difference when all you need to start a new app is to fill in few blanks like names, maybe some credential, and maybe declare "ok, it needs a >25G temporary storage, and 5G persistent shared storage, plus should be available under this URL" and have the templates fill in the rest.
3) Implement some standards on observability and administration. It doesn't really matter whether you set up your own EFK stack/Grafana/Prometheus/Loki/Victoria Metrics/whatever or go with Datadog (ow, my wallet) or another vendor. K8s generally makes it (comparatively) easy to hook together.
Apply templating and standardisation from (2) here. Whether you're a developer or sysadmin or both, having it "just happen automagically" leads to happier people, less burnout, and ultimately more productivity.
4) Abuse the hell of k8s flexibility to get multiple deployments going on. The more expensive the underlying gear is, the more important it is. If your whole deployment is parameterised, and with things like crossplane to integrate tasks like "allocate $VENDORs postgresql offering", you can quickly setup, test/poc/do something fun/do a dog&pony show/etc. then teardown an environment. Can you do it with more classical setup? Sure. Personally I found k8s both faster in wall clock time & development time for this... and definitely less expensive.
Making deployment easy and generic is indeed a key point I have been thinking about. Support won't be able to handle much customer-specific bespoke solutions requiring special hand-holding.
Observability will also be key. Our customers typically call us first whenever something is broken, even if 99% of the time it's their own ERP or similar that's the issue. So support needs to be able to quickly identify that the relevant integration is running as expected on our side, and be able to tell the customer what's wrong.
That said, ability to go from vague complaint to both specific component and bird's eye view of the right cluster is great force multiplier for people.
As for delivering in more on-site bespoke situations, it might be worthwhile to be able to go down with minimal k3s or similar setup that can run inside single VM that you can provide as deployable OVA file to be imported into VMware or similar environments - as well as for building small on-site test environments, or even make a Chick-Fil-A style cluster out of three NUCs taped together and use it for dog&pony road trip, excuse me, sales demo ;)
(though it is a pretty brief stream of consciousness)
Either that or Hashicorp Nomad, because it is also a pretty nice solution and seems to overall have a brighter future than Swarm does, even if the HCL which it uses is a little bit weird and needs porting over from Docker Compose files.
Admittedly, in the last year, i've also seen Kubernetes be tolerable for on-prem deployments in the form of K3s, which is a lovely project and wastes way less memory and CPU for smaller clusters: https://k3s.io/
I know that many out there won't necessarily have to or won't want to run it on-prem and instead will just run whatever Kubernetes solution will be offered to them by their cloud vendor, which is also a valid point, but not one that i have to deal with.
Instead, i often find myself with a VM that has about 8 GB of RAM and needs to run everything from the actual cluster, to PostgreSQL, RabbitMQ, Redis, a few Java apps and who knows what else inside of a development environment, because clearly that was possible without containers so also should be possible with them. There, running something like Rancher or even a K8s cluster that's made from scratch (or initialized with kubeadm or whatever) would be insanity. And having a shared cluster for the whole company might also necessitate people to just manage it, which isn't in the cards.
Also, in my personal hybrid cloud (homelab + cloud VPSes), running Kubernetes would be comparatively expensive and wasteful, since Docker Swarm still has a lower overhead and Swarm provides me with most of the functionality that i might want, oftentimes in the form of regular containers, e.g. Caddy/Nginx/httpd web server as an ingress with some configuration files.
Actually, anyone who has ever tried to get the K3s Traefik ingress working with a custom wildcard certificate will probably understand why i might actually prefer to do that, since in K3s you'll need the following for such a setup to work:
- a ConfigMap for Traefik, knowledge about the structure of the ConfigMap (tls.stores.default.defaultCertificate)
- a TLSSecret for storing the actual certificate/key
- a TLSStore (which i also needed to actually use the secret, spec.defaultCertificate.secretName)
- a HelmChartConfig for Traefik to load the ConfigMap with the mounted secrets and config
All of those weren't sufficiently documented because apparently the popular usecase was to use Let's Encrypt with either cert-manager or another solution altogether. All of it took digging through GitHub to find out and trial and error to configure it all.Versus just reading a page for a web server that i want to run and dropping some values in a config file.
- Use OpenID tokens instead of secrets
- Use Istio instead of ingress
- Use Knative
Sow the wind, reap the whirlwind.