Why Kubernetes is important for the future of data platforms
mertkavi.com
mertkavi.com
Also the complexity cost significantly adds to the 'configuration that outperforms a single thread'. 99% of the time you'll be better off putting a docker container on a single box.
It can do a lot of cool stuff, but there's plenty of complexity there, and there are quite a few potential nasty surprises in there (from a security standpoint anyway, which is where I focus)
This isn't too say that throwing more code and complexity can't help. But it is probably no different than any other help, in that most gains are probably slow to realize and more marginal than you'd hope.
This is changing. For databases the usual approach is to write an operator. Most databases these days typically have one (or more).
Sorry, what did you mean here? This doesn't make a lot of sense.
Kubernetes is so complicated and many organizations do not really think about what problems they are trying to solve before diving in. Kubernetes solves so many problems that most IT organizations don't actually have or that are perfectly solvable in a simpler way with the existing virtual-machine ecosystem.
Of course if you're not a human person it's different. Corporate people, being made up of multiple different people, care more about their human elements being able to interact with each other's work. There containers might help.
I am pretty sure they're supposed to use a somewhat "Cooperative Namespacing" schema, that makes both the application and the OS/container abstraction layer more complex, while attempting to deliver lower overhead for said namespacing. Sorta like Cooperative Multitasking, compared to Preemptive Multitasking. And even that doesn't work for some situations; as an example, old versions of Java (7 and older) will use the incorrectly namespaced Linux file system APIs to incorrectly read whole-system limitations instead of current container limitations to setup memory and CPU parameters.
> remove a lot of pointless attack surface.
Not really, containers are for lower overhead cooperative namespacing, not security. You want bare minimum lightweight VMs for security.
A kickass CICD setup and good developer tools probably should be done first for the typical company, otherwise the majority of the added engineering time will go into scaling the infrastructure setup.
Even if you don't use user or network namespaces it at least confers a desirable separation of application and its dependencies vs. underlying host at the filesystem level.
It simplifies administration and allows your app and its dependencies to move at a pace decoupled from the host kernel and its minimal dependencies required for providing ssh access and systemd.
There's a concept called COST - Configuration that Outperforms a Single Thread, which basically asks the question: "how many cores/servers does a system need to outperform a single thread/core". Often this means that to outperform a single beefy server you need to get to 100+ servers to overcome the coordination costs and actually do more processing or process more faster. Lots of workload specific details here, but the general application in data engineering is "How many nodes does your spark cluster need to outperform my laptop, or how many nodes does your cluster need to outperform a single turbo.large.beef instance?"
My belief is that most companies chase scalability wrong in data.
That said, there are some companies that handle telemetry at scale and actually need to record and analyze billions of events per second. In these cases you will trivially hit the scaling limits of a single server. 99% of companies are not here.
I could be convinced that this number is actually as low as 90% :)
> Resource efficiency
What does this mean? More efficient than what?
> Seamless scalability on containers or nodes
Not sure what “on containers” means. “On nodes”, sure. Do other platforms not solve this as easily? (Running multiple workloads per host)
> Built-in scheduler
I guess same question as above, there aren’t other solutions where you give something a fleet of hosts and it runs things on them?
> Cloud flexibility
Again compared to what I guess?
> Pushing to apply engineering best practices
Such as?
Than normal processes that are platform dependent.
> Not sure what “on containers” means. “On nodes”, sure. Do other platforms not solve this as easily? (Running multiple workloads per host)
On containers means your application is packaged to be used as a consumable service, that is in turn used to compose a more complex workflow service that solves a ‘real’ problem. I guess other platforms also solve this, but the concept of containers (and Docker streamlined it) was the first. Everyone followed later.
> I guess same question as above, there aren’t other solutions where you give something a fleet of hosts and it runs things on them?
There are years of bin-packing research (think operations research from 50’s). But none of them came with the integration that something like K8S brought. Mesos is another example, but comercial support is slowly killing it.
> Again compared to what I guess?
Compared to having heterogeneity, even in your local environment, let alone in a production system.
> Such as?
Agile. CI/CD.
* containerization is good in general
* k8s' binpacking is good for efficiency
---
> Than normal processes that are platform dependent.
Why/how is k8s more efficient? Do you just mean binpacking?
> I guess other platforms also solve this, but the concept of containers (and Docker streamlined it) was the first. Everyone followed later.
"We should use containers" is different than "we should use k8s".
> Compared to having heterogeneity, even in your local environment, let alone in a production system.
Sorry I meant "compared to what other systems" not "what's the alternative to cloud flexibility?"
> Agile. CI/CD.
Here again I guess you just mean via containerization in general?
You don’t need to use k8s. But if you want scalability and manageability, you have to solve the same kind of problems it tries to solve, because ultimately resources aren’t infinite. Of course, this will always depend on the size of your (multiple) service(s).
And while we needed something like Kubernetes this past decade to grow microservice architecture en mass, many of things that Kubernetes does right now will be abstracted away over time in a way that maintains much of that same flexibility.
In that respect, yes, Kubernetes is important for the future of data platforms - but as a stepping stone (an angle that was absent in the author's post).
I don't see how the second point necessarily follows from the first.
And because of this, there isn't a set of standard tools and practices that can work for all Web apps, and when things change it often requires big workflow changes as well. In other words, for anything other than a simple Web app (which could be workflowed with GitOps) you have to have a good SRE to keep things running with Kubernetes ;-)
Need to update your cluster? Rerun cdk/pulumi/terraform.
So are a gazillion other solutions that do no claim that they are essential for the future of data platforms.
A good Kubernetes cluster will also give you centralized logging/monitoring/alerting as well.
So if you already have a cluster (and an administrator), you do get benefits. But the benefits are only really worth it if you can amortize the cost of the cluster administration over many applications.