K8s Service Meshes: The Bill Comes Due
matduggan.com
matduggan.com
K8S has service routing rules, network policies, access policies, and can be extended up the wazoo with whatever CNI you choose.
It’s similar to Helm, in that Helm puts a DSL (values.yaml) on top of a DSL (go templates) on top of a DSL (k8s yaml), just that it is routing, authentication, and encryption on top.. well, routing (service route keys), authentication (netpols), and encryption.
It boggles the mind!
The promise is that all of the inter-app communication channels are fully instrumented for you. Four things mainly 1) mTLS between the pods 2) Network resilience machinery (rate limiting, timeouts, retries, circuit breakers). 3) Fine grained traffic routing/splitting/shifting. And 4) telemetry with a huge ecosystem of integrated visualization apps.
Arguably, in any reasonably large application, you're going to need all of these eventually. The core idea behind the service mesh is that you don't need to implement any of this yourself. And you certainly don't want to duplicate all of this in each of your dozens of microservices! The service mesh can do all the non-differentiated work. Your services can focus on their core competency. Nice story, right?
In reality, it's a little different. Istio is a resource hog (I've evaluated Linkerd which is slightly less heavy weight but still). Rule of thumb: For every node with 8 CPUs, expect your service mesh to consume at least a CPU. If you're using smaller nodes on smaller clusters, the overhead is absurd. After setting up your k8s cluster + service mesh, you might not have room for your app.
Second, as you mention, k8s has evolved. And much of this can be done, or even done better, in k8s directly. Or by using a thinner proxy layer to only do a handful of service-mesh-like tasks.
Third, do you really need all that? Like I said, eventually you probably do if you get huge. But a service mesh seems like buying a gigantic family bus just in case you happen to have a few dozen kids.
https://learn.microsoft.com/en-us/aspnet/core/grpc/loadbalan...
All major hacks are 0-days (well, not updated Wordpress is not necessarily 0-day; a lot of 0-days are exploited months or years later), stolen credentials (social engineering usually), brute force password hacks or applications that are left open (root/root for mysql with 3306 open to the world). Those have nothing to do with (un)encrypted traffic.
https://docs.docker.com/engine/reference/run/#runtime-privil...
Unless you perfectly drop all privileges from every pod you are open to attack.
Containers are not security contexts, they are namespaces, that require all actors that can launch a VM to actively drop privileges.
This is an intentional design decision and not a bug.
If you aren't a Google, Apple, Microsoft, ...etc scale company than a service mesh might be a tad overkill
There are plenty of languages and RPC frameworks where you can solve this without resorting to a service mesh.
Practically, and to your point, service meshes solve an organizational problem, not a technical one.
On that scale I'd expect people to use client-selected replicated services (like SMTP), and never something that centralizes connections (no matter where it's close to).
You can always add observability at the endpoints. Unless your infrastructure is very unusual (like some random part of it costing millions of times more for no good reason, as on the cloud), this is not a big challenge; you add it to the frameworks your people use. (I mean, people don't go and design an entire service with whatever random components they pick, or do they?)
Small note, unless it has changed recently, containerd default capabilities list includes CAP_NET_RAW, so hostNetwork=true pods can sniff all traffic.
Of course there's tradeoffs, but I think it's a specific perspective that says that Kubernetes is any more complex than virtualizing the hardware and scheduling multiple VMs across real hardware.
It's a curated, extensible API that provides a decent abstraction over a heterogeneous collection of hardware. Nobody's done that before, and it's extraordinarily useful for being able to define intent.
No-one forces you to use YAML, it's just a serialisation format.
Once you have a bunch of components that implement this API, it becomes trivial to deploy pretty much any level of complexity of containerised application, without having to care too much about the actual exact details of how scheduling, networking, storage etc. is implemented. Even better, I can hit two different clusters configured in two completely different ways with the same manifest and get roughly the same result.
It's the abstraction.
That could also tell you that you're in a bubble of converts and have forgotten what it's like to live without k8s. ;) There are always at least two sides of the coin.
> No-one forces you to use YAML, it's just a serialisation format.
Really? So when do I get to describe my singular app that needs Postgres and Kafka in 7-10 lines in a .txt file and not {checks our company's ArgoCD backend repo} ... not 8 YAML files? I ain't got all day or a week, where is that? Nice declarative MINIMAL syntax with all the BS inferred for me (like names and stuff, why should I think of 5-10 names of stuff? Figure it out, automated tool!) that only concentrates on what is being deployed. It can generate everything else just fine, or at least it should, in a more sane world anyway.
Excuse the slightly combative tone, I ain't trolling you here but you also come across as a bit blind about things outside the k8s holy land.
> Once you have a bunch of components that implement this API, it becomes trivial to deploy pretty much any level of complexity
Nope, nothing ever becomes trivial with k8s. I was on call with two senior platform engineers and we needed 3-4 hours to describe a single small app that needs Postgres and Kafka and to listen on a single port while needing a single env var and a few super small config files. These guys provisioned an entire network of 500+ pods working perfectly for years. They made 10+ mistakes while trying to help me deploy this small app on Argo. (And they've done the same dozens of times at this point.)
Do with that info what you will but I'd strongly disagree if your takeaway is "they are not that good" -- because they have done quite a lot very successfully (with and on k8s).
> It's the abstraction.
My 22+ years of programming have taught me that people enjoy abstractions too much and make huge messes with them. I am not convinced that having an "abstraction" at all is even a good selling point anymore.
You originally said "I don't see how it's useful", and I observed that some people find it useful.
I'm sorry you've had a bad experience, but it's a bit short-sighted to extrapolate from that to "this is universally useless".
> you also come across as a bit blind about things outside the k8s holy land.
You're not the only one with multiple decades of experience in writing code and managing infrastructure. I've used a lot of the tools in the toolbox, I know when each is likely to be more or less appropriate.
Could you post those 7-10 lines needed to fully manage a Postgres AND Kafka deployment? I'm by no means a master, but I have a decent amount of experience outside the "k8s holy land" and I have no idea how to accomplish that.
I think for most places, most of the time Kubernetes is overkill. Cloud in general isn't great bang for the buck. But speaking to your anecdote, having actually deployed these things in the past 3-4 hours to deploy a custom application, AND Postgres AND Kafka is a pretty compelling use case FOR Kubernetes. It would certainly take me a lot longer doing a proper job of it managing the system directly or using something like Ansible.
I as a dev should not waste days on this -- something I have no experience in. I can add the basic service / ingress / whatever and move on with life.
The infra / platform guys can add HA and backups later, right?
> I think for most places, most of the time Kubernetes is overkill.
Then we agree. I am not sure this app in particular was a good fit for k8s though; having it on a DO droplet + managed Postgres + managed Kafka would have taken me 1 hour, 40 minutes of which would be me cursing because I forgot to put a host name and port combo somewhere. :D
> I think for most places, most of the time Kubernetes is overkill.
True, but as you yourself said, backups and high availability are much harder on bare metal -- for devs anyway. I can do an amazing job at it because I do it at home for my stuff buuuuuuuut, it's going to take me days, maybe even two work weeks. Not a good time investment.
That's why I am resenting it when I have to do it: it's wasting my time with something that I never learn properly so I have to relearn it every time from scratch because I can't engrave it in my memory and that is because I never do it for long enough stretches of time to indeed remember it.
I already complained to my manager about it and he took it seriously, by the way. I am not raging at you or others about a deficient process in my employer's company (that I plan to help fixing, or at least mitigate somewhat). I am displeased with the fact how everyone just settled on a very, VERY low maxima -- and somehow nobody wants to rock the boat.
We can do better. We should do better. But alas, guys like myself have worked 22+ years on stuff they don't love -- and I already have burnout and hate working programming for money, which is another tragedy entirely -- and will NEVER EVER get the chance to change anything at all because I have no choice but work for the man... for now. Though at 40+ I am not sure I have the strength to manage to somehow command huge commissions for consulting or w/e else. Anyhow, off-topic.
Again, we should do better.
Helm is still useful for consuming 3rd party charts, but IMO it's status as the "default" is more due to inertia more than good design.
If you really want a concrete recommendation try https://cdk8s.io/.
Well, that's comparing apples to oranges. Product teams have completely different goals, e.g. adoption/retention/engagement, so naturally internal cluster encryption is so far out of scope that in fact only the platform team can reasonably implement it. I don't see how that statement is relevant. You don't send an electrician to build a brick wall
Too many times have I seen architects and developers completely ignore it to make their jobs easier, leaving it to operations/infrastructure to implement. It's easy to twist the arm of business people with a "I can't ship feature X if you want me to look at security Y".
If everyone took this seriously perhaps we would have fewer issues.
You can, but it's absolutely a pain in the neck. Services need to load the certs from the filesystem on boot-up and trust the certs provided by other services. To manage trust, you need a certificate authority. Now you need to load the certificate authority's cert, and you need to manage rotation of certs. You need to help developers set up laptop-local certificate authorities and get them to issue certs so that you have Dev/Prod parity. You need to ensure that developers are enforcing modern ciphersuites, not doing bullshit "insecure-skip-verify" kind of toggles that make their jobs easier (because remember, their job isn't security, it's shipping features), not accepting self-signed certs or other certs not signed by the certificate authority. You need to make sure all this stuff is put in the testing suite to make sure it keeps getting maintained, and you need the files for these tests marked in CODEOWNERS to be under InfoSec control to ensure nobody rips them out just because they're inconvenient. And you need to copy this for every single service you run in production and every single development team.
You know what else you can do? Write your own web server (/sarcasm). I mean, who needs nginx? Probably writing your own will have lower overhead and less total complexity, not running a bunch of features that you don't use. And probably it will not be anywhere close to as good as a battle-hardened web server used by millions of engineers that gets regular support.
Personally I think it's debatable whether services really need mTLS within a private network. It's mostly a question of what scale you're running at; probably there's higher benefit-to-effort-ratio InfoSec projects to tackle. But if you do decide you need it, unless you can prove that the overhead is unworkable for your requirements, really you need to bite the bullet and put in a service mesh.
I don't get how you don't still need to do all that? Are the local server proxy to the service in plaintext then? Encryption is just between proxies?
I mean most app servers abstract away https on the server level and most dev is done unencrypted. So this seems reasonable.
True in my opinion. Its very complicated and the docs are confusing af.
Recently an overly confident security engineer came to us and demanded that we get service mesh because thats a SOX requirement. I have no idea from where these people get pipe dreams like that.
Rather than injecting sidecar containers that set up networking and so on, having pods join an existing SDN that just works with no app-side config would be a much more elegant solution.
Other than Cilium, I'm not aware of an SDN like ZeroTier or Wireguard that works seamlessly with Kubernetes this way (and which also works on managed Kubernetes like GKE and EKS).
Are there other ways of going about this for FedRAMP moderate or high IL?
What makes you think that?
Most service boundaries at organizations are "I need a different version of a pinned package we can't upgrade.". This is common in languages where there is support for only using one version of a given package, and it's worse if there isn't a culture of preserving function APIs. E.g. any python company with pandas/numpy in the stack will need to split the environments at some point in the future, no ifs ands or buts!
um... You should have been pinning all your dependencies from the very start
> Virtual environments doesn't solve any of the problems as python authors either don't specify the version or their dependency author doesn't specify it.
This is a problem of lack of training or willful poor practice, not an issue with venvs themselves which absolutely do solve this issue.
Agreed, the problem would not exist at all if an approach like that of say, node development, existed, but here we are.
You mean only your direct dependencies, right? And yes, of course
These are very solveable (and have been solved for a while) problems. It really pains me that most people are either unaware or do not dive deep enough below the surface to find them.
If you depend on dependency pinning due to unmaintained code then you should go deal with the problem directly. Say you pin 10 external libraries to three year old versions, how many security holes does that expose you to?
That's really my issue with dependency pinning, you end up with software that are just allowed to rot, making upgrades more difficult with every passing year.
> That's really my issue with dependency pinning, you end up with software that are just allowed to rot, making upgrades more difficult with every passing year
You seem to be suggesting not pinning and just pulling in the latest (minor, I presume?) versions every time you redeploy.
A better way is to pin _all_ your direct dependencies and check for minor version upgrades on a regular basis, especially for publically available services. A proper CI system should allow you to do this with confidence.
Occasionally you will be forced into major version upgrades, but in my experience this is rare.
This should be a regular and essential part of software maintenance and security vigilance.
I'd agree with that, I just don't believe that to be happening on any significant scale. Reasonably I think you should be able to do pinning like:
library>=3.0,<4
Depending on the versioning scheme of the library, but yes, pulling in minor release. We currently do that using Debian packages and just rely on the OS to provide security updates to underlying Python packages.Pinning is fine is you manage updates, but if you're then pulling in three year old libraries, which may come with their own pinned third party dependencies, I still think you're doing it wrong. In the example from GP it appears that they are pinning versions as a way to avoid forking and patching a deprecated and unmaintained package.
I'm not oppose to version pinning as such, but if you don't have a plan to stay on top of security updates, then you're better off pulling in any minor update. It's not just at redeploy, you may reasonably need to pull dependency updates more frequently than you redeploy.
If you deploy using Docker e.g. you're CI system needs to be constantly be pulling in updates, rebuilding and testing container images. People just don't do that. Most developers I worked with don't even care to update their container images if they aren't updating their own code. I've on multiple occasions had to deal with developers who absolutely lost their mind because we didn't automatically pull in OS updates, yet they themselves shipped outdated Java libraries or at one point even relied on an out date version of an alpha release of Tomcat that hadn't been updated for three years. The only different was that their dependencies where in containers and therefor, in their mind, "safe".
Why not use a venv or Docker? It's a world of difference. There are multiple issues with relying on the OS package manager for Python dependency management.
Overall it sounds like you know good practice and you know the value of venvs, pinning and Docker but you've had to deal with some workplaces with really shoddy practices. Still, that's no reason to say "venvs don't work" or to give up altogether on the notion of good practice by saying "People just don't do that".
I've worked in shoddy shops, I've worked in shops with excellent practices. In some places I've been the one that made them transition from the former to the latter.
After the Py2 vs. Py3 thing settled down, almost all of the operations issues got away.
That said, Python has a really bad situation about dependencies upgrade on the development side. But Docker won't help you there anyway. Personally, at this point I just assume any old Python program won't work anymore.
./deploy.sh pull sync migrate seed restart
pull calls git pull, sync runs pipenv sync, migrate runs django migrate, seed runs django management command seed, restart calls systemctl --user restart
Target simplicity, it’s the ultimate sophistication.
This is an incredibly common occurrence, especially with ML systems which are not designed by people with an engineering mindset.
Your example talks of packaging issues.
K8s is a great platform with many options, but many decision makers have little knowledge (or don’t research) the implication of their choices.
I hate how true this statement is within the industry. To many of these C-level executives base decisions off whatever CEO summit they recently attended.
“Every app to use microservices!111!”
“Hybrid cloud. We are doing it”
“Serverless, let’s start using this”
“We are fully going to the cloud!11!!”
Then when the results come in, the complaints start rolling in:
1) wHy iS aPp SlOwEr (after MSA)
2) gUys, iNfRaStruCtUrE cost is SoArInG (shifting to “cloud”)
3) the ApP is ToO cOmPlEx (after MSA, and “serverless”)
Some of these aging dinosaurs need to be put out to pasteur
AFAIK, the virus of C-suite IT bad ideas doesn't discriminate on the basis of age.
Ideally you would be using immutable events over a message bus as a default pattern.
Just as anti corruption layers have a use case. So do service meshes.
But service meshes they are synchronous communication, with indirection typically through a sidecar, they will have the very real costs of synchronous communication in distributed systems.
IMHO the problem isn't that they are problematic when applied to the correct use case, it is that as a default pattern they inherently will cause problems with little to no benefit for systems that can use events.
It is simply the same as any tight coupling, it should be avoided unless a specific use case justifies the costs.
Cargo culting it in because it is superficially 'easier' will typically result in code that doesn't have the scaffolding in place to support events and is very expensive to pivot away from if you don't intentionally write your code to support looser coupling in the future.
Istio, Envoy, Linkerd, etc... currently cater to the synchronous request/response communication between microservices.
It doesn't matter how decetralized their implementation is, it is still essentially sync operations over a network.
We have known about the costs of that decision for a long time irrespective of the implementation complexity.
It is a 'sometimes' food.
I understand there are various advantages like metrics, etc, but the encrypted traffic between pods and services is the one that I can see many orgs demanding. If not a service mesh what other options are there?
Once again, the bad implications of a seemingly good idea come to bite us.
All of these go away with a service mesh and sidecar model where the everything is encrypted for you and certificates are completely managed by the service mesh. No need to roll your own PKI. Your app only needs to speak plain HTTP and from it's perspective every other app is also just speaking plain HTTP, no TLS in sight. It makes developers' jobs a lot easier, only concentrating on shipping features vs. wrangling with TLS.
Saying all of that, you could just roll your own fine. Depending on the competency of your company.
There's no way I'd trust random developer to implement mTLS properly.
If adding basic mTLS security to every service is that much of a pain, maybe you have too many of those?
It might take years of trying to avoid TLS to learn that all the alternatives are far worse. So—yes, just bite the bullet, it's not that bad once you internalize the model, and you really only need to solve this once per organization.
Reading that comment, I sort of understand why people don't do this because it takes some understanding of HTTPS and cryptography to do this properly.
If this doesn't look like a real thing to you, congratulations, you are in a well run organization that doesn't have this problem.
Anyway, whether the costs of this thing are worth the gain, I have no idea.
As organizations accelerate their moves to the cloud, they are, by necessity, modernizing their applications as well. But shifting from monolithic legacy apps to cloud-native ones can raise challenges for DevOps teams.
Developers must learn to assemble apps using loosely coupled microservices to ensure portability in the cloud.
This just isn't true. You don't need microservices to use the Cloud.Who on earth is giving you this advice? Service-messages are squarely in the “you’ll know when you need it, and it’s not day 1” bucket.
Is this the experience people are having with K8s? Welcome to Kubernetes, here’s a service mesh gl;hf. If this advice is common, no wonder some people think K8s is over complicated.
To anyone not aware: it _does not_ have to be this complicated. Default ingress controller and a normal Deployment will carry you really, really far.
The context of the content is obviously for Kubernetes enterprise platform teams and not deploying your first nginx pod. The author also gives specific examples of why a service mesh is useful and in what scenarios it shines.
The author's entire point was money which is weird as it is peanuts compared to the DevOps salary of people needed to work with this complex setup. Unless you are in a startup and good in tech, in which case service mesh doesn't make a lot of sense. Why do you need encryption or metrics collection above ingress. Nginx does a good enough job for metrics.
Thanks for the unneeded tone. But the solution that author likely uses is linkerd which is free for 50 users. And even then very likely it would be less than $1000/month. If it is too much for the org to pay, I don't think they need it.
Service mesh is not something a small or medium sized company needs.
> those solutions usually need tweaking & knowledge to be operated. And if any issues happen, swift support comes at a hefty premium.
Exactly my point. Base price is not something that is gonna cost them the most when using advanced tools like this. Even if it was free it would costed pretty much the same for any medium sized org.
You can have a guy banking 200k a year in base salary begging his superiors to have a fast laptop for work and compilation, and request is not even rejected, just hangs indefinitely. It would pay for itself within few days with increased efficiency. Or say Java dev fruitlessly begging for intellij idea license that cost less than 1 MD.
Been there, done that (in some way), now I just accept whatever boundaries are given, and let everybody know that due to this my work will take X amount of time. If they don't like it, escalate this so that management structure is aware of unreasonable expectations within given situation. If corp is so rotten that this all doesn't matter and its a long term situation, just leave such toxic place, you may feel scared of uncertainty but later you will thank yourself.
Health is a finite resource and leaking fast, you often don't notice it since warning signals come way too late.
Sorry for that, it was unnecessary indeed.
> But the solution that author likely uses is linkerd which is free for 50 users. And even then very likely it would be less than $1000/month. If it is too much for the org to pay, I don't think they need it.
Didn't check the Linkerd plans in detail but the author says 2 grands per cluster per month. Now, even a medium org (like where I work), can have easily 10 clusters between different environments or different workloads. It's not that weird. And that would make the price already going to 20k a month which needs to be justified. Especially if you were using the same thing but not paying anything before.
This kind of world arises if an org has, say 1000 engineers, or ~100-200 teams, and every team wants to do their own thing with accountability improperly mapped etc.
And infrastructure engineers turn into police with an obvious overemphasis on observability.
I guess a large amount of complexity in infra and wastage is done to hide human inefficiencies.
< Usual caveats apply, don’t bother until you need it, etc etc. Personally though, I’ve found dropping in linkerd, or Cilium CNI about an afternoon’s work, with no application code changes necessary>
Yes I do this for living snd know what I'm talking about.
Regardless of service mesh though, although it's a part of it, the complicated reputation is not undeserved and not solely due to meshes.
Retries for free are good I guess, but not something essential. MTLS is something I don’t want at all.
I mean, if even Google who were behind it in the first place, are saying it's complicated and have 3 levels of managed Kubernetes services, I'd say it's pretty clear to everyone it is indeed complicated. Some of the complexity is simply due to the complex problems it solves, some is footguns, some is layers upon layers of abstractions to glue around design deficiencies.
> To anyone not aware: it _does not_ have to be this complicated. Default ingress controller and a normal Deployment will carry you really, really far
You're absolutely right. However, Kubernetes is rarely chosen after a careful evaluation of requirements; it's more often than not because it's considered necessary or for resume driven development. In that case, service mesh is easy to tack on, regardless of need.
How is it better have service A request to a proxy, which requests to another proxy, which requests to service B? I get the security benefits of that, but the network architecture is boggling. How many PBs of data are sent each day for what could be a monolithic service?
Actually, to the end — companies that embrace microservices, which see the value in them. How do they manage network traffic for these kinds of K8s setups at scale? Surely they’re not using REST and HTTP. Is it as simple as protobufs over HTTP? Quic? Something else?
Edit: lots of great discussion below but I really meant how do they manage network TRAFFIC, not microservices in general :)
I find that microservices and this type of architecture have become a religion - you do it this way because you do it this way. You add another layer of complication because that's what you do now. You add this product because that's what you do now. Now you do it this way. Now you stop doing this thing and do this thing instead. It's all proclamations and a truly insane level of complexity and often a truly stunningly low level of performance achieved from some very powerful hardware because everything is behind at least twenty layers of abstraction and you're like, encrypting traffic which is just being passed between VMs which are on the same hardware, but because you can't guarantee that they're always on the same hardware you have to encrypt and use a proxy and... oh wow
Watching it from the outside is a bit exhausting, it just seems to be so much churn and overhead.
The tradeoff here is to decouple services in order to allow them to be developed somewhat independently of each other. Monoliths remove the overhead of network requests but they present their own challenges. You have a lot of implicit dependencies and feature development becomes complicated with changes having unexpected effects very far from the source.
Ultimately engineering organizations need to decide the model that works best for them. Neither is inherently better, they’re solving different problems.
I have seen, developed, designed, managed, deployed, operated, fondled, and otherwise been around thousands of large systems that are anything between trivial importance to “must always be running, in the national interest”. I was around when SOAs were a hot new thing, and SOAP was being rumoured as the thing that was going to save us from everything. A fondly recall an overpaid Compaq consultant talking about “token passing systems” when they were describing message queues.
I have seen exactly two systems that really had to be designed and built along a microservice architecture. Both of these had requirements that introduces a scale, scope, and complexity you simply don’t see very often. All other microservice architectures didn’t solve for requirements, they solved for organisational inefficiencies, misalignments, mismanagement, and - in no small part - ego.
When discussing this topic, proponents of microservice proponents talk about many of the advantages these architectures bring, and they are often not completely wrong. What is lacking from these discussions is often a sense of perspective. “Is this solving real problems we have?”, “What is the compound lifecycle cost of this approach?”, and, of course, “How much work is involved in displaying the users’ birthday date in the settings page?”[1], and “when will Omega Star get their fucking shit together?!”[2].
Don’t start with microservices as the default. I will typically work out the monolithic approach as a point of departure. Want microservices? I’m open to that, just demonstrate how that will be better.
Be developed, maintained, and operated by the real organization that exists and not an ideal organization which doesn't is, in fact, usually a practical if not a theoretical requirement, and its usually easier to adapt architecture than to adapt organization.
You don't need a service mesh for that, though? Heck, you don't even need an ingress for service to service traffic. ingress-nginx does the job well, without being overly complex, and most importantly to me, logs when something is wrong, which I cannot say the same for Istio which I was fighting earlier this week where it was just happily RST'ing a connection and saying nothing about why it was deciding to do that.
1. I can have a unified set of metrics for all services, regardless of language/platform or how diligent the team is at instrumenting their apps.
2. I can guarantee zero trust with mTLS, without having to rely on application teams dealing with HTTPs or certificates.
3. I can implement automation around canary releases without much lift from dev teams. Other projects leverage these capabilities and do it for you as well.
4. I can get the equivalent of tcpdump for a pod pretty easily which I can use to help app teams debug issues.
5. I can improve app reliability with automatic retries and timeouts.
Probably some other things as well... That said, it can be a big increase in complexity to your system the pains of which aren't always distributed to the folks getting the benefits.
But they're only going to be coarse metrics, like what requests/second. You're still going to be needing application-specific metrics.
> 2. I can guarantee zero trust with mTLS, without having to rely on application teams dealing with HTTPs or certificates.
I do like the idea of this feature of service meshes. It is a slog to get teams to do this responsibly. But, like I said: fighting Istio to understand why it was RST-ing a connection, for no apparent reason. Not logging errors is worse. Perhaps the idea is sound, but the implementation leaves one desiring more.
I should mention the same Istio service mesh above is a SPoF in the cluster it runs in, on account of being a single pod. I can't tell if the people who set it up were clueless, or if that's the default. I suspect probably the latter.
> 3., 4., and 5., as well as actually using mTLS in 2.
TBH, these are just benefits I've never been able to realize. I'm stuck slogging through the swamp of service mesh marketing and the people who want to bring the light of their savior the service mesh but without actually getting their hands dirty doing the work of deploying it.
The fact that you need other metrics does not substract from OPs original point. It's still good and much better overall to handle a set of comprehensive metrics at infra level than to orchestrate every app.
About the coarsness, i think it's not really true. Proxies are freaking powerful and they do a lot of stuff at l7, too much in fact (look at envoy, jesus christ). That's one of the reasons why despite the insane complexity of service meshes, they are paramount for observability.
ok, but what's your threat model for this? great you can tie service versions to each other, but they are just proxies.
2. I always wonder whether that's a timing thing vs. NetworkPolicies and encrypted inter-node traffic. Are the realistic attack scenarios where it's possible to read out intra-cluster traffic but not mess with the cluster, or even read the intra-pod traffic?
3. I've been quite disappointed with how little k8s provides here. I wish it was easier to move traffic off of an old version and only shut the pod down once the last connection was done :/ Maybe I need to look into a service mesh for that?
4. What's the difference to e.g. kubeshark, or just attaching a tcpdump debugcontainer to the pod? Another instance of first to market / potentially nicer ecosystem?
5. I get squeemish with infra-level activities like this. Yes, technically the http method and some headers should make it obvious whether that's save or might break at-least/at-most once or similar semantics. But that requires well behaved applications. While the premise here is infra imposing behaviour to allow applications to be looser around these kind of things.
But yeah I forgot this existed.
I don't really understand the question you're asking, but I think maybe the answer is that network pipes are just bigger than the scale most people are operating at? I don't think anything I've ever done has really had that many qps, and if it has, it is more likely to raise an eyebrow that says "who's spamming requests" more that "I guess we've made it to the big time".
REST & protobufs are orthogonal. Empirically, literally nobody is doing REST, and most things are just ad hoc, poorly to not-at-all defined JSON/HTTP with a few HTTP verbs sprinkled in to make everyone feel good. It could be protobuf, too, if you like, but unless you have some truly gargantuan JSON, it really won't matter in the end. Compression will make up enough of the difference in size on the network. Some languages don't have to allocate the keys a billion times, too, though even in Python, it's a while before it starts to hurt.
What I see more of is processes just inexplicably using gigabytes upon gigabytes of RAM, burning through whole years of CPU time for no particular reason before just dropping back to nominal levels like nothing happened, and dev teams that can't coherently understand the disconnect between just how much power a modern machine has, and what their design doc says their process is supposed to do (hint: something that shouldn't take that many resources).
I like microservices, but there should be strong areas of responsibility to them. For most companies, I think that's ~2–3 services. At my current company, it's ~2 services + a database, with the rest being things like cron jobs or really small services that are just various glue or infra tooling.
Our DevOps team starts off the monthly all hands meeting by leading a ritual during which they ceremoniously sacrifice an animal from the Fish and Wildlife Service's Threatened & Endangered Species list while the rest of the company chants:
Exorcizamus te, omnis immundus spiritus
omnis satanica potestas, omnis incursio
infernalis adversarii, omnis legio,
omnis congregatio et secta diabolica.-- Anthony DeBoer
"SCSI is *not* magic. There are *fundamental* *technical* *reasons* why you have to sacrifice a young goat to your SCSI chain every now and then."
-- John F. Woods
The companies that need service meshes couldn't possibly run everything in a single monolith. They already have several if not dozens of monoliths, each one typically coming into the architecture when a large enterprise acquires another company and its monolith, or simply different business units / product lines that don't talk to each other because of sheer organizational scale.
It's called a service mesh, not a microservice mesh. First you get the benefits wrapping, securing, and monitoring each of your monoliths, then you reap more benefit when you start to break up each of those monoliths so common concerns can be addressed by common services and provide a unified experience to the enterprise's customers/users.
We run 1800 services in our mesh. Most of them rest/http, a select few gRPC and graphql. We even have some soap/http services in it.
Or are the services just cogs in the machine, managed by individual devs / cogs in the machine? I imagine the latter and presume that's the main benefit of so many services, but genuinely curious as I've never experience that many services in an architecture
A well functioning service mesh (or even just a well maintained and discoverable ingress controller) is essentially invisible to the individual dev teams. Just think how modern stacks work from a frontend dev's perspective: team wants to use an additional feature, so they find the budgeted credit card, sign up to a random third-party provider, get their access token, and go. From the codebase standpoint, they merely added a new roundtrip to a random service and process the responses in their code.
From the dev team's point of view, having the same feature available internally, behind "just another URL", makes no big difference. Maybe less politics around vendor spend and, with luck, easier integration with the remote service auth. Almost certainly less wrangling with compliance and legal.
Whether that URL is provided by an ingress with a proper FQDN, or a service mesh entry with otherwise unresolvable name, is (and should be) irrelevant.
Modern distributed systems have long since become too large for any single person to fully comprehend them through and through. There is no Grand Design[tm], they are all results of organic changes and evolution. Service discovery and routing can be architected. Individual services within the system can be architected. The complete system where hundreds or even thousands of services interact can not.
I hesitate to tell you how many layers are between me an HN's servers right now.
How much traffic would be generated if, for every request to HN, six other requests fire? And every time you comment a cascade of requests fire in HNs imaginary K8s cluster?
It’s not that there is anything wrong with this, or that the tradeoff isn’t worth it… it’s just so much data flying back and forth over the wire.
Microservice architecture means there's lot of data flying around, but it keeps local resource utilization predictable.
A mesh doesn't add any extra data back and forth over a wire.
It adds some data being copied from one place in RAM to another place in RAM.
You check in configs into the monorepo and there is tooling to continuously sync the configs with the actual state of the infrastructure.
The advantage of microservices at scale is that team X breaking the build doesn’t affect team Y. This scale is probably not until you have 1000+ engineers however.
Plus, applications built on an external database (e.g., postgres not sqlite) will already have 1 hop. Going from 1 to 2 hops is less dramatic than going from no hops to 1. And I guess implicitly any service sending lots of data to the client already has 1 hop.
Imo the only case where a network proxy would be egregious is an in memory database kind of workload where the response size is small relative to the size of the data accessed (maybe something like a custom analytics engine), but that's pretty niche.
Complexity is very, very expensive.
Most organizations will run their services spread across multiple data centers / AZs.
On larger projects:
a) It is sometimes impossible to run the entire application on a laptop. And so having a micro-service means you can quickly iterate and test before embedding it into the wider system.
b) You will commonly run into conflicting transitive dependencies which you simply can't work around. Classic example being Spark on the JVM which brings in years old Hadoop libraries.
c) The inter-relationships between component can become so complex that you really want to be able to use canary or green/blue deployment techniques to reduce risk.
Good architects will know how and when to do what E.g. start with a modulith instead of a monolith, to ease refactoring into micro-services once user count goes through the roof and vertical scaling won't do it any more.
In my thinking, we might as well get used to the levels of indirection and complexity that AIs will be comfortable with. I suspect it will be more akin to what biological computation is "comfortable" with, and less like what our minds happen to prefer
But I digress :) yes, microservices strike me as wiiild
Come on, I just want to deploy my app somewhere. I don't want to know about yamls and kubernetes, I just want a button where I deploy my branch.
It often feels like people working on infrastructure and platform development do it for it's own sake, forgetting that they were supposed to be enablers for other teams building on top of it.
I find it annoying that there are a lot of young (and not-so-young) developers who don't take the time to properly understand where their apps are running in order to make the best of it and also not to write crappy and slow software.
I see people climb mountains in Java when they had all they needed in the database or in the OS.
Either way, you have to understand the solution at a high level, and I think much of the balking at infrastructure is really just balking at needing to understand the fundamental nature of their own problem.
I think there are at least two reasons for this: the first is that infrastructure engineers (e.g., SREs) just think about reliability and architecture a lot more than SWEs and the second is that the infra eng position generally selects for people who have a penchant for learning new things (virtually every SRE was an SWE who raised their hand to dive into a complex and rapidly evolving infrastructure space).
Also, the best SWEs have been the ones who accepted the reality of infrastructure and learned about it so they could leverage its capabilities in their own systems. And because our org allows infra to leak to SWEs, the corollary is that SWEs are empowered to a high degree to leverage infrastructure features to improve their systems.
No it's not. This was simply just not the case before k8s. Running a stateless app behind a (cloud) LB talking to a (cloud) DB has never been the hard part, and it's even easier today when something like Go is essentially just starting a single binary, or using docker for other languages. People seem to have forgotten how far so few components gets you. But incentives for too many people involved aligns towards increasing complexity.
I think most would be surprised how big chunk of modern apps are within that space without active intervention by stuff like microservices. No need to stop at a handful of VMs though, although I imagine that most companies could easily be covered by a few chunky VMs today.
And yes, if you're a PaaS, multi-tenancy something something, then sure, that sounds more like a suitable target audience of a generic platform factory.
And if that's the territory you're in, you need probably need to set up log aggregation, monitoring, certificate management, secrets management, process management, disk backups, network firewall rules, load balancing, dns, reverse proxy, and probably a dozen other things, all of which are either readily available in popular Kubernetes distributions or else added by applying a manifest or installing a helm chart.
I don't doubt that there are a lot of systems that are running monoliths on a handful of machines, but I doubt they have many development teams which are decoupled from their ops team such that the former can deploy frequently (read: daily or semiweekly) and if they are, I'm guessing it's because their ops team built something of comparable complexity to k8s.
One downside I also see compared to the "old-school" approach, albeit maybe an indirect one, is that it's also a very leaky abstraction that makes the environment setup phase stick around seemingly in perpetuity for everyone rather than being encapsulated away by a subset of people with that expertise. No normal backend/frontend dev needed to know what particular linux distro or whatever the VMs were running or similar infra details, just focus on code, the env was set up months ago and is none of your concern now (and I know there's some devops idea that devs should be aware of this, but in practice it usually just results in placeholder stuff until actual full-system load testing can be done anyway). So a dev team working on a particular module of a monolith should be just as decoupled as with microservices. Finally, for stateless app servers, the maintenance required was much rarer than people seem to believe today.
I realize that a lot of this is still subjective and includes trade-offs but I really think the myth building that things was maintenance ridden and fragile earlier is far overblown.
[1] https://cloud.google.com/compute/docs/containers/deploying-c...
I suppose we all live at the abstraction level that suits us. I'm of the opinion that a lot of the modern infrastructure we work with is over-abstracted. I'm not sure that most developers even know that an HTTP server exists thanks to the work of modern devops where you push a commit and it gets automatically deployed to some sort of container orchestration system. At some level it has to be this way; we stand on the work that was done before us; "the shoulders of giants" it's often said with more poetic flair. So I wonder if it's just me getting older and not understanding the new, or if we have genuinely gone astray in some ways. And those need to be mutually exclusive either.
I guess I've traditionally built and deployed all my apps myself, without the luxury of an infra person, so I had to learn all this, including Linux administration.
It's a wide-ranging skillset no doubt, but still built on the work that precedes it. Did you build the servers that ran your code? Rack them and set up the networking? Build redundant power systems for them? HVAC? I'm sure you've done some of that, but no one person can really do it all. I think the saying is something like, "fish don't know they're wet"; meaning that they're so adapted to the ocean that they don't really have any idea that there are other ways to live. I get that feeling when I hear older folks talk about working on mainframes - it just sounds like an alien world. I came of age in the era of distributed systems - Microsoft vs Open Source, but it was always on commodity x86 hardware and TCP networks. There's no law of the universe that said it had to evolve that way. It seems like the new way is "cloud-native" which I'm not sure I understand or like (as I've said, it may just be me getting old); my biggest problem is that it seems to give ever more power to Amazon and other members of the tech oligolopoly rather than any technical shortcoming.
Maybe it's the same as before, maybe I'm just less familiar and it seems huge to me.
I am afraid people is focused on company products that the will forget what infrastructure is.
If that is not possible, you have failed your job as masters of the infrastructure.
At my company, for this I need to copy-paste some terraform in repo 1, write some scaling config in repo 2 and put in new, encrypted credentials and other configs in repo 3, then finally I can deploy my docker app from repo 4.
I can all do it myself in two hours but it constitutes total failure on the infrastructure teams side.
I'm getting the impression that the whole industry shifted focus from solving actual problems in the most simple way to draining money and keeping things afloat as much as possible.
"Leaking" infrastructure details to developers is actually less work for them because they are deferring the finer details of the app deployment to the developers who made it. Your request for a magic button that doesn't require you to understand anything about infrastructure is exactly why we have so much complexity.
I think there's a broader point though, in that very often, infra teams will pursue solutions that solve problems in the perspective of their own lens and interests without good oversight from the broader organisation, for developer experience, and economies of scale.
I've seen and worked in environments at both ends of that scale and the gap in dev-ex, and agility as a result can be absolutely staggering.
So often its to avoid 'vendor lock-in', only for infra teams to become the 'vendor', its complex, so they grow by necessity to be expensive, but then still lack the resources to be able provide a clean experience that can be easily migrated off of, resulting in lock-in.
As a dev, the cost isn't my concern, but whats frustrating is knowing that its possible to deploy a new service, with all the bells and whistles, in an hour, and being unable to.
My most recent example: deploy a lambda function. Building the image: 1 day. Deploying it to prod: 2 weeks.
If you agree with the above statement, please go into management.
You shouldn't have to know about the yamls and the k8s, but you should know about infrastructure concerns and how they relate to you and your design choices, just as infra guys should be aware of what's going on on the front side. Having to work with close-minded counterparts who refuse to elaborate further than "make it work now, or else" is about the worse thing there is. Being willing to participate in good faith in back and forth discussions around those subjects is important and pays dividends.
It’s also worth noting that sometimes the details leak because app teams want them to leak—they want to build some feature that requires some particular infrastructure solution (e.g., we want our app to run some background task in response to requests, and we don’t want to do it in the VM lest the autoscaler—which is background-task-agnostic—kill it while the background task is running).