The Service Mesh: What Engineers Need to Know
servicemesh.io
servicemesh.io
Most immediately to API gateways (eg. Apigee, Kong, Mulesoft), which provide similar value to SM (in providing centralized control and auditing of an organization's East-West service traffic) but implemented differently. This is why Kong, Apigee, nginx etc. are all shipping service mesh implementations now before their market gets snatched away from them.
Secondly to cloud providers, who hate seeing their customers deploy vendor-agnostic middleware rather than use their proprietary APIs. None of them want to get "Kubernetted" again. Hence Amazon's investment in the very Istio-like "AppMesh" and Microsoft (who already had "Service Fabric") attempt to do an end run around Istio with the "Service Mesh Interface" spec. Both are part of a strategy to ensure if you are running a service mesh the cloud provider doesn't cede control.
Then there's a slew of monitoring vendors who aren't sure if SM is a threat (by providing a bunch of metrics "for free" out of the box) or an opportunity to expand the footprint of their own tools by hooking into SM rather than require folks to deploy their agents everywhere.
Finally there's the multi-billion dollar Software Defined Networking market - who are seeing a lot of their long term growth and value being threatened by these open source projects that are solving at Layer 7 (and with much more application context) what they had been solving for at Layer 3-4. VMWare NSX already have a SM implementation (NSX-SM) that is built on Istio and while I have no idea what Nutanix et al are doing I wouldn't be surprised if they launched something soon.
It will be interesting to see where it all nets out. If Google pulls off the same trick that they did with Kubernetes and creates a genuinely independent project with clean integration points for a wide range of vendors then it could become the open-source Switzerland we need. On the other hand it could just as easily become a vendor-driven tire fire. In a year or so we'll know.
Another intresting note is that Google did NOT recede control over Istio to CNCF.
I'd argue this is backwards. Envoy has a fairly tightly defined boundary with relatively strong guarantees of consistency given by hardware -- each instance is running on a single machine, or in a single pod, with a focus on that machine or pod.
The control plane is dealing with the nightmare of good ol' fashioned distributed consistency, with a dollop of "update the kernel's routing tables quickly but not too quickly" to go with it. It's "simple" insofar as you don't need to be good at lower-level memory efficiency and knowing shortcuts that particular CPUs give you. But that's detail complexity. The control plane faces dynamic complexity.
Kubernetes hasn’t changed the landscape as far as cloud providers market share. AWS is still the leader, Azure is still big in MS shops and GCP is an also ran.
And I think it is more akin to have an aircraft carrier and the protection fleet of smaller ships..
You have your big api gateway with management, billing, governance (for B2b, b2c, b2b2c) and your service mesh.
Istio is that project, but they'd rather it was Luxembourg than Switzerland.
What amuses me about this is back in the day everyone thought the Mach guys were crazy for thinking things like network routing and IPC services be implemented in user space... and others mocked the OSI model's 7 layers as overly complex (e.g. RFC3439's "layering considered harmful").
Now we've moved all our network services onto a layer 7 protocol (HTTP), and we've discovered we need to reinvent layers we skipped over on top of it. We're doing it all in user space with comparatively new and untested application logic, somehow forgetting that this can be done far more efficiently and scalably with established and far more sophisticated networking tools... if only we'd give up on this silly notion that everything must go over HTTP.
Link for others:
Now I have a name for that at least :)
We also want our retries to work at the application layer too so lower layers would at least have to understand chunks of a connection larger than a packet can fail even if all the packets were acked successfully... _and_ a completely different machine should receive that chunk instead.
We want all that for free and we don't want to pay for hardware. Http/2 is the only close to this and you still need app layer retries.
We initially choose istio because it seemed to satisfy all our requirements and more - mTLS, sidecar authz, etc - but configuring it turned out to be a huge pain. Things like crafting a non-superadmin pod security policy for it, trying to upgrade versions via helm, and trying to debug authz policies took up a non-trivial amount of time. In the end, we got everything working but I probably wouldn't recommend it again.
It's funny that I was at kubecon last week and there was a start up whose value prop was hassle-free istio and the linkerd people stressed that they were less complex than istio.
Istio may appear more complex but that's because it has a superior abstraction model and supports greater flexibility. We're beginning to migrate from Linkerd to Istio at this point. I had the same initial frustrations with podsecuritypolicy (and linkerd suffers from the same), but istio-cni solves the superuser problem, and I believe even the istio control plane is now much more locked down in the latest release.
However if I had my way I would be telling every team they don't need service mesh. We don't have any particular service large and complex enough to really take advantage of its sold features.
Would you mind elaborating on what those Linkerd issue are/were that were effecting reliability and troubleshooting?
The 10:1 ratio of microservices to developer sounds like hell though that's just too much to reason about.
Plus, it’s not either/or situation. A fat clients for Go + Node.js; and a proxy for all others. This way your core logic can enjoy increased introspection / more speed / higher reliability; while special purpose services get a proxy which allow interoperability.
Service meshes are proxies deployed alongside all the services, and used as much by devops (and potentially security) as by the application developers.
It seems like a very thin API Gateway that forwards calls directly to the microservice API without enforcing much would be much easier to manage and have all the same benefits.
Whats the purpose of deploying the reverse proxy alongside the microservice?
The idea behind a service mesh is that no matter how you design and implement a microservice architecture, you still have all these little services talking to each other somehow. By "default", they're communicating using native sockets and HTTP or gRPC.
There's lots that can go wrong any time you have a bunch of things talking to each other on a network (see Tanenbaum's "Critique of RPCs" for a classic explanation).
If everything was written in the same language, what you'd probably do is come up with a common networking/RPC library for all your services to use. That library would give you a common interface, so everything used the network the same way, and some measure of observability (maybe just by doing a standard log format), and maybe some security controls like a standard HTTPS connection and a standard token format.
But if you've got multiple languages, that option isn't very attractive anymore, because even if it's feasible to build that library once for each platform, now you're keeping an additional component (the library) in sync between the platforms, and that's a drag.
The insight behind a service mesh is that containers make it cheap and easy to just stack a tiny out-of-process proxy alongside your services. So you can just move all the logic you would have had in that common service library into the proxy. The proxy will give you the same features no matter whether your service is a Rust binary, Clojure running on the JVM, or a shell script.
And, because it's an out-of-process standalone component, it can be built and maintained by a third party, which means the features it provides can be a lot more ambitious than you'd build in your own library. So now instead of just hoping to get some TLS and maybe a standard token, everything can be mTLS with client certificates, a standard dashboard for managing certificates, rule systems for what client certs will allow you to talk to which systems, graphical maps of who's communicating with who and trace collection for specific pairs of services, etc.
This is all very un- like what an API gateway does, and the whole concept sort of revolves around having little reverse proxies running alongside the services.
If it were in-kernel and supported scatter-gather and zero-copy, that might be different (though some people even avoid going through the kernel).
I still think describing the concept more plainly would help a lot of the confusion.
A more general way to think about service meshes is that the network layer we code to right now is actually really primitive; its service model was fixated in the 1980s, and its programming interface hasn't much evolved from the early 1990s. We'd be happier if we could level up the whole network, so that it had QoS controls, a really expressive security model that didn't rely on magic-number ports and address ranges, and observation capabilities that communicated application-layer details and didn't just try to approximate them the way flow logs do. You can get all that stuff, internally at least, by putting all your services on the same service mesh.
Another thing to look at is Slack's Nebula, which was just released last week:
https://slack.engineering/introducing-nebula-the-open-source...
Nebula is a service mesh that runs at the IP layer (where Istio and Linkerd ride on top of HTTPS proxies, Nebula rides on top of a somewhat Wireguard-ish VPN). Slack has been using it internally for 2 years now. It's solving the same problems Linkerd is, but with a radically different implementation. You can get your laptop connected to a Nebula service mesh in ways that would be clunky to do with a Linkerd mesh.
I'm not sure what you mean by "instances had a local HAProxy running" but if you're thinking about a bunch of reverse proxies handling incoming requests, be aware that service meshes are handling both inbound _and_ outbound traffic to/from your service. For example, you might use a reverse proxy in front of your instance to terminate TLS, but you cannot implement something like two-way TLS authentication between your services unless you're putting something in at the client side as well.
I had trouble finding this under the given title. For those searching, I believe tptacek is referring to Tanenbaum & Renesse, A Critique of the Remote Procedure Call Paradigm [0], published in 1987 or 1988.
A fun read, especially since "it's just the same as local" is a myth each generation gets to revisit. Compare to Waldo, Wyant, Wollrath & Kendall's A Note on Distributed Computing[1], published in 1994.
[0] https://www.win.tue.nl/~johanl/educ/2II45/2010/Lit/Tanenbaum...
[1] https://www.cc.gatech.edu/classes/AY2010/cs4210_fall/papers/...
A better way to describe is "smart pipes, dumb programs". Imagine that all your circuit-breaking/retry/etc robustness logic was moved into another process that happened to be running right next to the program actually doing the work.
You can have both an API gateway and a service mesh deployment -- for example Kong's Service Mesh[2] works this way. They're saying stuff like "inject gateway functionality in your service", but that only make sense if you sent literally every request (whether intra-service or to/from the outside world) through the gateway. Maybe that's how some people used Kong but I don't think everyone thought of API gateways as a place to send every single request through. You'll have a Kong API gateway at the edge and the kong proxies (little programs that you send all your requests through) next to every compute workload.
[0]: https://linkerd.io/
[1]: https://istio.io/
Imagine that for every application there is one small binary that runs and serves all it's traffic, like a chauffeur. Your application stops talking to the outside world completely and sends all messages to the small chauffeur binary -- which then talks to other chauffeurs, over the network.
Keeping with the chauffeur analogy, there is a "head office" which calls the chauffeurs on CB radio at regular intervals that lets them know which cars go where and how to start them/etc.
"head office" => "control plane"
"chauffeur" => "side-car proxy"/"data plane"
In the end what this means for your application is that you just make calls to external services (whether your own or others) and since all your communication goes through this other binary, you get monitoring, traffic shaping, enhanced security, and robustness for free.
Another interesting feature is that if the side-car proxy can actually understand your traffic, it can do even more advanced things. For example you can prevent `DELETE`s from being sent to Postgres instances at the network level.
Once you have this underlying grid of services all connecting via a service mesh - in a reliable, secure and observable way - for some of these services you want to define governance and on-boarding rules, or you want to expose them externally as products that developer can consume and you then use an API GW to do those things.
In this regards, an API GW is just a service among the services you are running inside the service mesh.
This is really the future of distributed parallel computing, but we're still just bolting it on rather than baking it in.
The appeal of App Mesh for us was initially around using it to facilitate canary deployments. AWS Code Deploy does a nice job with Blue / Green deployments and that may suffice for us, but it doesn't support canary for Fargate. Is that enough reason to add the additional complexity in our stack? Not sure, looking for input.
Also, much of the documentation is focused on K8s. I'm murky on how to implement an internal namespace for routing. Most of what I've seen is like myenv.myservice.svc.cluster.local but its not clear to me that using that pattern is needed in the context of Fargate.
Consistent observability is valuable, but again Fargate can do that pretty well- it just doesn't mandate access logging so that would be left to the app itself.
We want to implement OIDC on the edge for some services, but App Mesh doesn't support that yet as other meshes like Ambassador, Gloo, and Istio seem to. Since App Mesh doesn't really act as a front-proxy on AWS, we'll still be using ALB to handle auth which is fine, I think. I get mixed messages about the need for JWT validation, but if so, that would need to be implemented in the app level with ALB fronting it.
Can anybody help me find resources to sort this out? I've been through the `color-teller` example time and time again, but it still leaves lots of open questions about how to structure a larger project and handle deployments effectively.
Maybe you should write a script for this? It sounds like you're about to take on a lot of complexity for just the ability to do canary deployments when you could probably hack up a script in a day or two.
> We want to implement OIDC on the edge for some services, but App Mesh doesn't support that yet as other meshes like Ambassador, Gloo, and Istio seem to. Since App Mesh doesn't really act as a front-proxy on AWS, we'll still be using ALB to handle auth which is fine, I think. I get mixed messages about the need for JWT validation, but if so, that would need to be implemented in the app level with ALB fronting it.
JWTs are only required for client-side identity tokens (you can use opaque ids and other kinds of stuff for backends) -- it seems like you're also at the same time looking for something to take authentication off your hands? App Mesh doesn't do that AFAIK, it's only the service<->service communication that it's trying to solve.
I think it might be a good idea to make a concise need of what you're trying to accomplish here, it seems kind of over the place. From what I can tell it's:
- Ability to do Canary deployments
- The ability to shape traffic to services (?)
- Observability, with access logging
- AuthN via OIDC at the edge
A lot of meshes do the above list of things, but the question of whether it's worth adopting one just to get the pieces you don't have already (which is only #2 really, assuming you scripted up #1), is a harder question.
> Sure, it only worked for JVM languages, and it had a programming model that you had to build your whole app around, but the operational features it provided were almost exactly those of the service mesh.
The thing is, all of our microservices communicate with each other using Kafka. Envoy has an issue open for Kafka protocol support [1], but it's a fundamentally difficult issue because adopting Kafka forces you to build out "fat client" code and building a network intercept that can work with pre-existing Kafka client code is non-trivial. On observability, Kafka produces its own metrics.
Granted, Kafka doesn't offer the same level of control. But Kafka does offer incredible request durability guarantees. We don't have "outages" - we have increased processing latency, and Istio/Envoy and other service meshes can't offer that because they do not replicate and persist network requests to disk.
[1]: https://medium.com/airbnb-engineering/smartstack-service-dis...
But if you're going to do the fine-grained microservice thing, the service mesh concept makes sense. You might choose not to use it, the same way I choose not to use gRPC, but like, it's clear why people like it.
* Codebase should be defined as 'the platform'. where one team will most likely never look at the code of other team's microservices. * this communication problems and overhead start the moment you go from 2 to 3 or more teams. * the term 'team' in this context should be interpreted very broadly. One dev working alone on a microservice should be considered "a team".
Also, things mentioned in the article: you don't want to implement TLS, circuit breakers, retries, ... in every single microservice. Keep them as simple as possible. Adding stuff like that creates bloat very quickly.
Unfortunately these technologies are at peak hype so everyone seems to be implementing them for their small to medium crud apps. But get very sensitive if you try and point it out.
One meta call out on the writing - I read and scrolled at least 30% through the page on my iPhone until the author explained why I should care about a service mesh I.e. what problems it tries to simplify or solve.
It seems to me there are some strong use cases here, but it’s only worth your while if you’re operating at sufficient scale.
For instance, if my team at some FAANG scale company is responsible for vending the library that provides TLS or log rotation or <insert cross cutting/common use case here>, and it requires some non trivial on boarding and operational cost, migrating to this kind of architecture longer term where these concerns are handled out of the box may be beneficial.
Still - it doesn’t mean the service owners are off the hook. They still need to tune their retry logic, or confirm the proxy is configured to call the correct endpoints (let say my service is a client of another service B and for us, B has a dedicated fleet because of our traffic patterns). This is an abstraction. Abstractions have cost.
Trust but verify.
The trap people fall into is, “Here’s a new technology or concept. Let’s all flock to it without considering the costs.”