The Ambassador Pattern
learn.microsoft.com
learn.microsoft.com
Hmm, that request didn’t make it through! I wonder if it was dropped by the database, the business logic, the spring http-handling framework, the retry annotation, nginx reverse proxy, the authentication sidecar, the “ambassador service”, the SQS queue, the internal VPN, or cloudflare?
Just writing it out makes my brain hurt.
But otherwise it's not required.
Or maybe scale doesn't matter. Even if you replace thousands with tens you still often have a huge variety in skillset and focus.
this is a business thing. you need this when you're hiring a contractor to add some functionality, and you don't want them to touch the service behind the ambassador. probably because you've hired a different contractor to build the original service, nobody at your company knows how it works anymore, and making the new contractor figure it out is out of budget.
sidecars bad
What does this have to do with microservices?
I'm trying to debug my monolithic service:
Was it the database? My own business logic? The HTTP framework? The reverse proxy we run to terminate TLS? Was it an authentication or authorization issue? The network appliance in front of all of this? Did the queue fail to persist it at a layer above this? Was it the hosting provider?
You outlined any distributed service.
If the call is done in a synchronous context (for example original client executes a REST call), you are adding latency by inserting an extra service in between. When the remote service is not immediately available or responsive or fails, and you do a retry, in the end your call might take too long and the caller might cancel the request. I have seen this behavior in practice and it makes me wonder how useful this implementation is in a synchronous context. You add complexity (on the infra level) and you add latency in the happy path. When a retry is needed, most of the time (in my experience) the call to the remote service does not succeed in time, and the original call still fails.
If the call is done in an asynchronous context (for example the original client picks up a message from a queue and processes it), you are adding latency and complexity by inserting an extra service and extra logic. However, when the remote service is not immediately available, you can just let the processing of the message fail. The queue or bus should contain retry logic that could be finetuned based on the type of error you get. So, in that case, you should already have a retry mechanism out of the box, then is the added latency and complexity really worth it?
I understand about the circuit breaker, and I understand that might be useful, because it could prevent the remote service being overloaded with requests (well, at least if every caller to the remote service implements circuit breaker... but then, the ambassador service would better be placed on the side of the remote service and every client should be forced to pass through it, and then it might just be implemented inside the remote service instead of in a separate service adding extra latency/complexity/... basically, the remote service should protect itself against this).
Thoughts? Does it make sense what I am thinking?
Consider 10-15 applications running on a host, and all of them are listening to data being distributed by another service. Instead of all of them opening a connection to that service, instead they would all be connected to this sidecar, and the sidecar would merge the distribution of data (and subscriptions) to the pubsub system
Use this pattern when you:
* Need to build a common set of client connectivity features for multiple languages or frameworks.
* Need to offload cross-cutting client connectivity concerns to infrastructure developers or other more specialized teams.
* Need to support cloud or cluster connectivity requirements in a legacy application or an application that is difficult to modify.
This pattern may not be suitable:
* When network request latency is critical.
* When client connectivity features are consumed by a single language.
* When connectivity features cannot be generalized and require deeper integration with the client application.
If you want to question the usefulness of the pattern, it is best to argue against the use-cases in which it is recommended to be used instead of using a scenario the pattern _explicitly_ states it is not meant for (i.e.: low latency).Now you also had some thoughts on complexity. Regarding what you said: (retries, resiliency, extra logic) it will have to live somewhere. You don't add meaningless complexity for the sake of it after all. How are you going to add another retry policy to a blob you have no control over otherwise? Shifting the burden of networking is also an explicit option listed in the "suitable for" section. You can decide where complexity lives after all. If accidental complexity is your issue I am inclined to ask where you see it here in general. Both the ambassador pattern and your proposed alternatives add overhead in terms of complexity somewhere. I'm struggling to see a clear favorite here.
Explicitly regarding your last statement that "[...] the remote service should protect itself against this).": Retries and Monitoring are things the remote service can't do by itself by definition. Even load balancing/shedding might not be solvable by it depending on the situation. Notice that circuit breaking is not the only thing the ambassador is used for. Network related configuration updates (see section "Context and problem") are something that might not be done by the remote service either.
> You don't add meaningless complexity for the sake of it after all. > You can decide where complexity lives after all.
Good points and something I should think about when designing systems.
You have good and interesting points, and it is true that I am very wary about introducing extra latency in the context of an http api that is used by a webapp. With the infra that is available today, it is possible to build snappy webapps, but my feeling is that you have to be wary about introducing extra "hops" in the execution of one http call, even though that is not strictly a "low latency" requirement.
But it should obviously run in the same container/pod/etc.
I always thought this was a particularly interesting approach from Microsoft where they use this pattern to essentially take the complexity of micro services and instead try and keep it as simple as a normal .NET application but (and I think this is the clever part) in both a vendor and language neutral way.
But all of a sudden it means you can start removing all kinds of cruft and random SDKs from your codebase and push almost all of your interactions with the outside world into something like this .
The problem with having Dapr being a separate thing that my code has to communicate with over the network is that this basically triples the number of hops in an environment that already has too many network hops.
As a former game developer used to replacing divisions by multiplications with a reciprocal to save precious clocks, the design of Dapr makes me... itchy.
Also, I would argue that latency almost always matters.
Those microseconds add up faster than you think. You have to factor in things like TCP slow start, congestion issues, buffering, the extra CPU spent encoding/decoding, etc…
Not to mention that it rapidly becomes impossible or impractical to do things like process requests in a streaming fashion, so for large requests you can’t overlap the stages of a pipeline. You can easily bloat out a “1 second task” into minutes without ever knowing where you went wrong.
“We’re following best practices!” you say.
“Oh, is that why I’ve been waiting ten seconds for the little spinner while the site is processing a 1 KB transaction using 5 GHz CPUs?”
Yes, but you wouldn't use dapr for game dev.
> Also, I would argue that latency almost always matters.
Strongly disagree there. For every interesting business out there, there's 99 boring ones for whom sub-second added latency absolutely does not matter (not that dapr would add that much latency).
> This pattern can be useful for offloading common client connectivity tasks such as monitoring, logging, routing, security (such as TLS), and resiliency patterns in a language agnostic way. It is often used with legacy applications, or other applications that are difficult to modify, in order to extend their networking capabilities. It can also enable a specialized team to implement those features.
Not surprised this is a Microsoft page, given their legacy of long lifetime support for their software products.
It’s not for microservices, but rather for software maintenance of systems that other vendors would consider past EOL.
(You could also argue that obfuscated monolithic programs are be easier to reverse engineer, breakpoint, replay, emulate, time-travel-debug, trace, etc because you can completely control them in your test bench and aren't then working against a hostile distributed system)
20 something years ago "mobile objects" and "agent systems" were a thing and one of the practical incarnations was Jini: https://jan.newmarch.name/java/jini/tutorial/Jini.html
I understand that at different points, you'll frequently run into situations where you need one or several of those, but the way this sentence is phrased, it sounds more like chastizing people for not following the latest fad (the components of which they are conveniently selling too).
"what, your deployment doesn't have advaned metering and circuitbreaking abilities and I have to restart the service to make config changes? get back to me when you have a real deployment..."
And yes, this library usually also covers things like retries and handling errors and outages.
Maintaining an entire service to achieve the same seems like really expensive way to achieve this...
Having such "design patterns" seem like a side effect of our current tools' limitations. Maybe new technologies like Wing[0] (no affiliation) will help pushing us to the next generation of tools.
Actually it was an entire new paradigm.
Procedural programming!
i like this observation - it's not uncommon for patterns to turn into language primitives when they end up becoming a very common thing in your language.
this is reflected both in wing as a language but also in the fact that wing's preflight phase in a sense is a programming model for turning infrastructure patterns into composable components.