Linkerd 1.0
blog.buoyant.io
blog.buoyant.io
As soon as you're not allowed to rely on the service you talk to saying on the a static set of machines (and even on a static set of ports) over time, using standard HTTP libraries becomes really awkward. (Did this request fail because an instance shut down and I should talk to another one? How do I convince my DNS resolving library to ignore its cache and re-lookup the service? Ok I found 3 instances, which one should I talk to? How do we prevent all the other copies of me from spamming the same instance? Gee, it'd be really nice to have something push a notification to me when the services change!)
Not-so coincidentally, this constraint is really common in "cloud native" setups like Kubernetes and Mesos, where service instances come and go as part of normal operation. This is because, in this world you don't update your code by copying new files to some servers and SIGHUP'ing some daemon, you update your code by deploying new containers, waiting for them to be healthy, and tearing down the old ones. This means "failure" is a normal part of your normal software lifecycle.
Linkerd is one of many attempts to make "normal software" behave well in this world, without relying on too much intelligence in the HTTP client libraries (or other non-HTTP stuff, for that matter.)
In a perfect world, everything would respect and obey DNS SRV records. Oh if only most client software knew what they were, and how to handle them! Right there you can see which backends are available for a service, what ports they're on, what their priorities are, and how long that information should be considered valid. But alas, nothing really supports SRV records. So we need something like this until client software becomes more intelligent (if ever.)
Ex: foo.some-delim.baz.bam.example.com would resolve to baz.bam.example.com regardless of the foo prefix. The client could then iterate or randomly generate values for foo to ensure fresh DNS lookups.
One or more decades ago, these problems were proposed (and solved?) by IETF working groups. I could believe that linkerd offers some new features not considered by those RFCs. But it would be nice if it were extensions of existing protocols instead.
Having a transparent proxy could offer a level of convenient integration rarely found among disparate internet protocols. Then again, there might be really good reasons why collapsing too many features into one layer is a bad idea.
Corba had service location, and I assume many other predecessors did as well.
Doesn't bother me though, stuff gets reinvented and rediscovered all the time.
The more recent approaches, like linkerd tend to also include ways to easily wrap existing services without changing them, which is a plus.
Service Location Protocol: RFC 2608
RSerPool: RFCs 5351-6
It's a really interesting pub/sub style IPC. But really it's like a swiss-army knife. CORBA lovers (and haters) will recall IDL, it's used as the description for the interfaces used. IMO the commercial implementations are more feature complete than the open source ones.
If you looked at CORBA in the 1990s and have a sour taste in your mouth, don't hold it against DDS. It's a really great approach. It describes interfaces with a dynamically negiotiated Quality-of-Service among participants. The set of reactions to QoS policy violations effectively create hooks for all of your interesting cases.
IMO DDS works best when all of your participants share a broadcast domain, but it's not necessary to set it up that way.
Which never seemed to find wide, mature use.
It sounds like linkerd is closer to 'onhandig' - does it really imply an evil touch? (sinister - there we go with the left hand thing again)
1) Fraud 2) Adult Person 3) Most common guy 4) Convenient guy 5) Linkmichel 6) Nice handy guy 7) Beautiful guy 8) Smooth 9) Smarter 10) Slapped person 11) False nature
I prefer my own translation
Just musing out loud.
What is a cloud native application?
https://blog.buoyant.io/2017/04/25/whats-a-service-mesh-and-...
So, if you're elastic, you now need some dynamic way of finding the various services...they are spinning up and shutting down all the time, so you can't have a static list of IP's and ports.
So, some framework that maps questions like "what host and port should I use to connect to the pricing service" is desirable. There can be many instances of that service for scale, and when new ones spin up, they are added to the directory. When existing ones fail, or are purposefully shutdown, they are removed.
This is more complex than a normal load balancer, as it has to handle service-to-service connections versus just client connections. And the number of instances of each service is highly dynamic, based on load.
By using linkerd, you manage and configure various things like load-balancing strategies, retry-windows, circuit breaking etc etc. Instead of doing these things with a library inside the application like hysterix, you move it to it's own layer. This also has the benefit of being code agnostic, so you can leverage the same service concepts for any type of app. You can also run linkerd outside of kubernetes.
Linkerd applies dynamic routing rules to determine which service the requester intended. Should the request be routed to a service in production or in staging? To a service in a local datacenter or one in the cloud? To the most recent version of a service that’s being tested or to an older one that’s been vetted in production? All of these routing rules are dynamically configurable, and can be applied both globally and for arbitrary slices of traffic. Having found the correct destination, Linkerd retrieves the corresponding pool of instances from the relevant service discovery endpoint, of which there may be several. If this information diverges from what Linkerd has observed in practice, Linkerd makes a decision about which source of information to trust.
Linkerd chooses the instance most likely to return a fast response based on a variety of factors, including its observed latency for recent requests. Linkerd attempts to send the request to the instance, recording the latency and response type of the result.
If the instance is down, unresponsive, or fails to process the request, Linkerd retries the request on another instance (but only if it knows the request is idempotent). If an instance is consistently returning errors, Linkerd evicts it from the load balancing pool, to be periodically retried later (for example, an instance may be undergoing a transient failure).
If the deadline for the request has elapsed, Linkerd proactively fails the request rather than adding load with further retries.
Linkerd captures every aspect of the above behavior in the form of metrics and distributed tracing, which are emitted to a centralized metrics system.
Let's say you have authenticated a request and need to submit it to the search service. Instead of opening a socket to a particular IP address, you use the service mesh to send the request to 'service=search' on your behalf. The service mesh keeps an up-to-date list of instances that implement the search service in the background, picks one of them to send the request to and gives you the response.
In addition to managing the pool of servers, it can do a few fancy things:
* have connections primed and ready to go before you had to make the request
* balance the load between instances, potentially based on their latencies and error rates
* handle timeouts and retries against different instances
* fail requests without making the remote call if the error rate of the downstream service is too high (circuit breaking)
* expose metrics about request volume, error rates and latencies
it seems unlikely that you make your software faster or more reliable by adding more hops
We could test out Kong, Consul, Fabio, and forget the event-driven callbacks; We could use Route 53 and ELBs; or we could just use Linkerd to drive this and get some pooling features and monitoring features as well.
I think of it as a layer both on top of the services and also a pool that the services swim around in and it ensures that there's one place to manage traffic, the services don't have to care how to do service discovery for other services and it's all being load balanced at a reasonable price.
monitoring / "circuit breakers" would probably require a proxy I guess
you can use existing tools to introspect state (e.g. dig), and with sufficient logging it's perfectly debuggable
oxymoron?
By sending specific headers (and having your application forward them) one can have per request routing[1] and/or tracing[2].
[1] https://blog.buoyant.io/2016/11/04/a-service-mesh-for-kubern...
[2] https://blog.buoyant.io/2017/03/14/a-service-mesh-for-kubern...
Helps the instances doing X find and talk to the instances doing Y.
Need I say more? Paging Alan Kay to tell these folks what the origins of an unnecessary complicated mess looks like.