go-micro seems like it does a bit too much, like service discovery and balancing within the framework when that's likely better handled by an Envoy/Istio.
go-micro seems like it does a bit too much, like service discovery and balancing within the framework when that's likely better handled by an Envoy/Istio.
Sidecar proxies also have the fatal flaw that they can't propagate a distributed trace, so if incoming RPC call triggers outgoing calls (common in microservices, for better or for worse), then you lose the link between them. To do that, your application has to remember the trace ID and pass it along from the server to the client. For reasons like that, I'm kind of bearish on premade service meshes. But, they are OK for mTLS.
gRPC has added a lot of code in recent years to support xDS, which lets clients discover endpoints, load assignments, locality, etc. and lets your app do the right thing without an intermediate proxy. It can just connect directly to the control plane and make the same routing decisions a proxy would. That's the ideal situation to me.
I've also seen it cause a lot of pain during upgrades since upgrading the service framework is implicitly upgrading so many other components and the framework is probably not being diligent about only improving one component (ex. Auth) per release. Sometimes that leads to basically never upgrading to avoid the risk.
It's true you need to do propagation for tracing but open telemetry has pretty broad support these days. That cost seems worth the stability and control.
Netflix eventually moved away from HTTP+JSON to gRPC for backend services, and also started to expose internal DNS-based service discovery. This removed the need for that sidecar from _most_ places, since you could just generate code from protobufs, but it created other problems. Specifically, it was hard to apply the same resiliency patterns across multiple languages when calling those APIs, and so that would then contribute to weird incidents.
I'm not there anymore, but last context I have is that things are moving to an Envoy-based model now. The idea is that everything will use the sidecar, and resiliency patterns will be baked in at that layer. I suspect that'll be a better architecture in the long run.
> also started to expose internal DNS-based service discovery. Then this is similar to Consul-based service discovery. In that case, do you know how Netflix continued to support their client-defined routing rules, like selecting a specific set of nodes of a given service to route traffic to? I'd imagine that the number of IP addresses in an ARecord is limited too, so a service won't retrieve the entire list of IPs from a DNS record?
This is common even in xDS implementations; client services don't need to know every possible endpoint.