Coding these higher-level concerns yourself is laborious and error-prone, so if your application runs on Kubernetes, then service mesh can be a substantial help.
Now this all worked, since "grpc" (stubby) did all this (I guess) - then each borg machine, magically collected these metrics/logs/etc. and if these were not super useful to us - normal developers, they were the first thing an SRE would ask (especially as we moved to spanner later). Often on call, if you have an issue, you'll just have to increase the sampling to 100% for some time (say 30 seconds), and it'll capture quite enough for SRE to take a look (as it happens, this would be done when there is an incident, or things are going slow, bad, etc.).
Now back, in a gamedev company, where we started having (without accepting yet) "micro"-services - like things talking to various caching backends, things sitting behind rancher, postgres, mysql, custom breed services - but they all talk directly through sock() api, or http, etc. - e.g. collecting metrics from them is usually - whatever they expose (if they) to prometheus - but you don't get the whole idea.
Now, and I could be wrong, but this is my understaning - rather than rewriting these services to use GRPC, or something to plug same metrics, you can instead make these piece of software don't talk directly to each other but talk to a "service mesh" sidecar (daemon, etc.) then itself it'll talk to another "service mesh" (itself) on another node/machine, but by doing this (and I guess some proper configuration), you'll get the tracing information you need.
So at least to me, it's an escape hatch to place in front of something like mysql, postgresql, etc. but still get it in the e2e picture - e.g. user have sent a request, and we want to see where it went everywhere...
And this is where I see the value of the service mesh. There is also the cases of handling retry errors, "flaky" servers (unhealhty) through circuit breaking and passive health, or cumulative timeout (better wording here?) - e.g. 500ms timeout from the first request, decreasing the timeout with each subsequent call, thus dropping eventually.
Then authentiation/authorization, rather than re-implementing in several languages (or much of the above) you do it in one language (C++ for envoy!)
The elephant in the room - is how much is spent on this extra "hub" in the communcation. Also things like UDP support, and who knows what else (simple app developer here, I don't know network details)