Announcing Envoy: C++ L7 proxy and communication bus
eng.lyft.com
eng.lyft.com
Disclaimer: I work on Proxygen at FB.
Whether it's true or not, you look like you're just using the excuse to announce to everyone that you work at FB. As a person who works at Google, I understand the temptation... but you should fight it. It's super annoying.
The one thing I must have is SNI. The docs only have a short blurb [1], does someone know the full status of SNI support?
[1]: https://lyft.github.io/envoy/docs/intro/arch_overview/ssl.ht...
EDIT: Also is there any kind of visualization for the resulting network topology? It looks like Envoy should know everything about who is talking to who.
Many existing RPC systems treat service discovery as a
fully consistent process. To this end, they use fully
consistent leader election backing stores such as
Zookeeper, etcd, Consul, etc. Our experience has been
that operating these backing stores at scale is painful.
[1]: https://lyft.github.io/envoy/docs/intro/arch_overview/servic...We have had zero outages caused by our eventually consistent discovery system with active health checking (knock on wood), and haven't really touched the discovery service code in months. It just runs.
I'm not saying that a system using ZK, etc. can't be made to work. It certainly can since many companies do it. It's mostly that I think those solutions are actually making the overall problem a lot more complicated and prone to failure than it has to be.
However it also does load balancing. But doesn't that defeat the purpose a little bit? If your monitoring tool is the same as your load balancing tool, then who's monitoring the load balancer? :) I might be misunderstanding the architecture here.
Because all inter-service requests are going through Envoy, it is really easy to keep incredibly detailed stats about network health, request success rate & more.
Envoy performing the task of load balancing does not defeat the purpose, because it provides extremely detailed stats for ALL THE THINGS, they reported it helped them find problems much quicker, instead of checking service code, EC2 networking, or the ELB. Essentially by creating a supersolution with better stats reporting for all, troubleshooting seems like it would be easier.
#1: cost. Doesn't it basically cost double to move the request from the client application to the proxy, and then from the proxy to the backend?
#2: upgrades. What happens to the clients when the proxy is being rolled out?
#3: head-of-line blocking. If an application has two streams to the same backend, one stream which is low priority and one which is higher, how does the proxy handle that?
#1: the cost of local proxying is very small, on the order of tenths of ms. it's going to be hard to distinguish against the backdrop of the broader network latency. also, in our environment we've been thinking of only running the proxy on the client end for now, to make the migration easier.
#2: as with haproxy, you leave the existing connections handled with the running proxy, and new connections are taken up by the new proxy. unlike haproxy, envoy actually supports live config reloading, so you don't have to pay that penalty on every config change
#3: multiple connections are multiplexed, but i'm not sure how you specify priority to the proxy. do the docs specify a mechanism? i imagine both requests are served in parallel, in no particular order, and the data is buffered for the client.
I work at Lyft. To answer your questions:
1) There is added cost, though it varies depending on how many things Envoy is configured to do (e.g., logging, tracing, stats, rate limiting, health checking, etc.). Even in complex scenarios (Envoy being used to proxy both inbound connections into a service, as well as proxy outbound connections to Mongo or Dynamo), we measure Envoy overhead to be < 1ms, which for almost all applications is negligible. There are definitely certain cases where this might be prohibitive, but in general we find the common functionality we get (again stats, tracing, etc.) to be invaluable in a production setting.
2) Envoy supports hot restart (https://lyft.github.io/envoy/docs/intro/arch_overview/hot_re...), as well as graceful drain of existing connections, so there is limited/no impact to existing clients. There is one enhancement that we would like to make to our HTTP/2 graceful draining to make it even more seamless, but that is more complicated than I can type here. :)
3) Right now Envoy does not support HTTP/2 priority so all streams are treated equally. We are currently working on priority support at the routing layer, with different connection pools available for high and low priority traffic, as well as circuit breaking settings. In the future we will likely merge this back into a single HTTP/2 connection with proper priority support. In practice though, within the DC, head of line blocking at the TCP layer isn't too much of an issue.
However, I couldn't find a performance benchmark test or something compared to alternatives such as haproxy, nginx, etc. So I'm going to make my hands dirty now. ;)
I find it odd that they did not include Rust in their list of preferred languages. Rust is safer than C++ and Go.
All those other languages have their niches and ecosystems and are safe enough.
> very productive but not particularly well performing languages such as PHP, Python, Ruby, Scala, etc
Which seems to be missing a certain more-productive-than-C++ and very well-performing coffee-themed language.
I mean, you would never use Java for this, because although it could go fast enough, it would need way too much memory to do it. But i would have liked to see it dismissed for that reason rather than glossed over!
For this we have a secret management system, called confidant (https://lyft.github.io/confidant/), that we use to distribute any necessary secrets. So, yes, you may need to have keys on every node (depending on your monitoring system), but assuming you securely distribute them, it's not a big deal.
This is, of course, a general problem that's not necessarily related to envoy.
I agree it's a general problem. But sometimes certain architectures would require more vulnerable approaches vs others.