Envoy Proxy at Reddit
redditblog.com
redditblog.com
Others have mentioned that there are some gotchas with Envoy, and you mention a few about the migration bumps. Did you encounter other gotchas? And do you have any suggestions on how to avoid/mitigate their impact?
Aside from what what discussed, there weren't many things with Envoy specifically. We did have a few minor issues with the Thrift tooling, e.g.
https://github.com/envoyproxy/envoy/commit/a3c744294bca2cbf6...
but as the above indicates, they were resolved _very_ quickly.
The most important thing when making a transition like this is to have as much monitoring and observability as possible without the new tech. We were able to quickly identify and respond to issues we had with Envoy based on existing application and system instrumentation that weren't directly provided by Envoy, along with the vigilance of our engineering team.
Are you using envoy at all in your main http ingress path? You mentioned haproxy and AWS ELBs, but it wasn't clear if envoy is also being considered for public ingress traffic.
Keep up the great work!
Our HAProxy layer that routes ingress traffic to the core backend infrastructure has considerable routing logic that can be moved to Envoy and then further extended. We'd love to explore that path in the coming months.
I'm really excited about what we're building out for next year and can't wait to share as well. Feel free to reach out on reddit (u/wangofchung) or directly at courtney.wang@reddit.com for more in-depth discussion!
Quite a bit of complex routing and dynamic configurations can be provided by map files [2] and these and many other settings can be updated directly from the Runtime API [3].
With that said -- we are actively working to make things even better and intend to introduce support for updating SSL certificates/keys directly through the Runtime API as well as introducing a Data Plane API for HAProxy.
We have a new release coming any day now and this will lay the foundation that will allow us to continue to provide best-in-class performance while accelerating cutting edge feature delivery.
[1] https://www.haproxy.com/blog/hitless-reloads-with-haproxy-ho... [2] https://www.haproxy.com/blog/introduction-to-haproxy-maps/ [3] https://www.haproxy.com/blog/dynamic-configuration-haproxy-r...
Have you been thinking about adopting istio? If yes, why didn't you?
We didn't do so immediately because we did not want to immediately update all of our technology at once and felt that a piece-wise migration would be both the least disruptive to our infrastructure and safest. I think of Istio as like Smartstack in that it's not actually a complete "thing" so much as a suite of technologies that can be individually evaluated and deployed. It's very easy to fall into the trap of wanting to do everything at once, and we opted to make small progressive steps for this initiative.
1. Have you considered/are considering ISTIO control plane for your Envoy fleet? Why or why not?
2. Did you containerize your applications before using Envoy? The blog post talks about running them on autoscaled ec2 instances but its not clear if you're running application binaries on those vms or serving from containers
2. We have not containerized prior to Envoy. We're running application binaries provisioned with Puppet on EC2 for most of our infrastructure still.
In the period you had parts of your system with envoy and parts without, have your routed the outbound traffic from envoy-equiped services through their local proxy before reaching its envoy-less destination? Or did you omit envoy then?
But, K8s is the new hotness and people are going to use it.
Which I think is kind of OP's point, i.e. that maybe reddit is focusing its resources on the wrong issues. Afaik the website was and still is generally responsive from a networking point of view, there are some 503s here and there from time to time but nothing that could throw me away as a regular user, on the other hand the redesign (if it manages to overwrite all the present ways of getting past it, such as using old.reddit or i.reddit) will definitely turn me away as a regular user, and I say that as a really long-time user of that website.
The more general issue is that there are a lot of technical people around SV that do, well, technical stuff, because they are really, really good at what they're doing (so I'm in no way downplaying this post). The issue is that focusing only on technical stuff and ignoring how users actually use your website/product might turn those users away and you're left with a technical behemoth which has no users (see Google+ for a relatively recent example).
Generally speaking companies do have limited economic resources and also generally speaking yes, there is "a zero sum situation" when it comes to allocating resources inside a specific company. The management doesn't generally have infinite time and resources at its disposal and a focus on the technical side of things (or on any other specific side of a particular business) has many times resulted in neglecting the rest of the business not related to that particular topic.
A very good such example is the same Google+ case, with the now famous motto "all arrows pointed in the same direction" or some such which imho made them lose focus on a ton of other important stuff going on at the same time (Amazon and aws, most of the google products have become a chore to use since then etc)
[1] https://redditblog.com/2017/12/19/the-best-of-reddit-in-2017...