Why Segment Chose Go, gRPC, and Envoy to Build Their New Config API
stackshare.io
stackshare.io
Expect more on the state of the art as KubeCon vids come online ;)
https://www.youtube.com/channel/UCvqbFHwN-nwalWPjPUKpvTA/vid...
1. Protobufs importing each other becomes effortless. You don't need to download multiple repositories, if you have the proto repo, you have all the dependencies.
2. Decoupling "ownership" allows for a bit more freedom when creating proto files that describe shared data structures.
3. You get a singular location to browse infra. It doesn't describe relationships of course, but it's handy to be able to see what services you have and what their API looks like.
Disclaimer: I don't use gRPC, but I do use Protobufs & an RPC, Twirp.
It's probably best to separate the proto file for server and client and treat it as a dependency.
If you're using a mono-repo the solution is possibly a little simpler.
We keep them all in one repo.
Right now the entire API is a monorepo so it’s easy. But I plan to split it up soon...
All the protos in one repo. Ideally public so Segment users can use it as a reference too!
Then all the generated Go clients, servers, mocks, typescript, swagger, etc in another repo. This is what programs can import.
Then we can build gRPC services in their own repos if we want to. And it’s easy to produce and consume messages for all the other services.
So before pulling the switch too far, I decided to take a step back and assess with benefits we all got from gRPC. Mainly it was Protobuf spec, RPC design (ie not REST verbs and whatnot), and code generation.
Because those were the things we all agreed upon, I went with Twirp[0]. It's been a great experience so far. The rollout is just a bit ahead of where our gRPC was, and it's been mostly effortless.
The most effortless thing? I've got an (internal) API user switching from the old REST API to the proto RPC spec, and all they're changing is the URL and a few field naming schemes. It's still JSON for them. They're not generating code, they're not doing anything. It all "just works", which is pretty impressive.
TL;DR: We love simpler options that let us gain the benefits of Protobufs, without the complexity of gRPC when we don't need it.
[0]: https://blog.twitch.tv/twirp-a-sweet-new-rpc-framework-for-g...
For me, the reason to use GRPC is because there is a very good story about the boring details of networking. The stuff that would make other programmers eyes glaze over but are nonetheless important, has basically been solved by GRPC. What happens when the connection gets dropped but you haven't sent the RPC yet? Does it try to find another server to send it to? Does it time out? What happens if the DNS entries change after you connect? Does it even use all the addresses that DNS returned? If the connection has been idle for for 30 minutes, does it keep it around or shut it down? How does it even know if the connection is good after being idle so long? Home routers have a bad habit of breaking connections due to their NAT.
For most of these problems, TCP and UDP only gives you the tools to solve them, but don't do so out of the box. Coming back to Twirp, I don't recall seeing these things addressed in their code, (except what Go's net and http library already provide). If you don't care about those networking details (and it's completely legitimate if you don't, some people don't need to), then I think Twirp's API is much more pleasant to use, and still gets you 80% of the benefit of GRPC.
There isn't anything in the world that can match grpcs features and still perform at the same level, plus Google's decades of large scale use.
Guys, I am so disappointed at how Google push it's internal technology. People won't come to you without pitching, be smart...
Agree completely that .proto offers big advantages on its own.
Also agree that grpc is complex and challenging. We considered twirp too.
A few things that pushed us towards gRPC...
Canonical error codes and semantics. Standards here go such a long way to building reliable systems.
The ecosystem. There are a lot of middlewares, validators, transcoders written specifically for the Go gRPC stack.
Envoy. This has a lot of things we wanted like rate limiting and auth which are implemented over gRPC. With AWS adopting Envoy for their service mesh offering it’s a safe technology bet now.
Performance. This doesn’t matter for our API stack but could be a huge gain in Segments data plane.
That said I still wish all this was simpler. Twirp looks like a great choice in that regard.
First and foremost our customers universally favor REST APIs. GraphQL and gRPC are better tech wise but REST is more familiar.
Next gRPC uses pure HTTP/2 which doesn’t work with an AWS ALB.
Thanks to the grpc ecosystem the HTTP/1 and REST transcoding is just an Envoy config.
We use the envoy authz filter.
For every incoming request envoy first calls out to a custom auth check service with all the request metadata like path and http headers.
The auth service can return a “fail” response which indicates to not forward the original request any further. Or it can return a “pass” response plus data to add to the original request headers.
Docs here:
https://www.envoyproxy.io/docs/envoy/latest/configuration/ht...
In analytics space, for example, we built end points in Go which collect data and post it to Bigquery or Redshift.
Helping startup create their data pipelines.
We use Go everywhere whenever we need to glue things together.
One cool thing about Go is that it doesn't take much time to understand what's happening in the codebase which we've is good for dodging mistakes in app receicijg billions of events per week.