Many reasons to switch from JSON to gRPC: * gRPC uses HTTP/2 which means you can concurrently make multiple requests on a single TCP connection, while on a typical JSON API which probably uses http/1.1, you can't. * Type safety of protobuf types. * Bi-directional streaming e.g. Kubernetes controllers use the Watch API which notifies the object changes/add/deletes in etcd to the control plane. * Client libraries are automatically generated, and not error prone. * RPCs are already optimized for bytes on the wire efficiency, whereas JSON is not. This matters a great deal as Kubernetes objects get large in size/quantity over time but controllers still work effectively by not spending too much CPU on encoding/decoding like they do on JSON. * Similar to the previous point, most json decoders don't reuse objects, so every decoded object is a new alloc, whereas gRPC can Reset() and reuse the same object while decoding/encoding. * Builtin authentication primitives (such as JWT/tokens or even adding TLS to client and/or server). * gRPC has support for interceptors which are like middleware functions you can inject to requests/responses on both client and server-side for logging, authorizing etc.
The list goes on, but something to note is that etcd was not developed by Google (and as far as I know, not by ex-Googlers). Both etcd and gRPC are owned by the same open source foundation, so it's natural that they make use of an existing technology.
What are you actually talking about? An HTTP API by itself doesn't actually care about which http version is used, that's something only the HTTP server that serves the API should care about (nowadays, pretty much any http server in any language supports HTTP/2 and 1.1).
> Bi-directional streaming e.g. Kubernetes controllers use the Watch API which notifies the object changes/add/deletes in etcd to the control plane.
Even HTTP/1.1 has mechanisms for that, like long polling and chunk encoding (which allows infinite streams to keep a dual channel of communication open where each chunk can be treated as a message - with "headers" and all).
> Client libraries are automatically generated, and not error prone.
This is true, but we've had similar technology since the SOAP days... you could easily do the same with a XML-based HTTP API.
By the way, most of your points are against JSON, not a HTTP API per se, which usually would allow a number of formats , including XML and JSON at least.
> Builtin authentication primitives...
HTTP also has that when you include cookies and something like OAuth/OpenID.
> gRPC has support for interceptors which are like middleware functions
This kind of thing is better done by using the specific platform you're running on (depends on the language) so I don't see something like this as something desirable on the specific RPC framework you're using. With HTTP, caching, logging etc are trivial to do and one of the strongest advantages of using HTTP in the first place, so I think you're confusingly making a point for HTTP here, unless your focus is on the RPC side of things? In which case, you would have to make a case on why RPC is a better fit than HTTP for etcd, which you haven't.
So all in all, I found your points utterly unconvincing, but presented with so much conviction that you're actually right that I could not resist to respond (even if I don't want at all to get into a pointless discussion on the merits of HTTP VS RPC or JSON VS Protobuffers)!
For example, whole etcd API is located at https://github.com/etcd-io/etcd/blob/4c6881ffe4b3bae257c0720...
It is pretty straightforward and not too hard to understand the API. Designing APIs with protobuf makes things convenient and it brings lots of already written tools with it.
Since gRPC built on top of HTTP/2, we thought gRPC as easiest and performant way of writing a HTTP API with good defaults.
Yes it's easy to make that RPC call. But should you?
I've been off doing other things in the intervening period, but while I had my back turned the industry seems to have turned its back on REST and gone whole hog on RPC. Again.
At first I was thinking this was just internally here at Google, where protobufs and gRPC reign supreme. But it seems to have taken hold everywhere.
What did I miss? Why have we swung this way. Again. Is the pendulum going to go back?
I've heard this argument before, but how does gRPC itself cause this issue to manifest? I'm curious to hear what your opinions are on a better alternative, and how not to represent a remote resource as a local one.
But the fact that it presents the remote resource in an API which resembles a local object means that programmers often get lazy in the manner of which they think about these things. The REST semantic is supposed to make this more explicit.
A remote object is not an object in your program or even your computer. It's something you're taking from something that is computationally miles and miles away. Compared to the microseconds it takes to dispatch a local call, it's an eternity away, even on a local network of the highest speed.
Accumulate those latencies over thousands of dispatches, and trouble can ensue.
I am reminded of an observation from when I first joined Google, coming out of their acquisition of the scrappy awesome ads company I worked at (Admeld).
We had a little service that kept track of ad impression caps / budgeting. We didn't want to serve an ad a single time more than the customer wanted us to, etc. Serving many thousands of ads per second, a process distributed across multiple machines in multiple data centres, this is a bit of a tricky shared state problem. The people who came before me had designed a rather clever solution which used a form of backoff to trickle down the number of ads served as they got closer and closer to budget cap, and to synchronize this state across clusters (this was before there were Rafty services to make this kind of thing easier, BTW).
We had a bit of a show and tell with Google when we first joined. I wanted to know how they were handling this problem, since their scale was many times ours, so I asked the question and got a puzzled/annoyed look:
"Oh we just make an RPC call to our budgeting server."
Summary: if you're at Google you don't have to worry as much about these problems. You still do, but there's an insane amount of infrastructure and horsepower and an army of SREs to help make it happen.
So, yeah, my point is -- just because Google does something or has invented something doesn't mean it's the best way to do it, especially in a smaller more cost conscious organization.
What I don't see though, is how is making a (g)RPC call any different from making a REST call? Like you said, the REST call is supposedly more explicit, but at the end of the day, it seems more of a convention and some hard underlying difference. What's the difference between `httpClient.get("...")` and `grpcClient.foo(...)`?
I mean, internally at Google we have load balancing for grpc (I'm sure the outside world does now, too, but it was new to me when I joined) -- but load balancing HTTP requests containing readable JSON or XML documents, that's far more sysadmin friendly, wouldn't you say? Off the shelf infrastructure, nothing exotic, easier to monitor. Same goes for caching, for proxies, etc.
Being able to just stick a URL for a given resource in your browser, or hit it with wget/curl to read it, that's a serious bonus.
In general URLs follow conventions similar to those down by our Unix forefathers, when they designed the filesystem API. We are all familiar with this model. And in some ways REST done right is very Unix philosophy -- provide a consistent model upon which a bunch of little tools can interoperate.
I could go on... have to go put my daughter to bed tho
I don't think that's accurate... as the article claims (and it seems believable) it has only appeared in etcd because of Xooglers interfering with the project in the name of Kubernetes (also from Google). Or do you actually have examples of gRPC being used in many other non-Google-related projects?
Was a "JSON mapping for etcd's protocol buffer message definitions" the original "HTTP API"?
"A Critique of the Remote Procedure Call Paradigm" (https://pdfs.semanticscholar.org/e125/57a7582881a62040eee68b...)
"A Note on Distributed Computing" (https://github.com/papers-we-love/papers-we-love/blob/master...)