Why We’re Switching to gRPC
eng.fromatob.com
eng.fromatob.com
With grpc... It's designed by Google for Google's use case. How they do things and the design trade-offs they made are quite specific, and may not make sense for you.
There are no generated language interfaces, so you cannot mock the methods. (Except by mocking abstract classes, and nobody sane does that, right)
That's because grpc allows you to implement whatever methods you like of a service interface, and require any fields you like - all are optional, but not really, right.
Things that you might expect to be invalid, are valid. A zero byte array deserialised as a protobuf message is a perfectly valid message. All the strings are "" (not null), the bools false, and the ints 0.
Load balancing is done by maintaining multiple connections to all upstreams.
The messages dont work very well with ALB/ELB.
The tooling for web clients was terrible ( I understand this may have changed )
The grpc generated classes are a load of slowly compiling not very nice code.
Like I say, if your tech and business is like Google's ( it probably isn't) then it's a shoe-in, else it's definitely worth asking if there is a match for your needs.
For example in Go, the service definitions are generated as interfaces and come with an "Unimplemented" concrete client. We have a codegen package that builds mock concrete implementations of services for use in tests.
Zero values are also the standard in Go and fit most use cases. We have "optional" types defined as messages that wrap primitives such as floats for times when a true null is needed (has been mostly used for update type methods).
The web clients work, but generate ALOT of code. We're using the improbable package so we can provide different transports for broswer vs server JS clients btw.
The big win we've seen from grpc is being able to reason about the entire system end to end and have a central language for conversation and contracts across teams. Sure there are other ways to accomplish that, but grpc has served that purpose for us.
As a result of the above, we have been exploring twirp. You get the benefits of using protobufs for defining the RPC interface, but without quite as much runtime baggage that complicates debugging issues that arise.
I'm curious what languages you were using gRPC with. The batteries includedness across tons of languages is a big part of gRPC's appeal. I'd assume Java and C++ get enough use to be solid but maybe that's wishful thinking?
[edit: The flow control is actually done at the HTTP/2 level. However, the Go gRPC library has its own implementation of HTTP/2.]
For Java that isn't true. Java ships with a lightweight InProcess server to stub out the responses to your client.
> Load balancing is done by maintaining multiple connections to all upstreams.
Load balancing is fully pluggable. The default balancer only picks the first connection.
> The tooling for web clients was terrible
Agreed. This is almost entirely the fault of Chrome and Firefox, for not implementing the HTTP/2 spec properly. (missing trailers).
On the one level you can say new X(new MyService()) or new X(mock(Service.class))) if you have to, and on the other its just loads of jibber-jabber.
The extra "jibber-jabber" is what makes people confident that their Stub usage is correct.
Some sample bugs NOT caught by a mock:
* Calling the stub with a NULL message
* Not calling close()
* Sending invalid headers
* Ignoring deadlines
* Ignoring cancellation
There's more, but these are real bugs that are trivially caught by using a real (and cheap) server.
Agreed. It's always important to try to pick technologies that 'align' with your use-cases as well as possible. This is easier said than done and gets easier the more often you fail to do it well! I do think people will read "for Google's use case" and hear "only for Google's scale". I actually think the gRPC Java stack is pretty efficient so it "scales down" pretty well.
I want to skip over some of what you're saying to address this:
> Things that you might expect to be invalid, are valid. A zero byte array deserialised as a protobuf message is a perfectly valid message. All the strings are "" (not null), the bools false, and the ints 0.
Using a protobuf schema layer is wayyyy nicer than JSON blobs but I agree that it is misconstrued as type safety and validation. It's fantastic for efficient data marshaling and decent for code generation but it dosn't solve the "semantic correctness" side of things. You should still be writing validation. Its a solid step up from JSON not a panacea.
Do you consider protobuf superior to those alternatives for web-based (rather than server to server) projects?
My impression is that if you're going to talk to a browser, that edge stands to gain much more from conforming to HTTP standards. If your edge is more "applicationy" and less "webpagy" then maybe a browser facing gRPC (or GraphQL?) might be more appealing again.
As to the other JSON schema systems, I kinda wish one of them won? It feels like a lot of competing standards still. Not really my area of expertise.
Also: why wouldn't grpc work well with load balancers? It's based on HTTP/2. It's well supported by envoy, which is fast-becoming the de facto standard proxy for service meshes.
You answered your own question.
There is always some bit of older infrastructure, like a caching proxy or “enterprise” load balancer that doesn’t quite understand http/2 yet - it is the same reason so much Internet traffic is still on ipv4 when ipv6 exists - the lowest common denominator end up winning for some subsection of traffic.
We use gRPC in our tech stack but some pods handle far more connections than others due to the multiplexing/reuse of existing connections.
Sadly Istio/Envoy solutions are in our backlog for now.
We can't fault gRPC otherwise. It's way faster than if we were to encode/decode json after each microservice hop. It's integrates into golang nicely (another google coolaid solution!) so a win-win there.
Doesn't work on ALB because its HTTP/2 support is trash, not gRPC's fault here. Works fine with NLB btw.
> Load balancing is done by maintaining multiple connections to all upstreams.
Again, this is a "feature" of HTTP/2. Use linkerd or envoy that support subsetting among other useful things.
Don't blame your misunderstanding of how technology is meant to be used on said technology.
How does this work? How do you make, say, all fields but the second null? Do you just send a messages that's (after encoding) as long as the first two fields, where the first field is 0x00 and the second contains whatever data you want?
Disclaimer: Java fanboy bias. For services internal to a company, I think gRPC is an all around win. If you need to talk to browsers integrations, I don't have as many opinions.
Personally, I really prefer working at the RPC layer rather than at the HTTP layer. It's OOP! It's SOA! Pick your favorite acronym! HTTP's use as a server protocol (as opposed to a browser) is mostly incidental. It works great but most of the HTTP spec entirely inapplicable to services. I like named exceptions to 200 vs 5xx 4xx error codes. Do I really care about GET, PATCH, PUT, HEAD, POST for most of my services when all of my KV/NewSQL/API-over-DB services have narrower semantics anyway.
Out of band headers are nice though.
Between protobufs, http2, and a fresh, active server implementation we see pretty solid latency and throughput improvements. It's hard to generalize but I suspect many users will. Performance isn't the only driving factor but it's nice to start from a solid base.
I'm sure missing all the tools like curl and friends is an annoyance, but I like debugging from within my language, and in JVM land at least it's been easy enough.
Only downside I can think of is that there's no analogous mechanism to gRPC streams; you have to implement your own pagination.
From what I understand of it, the big idea is that instead of passing parameters from the client to the server and fully implementing the query logic, stitching, and reformatting etc. on the server side, you now have a way to pass some of that flexibility out to the client. Instead of updating both the server and the client as uses change, more can be done from the client alone.
I spend most of my time on the infra side of things and rarely if ever make my way out to the browser so I can't speak to WebSockets/SSE or web friendliness. Being the "backend-for-backend" I just prefer being more tight-fisted about what my clients can and can't do. I mostly deal with internal customers with tighter SLAs so I like to capacity plan new uses.
Maybe I'm just old fashioned.
Until now it works perfectly as expected. The C++ code generator provides a clean abstraction plus it saved a lot of time (both in programming and debugging). The gRPC proto file syntax also nudge you in the right direction wrt protocol design.
When trying to "sell" gRPC it helps that there are generators for plenty of languages and it's backed by a major company.
thrift --gen <language> <Thrift filename>
This handles the following languages: C (depends on GLib), Cocoa, C++, C#, D, delphi, Erlang, Go, Haskell, Java, Javascript, OCaml, Perl, PHP, Python, Ruby, Smalltalk. From this one command I can instantly integrate this into almost every build system I know of.On the other hand gRPC has the same features but it's a slightly different workflow. For each language you go to the language's code generation page, find the command line option for generating your language's code, read up on some decisions that were made for you, etc. All of that is fine, the part that annoys me a bit is each language needs a language module for the compiler (if it's not one of the core few languages). For example in the documentation for generating Go [1] they have you download and install protoc-gen-go from http://github.com/golang/protobuf assuming that you already have golang installed.
gRPC seems much more focused on the idea that I want to define an API for my code, I want that specification to live inside the project that I am writing, and that you can figure out how to generate stubs on a language by language basis.
What I want is something I can write a set of system specification files, type one command and get modules built for all languages. From there I can import those modules using my native language's favorite package manager (npm, composer, Hunter for CMake, etc). Ideally the Protocol Specification, the Library Generation, and the Library Usage are three components that are separate.
[1] - https://developers.google.com/protocol-buffers/docs/referenc...
This way your protocol for your infrastructure is just another library.
Addition: Borg and kubernetes are designed with similar goals but different emphasis. They are like complementing twins had different personalities. For this I recommend Min Cai's Kubecon'18 presentation about peleton [1], the slide is titled "comparison of cluster manager architecture".
[1] https://kccna18.sched.com/event/GrTx/peloton-a-unified-sched...
Google aside, many other companies like Dropbox rely on gRPC extensively to successfully run infrastructure: https://static.sched.com/hosted_files/grpconf19/f7/Courier%2...
Kubernetes is not a bad system but it’s not designed to run Google.
More specifically, "No more than 5000 nodes" and "No more than 150000 total pods" is fairly limiting to large (Google-large) clusters.
"ClusterData2011_2 provides data from an 12.5k-machine cell over about a month-long period in May 2011."
There's plenty of internal teams that use GCP. Increasingly this might be the direction things are heading.
[1]: no-one that we care. At Google this is obviously always incorrect. There’s always that someone who uses weird things like mongoDB and AWS.
Indexing doesn't run on GCP primarily because it's legacy (as in, the first product Google ever did) and thus long predates GCP itself.
Borg is cirra 2003, pb/stubby was before that. Gfs probably was similar in timing as stubby. And many other cluster level foundations. In the end, Borg is the true corner stone that ties everything together and completes the Google infrastructure puzzle (or the modern global scale cluster computing).
This is incorrect. I suspect you're overextending proto3's treatment of unknown fields to include discarding incorrectly typed fields too. If A has field 1 types as an int, and B has field 1 typed as a string, an A message with field 1 set will not parse as a B message. However, if the A message has no fields set, or sets a field number unknown to B, that could parse successfully with "leftover" unknown fields.
message enc {
int foo = 1;
SomeMessage bar = 2;
}
message dec {
bool should_explode = 1;
string why = 2;
}
You can successfully decode the latter from an encoding of the former.Change field 2 to a bytes field instead of a string field and then yes.
Protocol Buffers should generally be non-destructive of the underlying data. That means even if it encounters the wrong wire type for a field, it should simply retain that value in the unknown field set rather than discard it.
In the C++ reference implementation, which I wrote, this is not true. The field 1 with the wrong wire type would be treated as an unknown field, not an error.
It's possible that implementations in other languages have different behavior, but that would be a bug. The C++ implementation is considered the reference implementation that all others should follow.
However, shereadsthenews' assertion is not quite right either. Specifically, a string field and a sub-message field both use the same wire type; essentially, the message is encoded into a byte string. So if message A has field 1 type string, containing some bytes that aren't a protobuf, and message B has field 1 type sub-message, then you'll get a parse error.
But it is indeed quite common that one message type parses successfully as another unrelated type.
Yeah I was about to say, protobuf C++ implementation will definitely treat it as an unknown field. I just had it do that a few days ago. :)
[1] https://github.com/protocolbuffers/protobuf/issues/2497#issu...
I'm personally strongly in the "required" camp because at least the interface makes an attempt at giving clues to a user as to what fields are important. If everything is optional, there's no information being passed as to what is important anymore.
Depending on your use case, you may have to do a lot to work around protocol buffers not being self-describing. I haven't seen a good description of the problem online, but if you find yourself embedding JSON data in protobufs to avoid patching middleware services constantly, you should look at something like Avro or MessagePack or Amazon Ion.
The great thing about Clojure is that you can make holistic systems that are end-to-end immutable and value-centric, which means the side effects can go away, which means gRPC stops making sense and we can start building abstractions instead of procedures!
I'm writing this as I take a break from working on a polyglot project made up of Kotlin, Rust, Node, where we use gRPC and gRPC-web. We're slowly stealing endpoints from the other services/languages to Rust.
Without focusing on the war of languages, the codegen benefits of protobufs have made what used to be a lot of JSON serde much easier.
[0]: https://cloud.google.com/pubsub/docs/reference/service_apis_...
[1]: https://github.com/googleads/google-ads-java/blob/master/goo...
Phase 1: Ad hoc, free-for-all chaos.
Phase 2: SOAP tries to bring order. It fails mainly because the "S" ("Simple") is a lie.
Phase 3: Pendulum swings hard toward simplicity with HTTP plus JSON plus nothing else, thanks.
Phase 4: Things shift possibly more toward the middle (a little structure), but none of the competing systems have become obvious winners.
“binary” is not necessarily a benefit, particularly for developer tooling/debugging
Streaming is one area it may have a benefit but honestly the other issues outweigh that possible benefit (and it’s not like there aren’t other ways to stream data to a browser without resorting to polling)
Why can't you stream a JSON array?
Edit: Here's a (hastily created and untested) node.js example, even:
class JSONArrayStream extends Transform {
constructor() {
super({readableObjectMode: false, writableObjectMode: true});
this.dataWritten = false;
}
_transform(data, encoding, callback) {
if (!this.dataWritten) {
this.dataWritten = true;
this.push('[\n');
this.push(JSON.stringify(data) + '\n');
} else {
this.push(',' + JSON.stringify(data) + '\n');
}
callback();
}
_flush(callback) {
this.push('\n]');
}
}Editing, because I can’t reply: I was specifically going to mention SSE/EventSource but expected an immediate “not everything is a browser” response.
Editing the 2nd: yep, that’s kinda what I meant by “tricks” - essentially splitting out chunks to pass to the json parser.
- how can a client send an error message if it runs into problems mid-stream? You can invent a system, but you’re walking into an ad-hoc protocol pretty fast; why not use something others wrote?
- what if the remote end wants to interrupt the stream sender to say “stop sending me this” for any reason? For example, an erroneous item in the stream, or a server closing down during a restart.
- grpc supports fully bidirectional streams, interleaving request and response in a chatty session; how do you do this?
Not that the original article mentioned these. I bristle though when I hear the engineer’s impulse to “why don’t you just”-away at something.
As someone currently responsible for migrating services to gRPC at my company, I always say that the main reason people are switching is "because Google."
While there are merits to gRPC and protobuf, I don't think there are enough advantages to throw away REST and JSON and all the tooling around it. The moment you start switching, you start feeling the pain.
"Because Google" is also the main reason everyone wants to run their crap on Kubernetes. It's pure hype.
Always makes me think of "You Are Not Google": https://blog.bradfieldcs.com/you-are-not-google-84912cf44afb
That’s however just an implementation detail. JSON can easily be written and read as a stream.
Switching your whole architecture, dealing with a binary protocol and the accompanying tooling issues just because of your choice of JSON parser feels like total overkill.
JSON over HTTP is ubiquitous, has amazing tooling and his highly debuggable. Parsers have become so fast that I feel they might even have the opportunity to be faster than a protobuf based solution.
Finally I don’t buy the argument about validation. You have to validate input and output on the boundaries no matter what.
Even when your interface says “this is a double”, it says nothing about ranges (as seen in the article where valid ranges were specified in the comment) for example.
Not even close. Event new JSON serializers/deserializers aren't magic. Protobuf is a LOT easier to parse, so it's naturally a LOT faster.
First two duck results for "json vs protobuf benchmark":
https://auth0.com/blog/beating-json-performance-with-protobu...
https://codeburst.io/json-vs-protocol-buffers-vs-flatbuffers...
Even at a 5x improvement, most projects will never reach a point where the transport encoding is a bottleneck. Protobuf has a lot going for it (currently using in a project) but can’t be sold on speed alone.
True, but if you're wanting an implementation you can use in Javascript running in the browser, it may accurately reflect reality. You have a high-quality browser-supplied (presumably native) implementation of JSON available. For a protobuf parser, you've just got Javascript. (You can call into webassembly, but given that afaik it can't produce Javascript objects on its own, it's not clear to me there's any advantage in doing so unless you're moving the calling code into webassembly also.)
I don't think browser-based parsing speed is important though. It's probably not a major contributor to display/interaction latency, energy use, or any other metric you care about. If it is, maybe you're wasting bandwidth by sending a bunch of data that's discarded immediately after parsing.
> Simple service definition
> Define your service using Protocol Buffers, a powerful binary serialization toolset and language
It's a little unfair to call it that orthogonal.
We've been working hard on OpenRPC [0]. An Interface Description for JSON-RPC akin to swagger. It's a good middle ground between the two.
[0] Instead of passing in a plain object, you build it as such:
const userLookup = new UserLookupRequest();
const idField = new UserID();
idField.setValue(29);
userLookup.setId(idField);
UserService.findUser(userLookup);
The metadata field doesn't seem to mind though..."While more speed is always welcome, there are two aspects that were more important for us: clear interface specifications and support for streaming."
This offers a quick exit for anybody who already knows about these advantages.
i _suspect_ for google-scale it should all be fine, where available cpu's are essentially limitless, and consistency of data gets handled f.e. due to multiple updates etc. at a different layer.
writing safe, performant, multi-threaded code in presence of signals/exceptions etc. is non-trivial regardless of how your 'frontend' looks like. async-grpc is quite unwieldy imho.
i have heard folks trying grpc out on cpu-horsepower-starved devices f.e. wireless-base-stations etc. and running into aforementioned issues.
gRPC adds bi-directional streaming which is not possible in http, but the use cases for that are more specialized.
That’s why formats like FlatBuffer where written; however parsing is likely not going to dominate your application, so other factors should influence your decision instead.
Be careful not to back broad arguments with outlier benchmarks.
In general, it is plainly true that JSON is much more computationally difficult to encode and decode than Protobuf. Sure, if you compare a carefully micro-optimized JSON implementation against a less-optimized Protobuf implementation, it might win in some cases. That doesn't mean that Protobuf and JSON perform equivalently in general.
You’ll never need to move to protobufs due to parsing speed.
Exactly. If you need performance, you'll use Cap'n Proto or FlatBuffer, which use a native binary interface, so you don't need to create and copy objects, you just map them them in from IO.
Sure it does. You can implement "streaming" in Cap'n Proto by introducing a callback object, and making one RPC call for each item / chunk in the stream. In Cap'n Proto, "streaming" is just a design pattern, not something that needs to be explicitly built in, because Cap'n Proto is inherently far more expressive than gRPC.
That is, you can define a type like:
interface Stream(T) {
write @0 (item :T);
end @1 (); # signal successful end of stream
}
Then you can define streaming methods like: streamUp @0 (...params...) -> (stream :Stream(T), ...results...)
# Method with client->server stream.
streamDown @0 (...params..., stream :Stream(T)) -> (...results...)
# Method with server->client stream.
Admittedly, this technique has the problem that the application has to do its own flow control -- it has to keep multiple calls in-flight to saturate the connection, but needs to place a cap on the number of calls in order to avoid excess buffering. This is doable, but somewhat inconvenient.So I am actually in the process of extending the implementation to make this logic built-in:
https://github.com/capnproto/capnproto/pull/825
Note that PR doesn't add anything new to the RPC protocol; it just provides helpers to tell the app how many concurrent calls to make.
[0] https://github.com/reactive-streams/reactive-streams-jvm
Problems I don't have when using Erlang/Elixir umbrella apps + OTP.
I do remember reading quite a lot of articles about the inverse of this: "we don't hire people using Windows/IDEs/ because it says a lot about them, a craftsman should chose his tools wisely, ...", but never the positive.
I don’t remember any “we don’t hire people on Windows/using an IDE” (the last part would be particularly weird IMO), but I wouldn’t be surprised if somewhere said “if you want to use Windows you’re on your own (support wise) and if it becomes a time sink you switch or find work elsewhere”.
I’ve supported (in terms of dev environment/tooling) people on Macs, Windows and Linux. Windows by far had the weirdest issues to solve/avoid.
Two famous examples:
http://charlespetzold.com/etc/DoesVisualStudioRotTheMind.htm...
> If you are a startup looking to hire really excellent people, take notice of .NET on a resume, and ask why it’s there.
https://blog.expensify.com/2011/03/25/ceo-friday-why-we-dont...