gRPC in Production
about.sourcegraph.com
about.sourcegraph.com
Insert: Object comes in (built in RPC system for both), object is assigned ID (which is added before serialization), object is written to RocksDB, object is serialized and indexed.
Query: List of strings come in, strings are looked up in tag index, intersection of results is done, objects are retrieved from RocksDB, objects are deserialized, objects are added to list, objects are returned to client.
FlatBuffers was unfortunately a no-go since it doesn't permit modification of fields that haven't yet been set and it didn't have a sane way of making a copy (which would also be slow).
Protocol Buffers worked but the pure Python driver is incredibly slow and the C++ driver for Python was non-default and marked "experimental". I eventually adapted to it but it was still quite slow even on the server side. From memory a response with C++ client and server took ~9ms, 4-5ms of which were spent on serialization/deserialization.
Cap'n'Proto eventually won for me. The Python driver is unstable and memory usage soars like an eagle until Linux shoots it down with an OOM but the server was much faster. Typical response times were closer to 3ms, most of which was spent in RocksDB or the actual index lookup. A downside though was that Cap'n'Proto has its own built in and odd library for async stuff and doesn't really support threading.
The generated code is not particularly efficient(maps of strings to function pointers) and not very modern. The code itself relies on boost shared pointers everywhere, which is hugely wasteful, and the actual process of getting thrift up and running was on trivial on windows. (Still, easier than grocery was and capn porto at the time). The compiler itself is a bit funny, and hasn't been particularly easy to integrate into our build system (cmake).
The main argument I've heard in favor of gRPC is that it is supported in more languages. Otherwise, I'm not sure what would make gRPC "much better at RPC" -- Cap'n Proto is actually much more expressive, and has features like capabilities and promise pipelining which gRPC lacks.
Do you remember the arguments you saw? I'd be curious to know what they were.
(Disclosure: I'm the author of Cap'n Proto and also Protobuf v2.)
Why was gRPC nicer to work with in my opinion?
- Comprehensive and readable documentation. Cap'n'Proto has a high level overview on the main page but the documentation is far from complete. For example regarding I/O, a core part of many RPC systems, "Function calls that do I/O must do so asynchronously, and must return a “promise” for the result." is said, then no explanation is given on how to implement a function that does it (all the examples assume an existing library returns kj promises and since that's custom, no such libraries exist). For serialization, how do I build a list of integers? The "Lists" section on the serialization page doesn't explain this at all, it only mentions the init method for lists of lists or lists of blobs. The gRPC docs on the other hand are quite good, with numerous easy to read tutorials and examples.
- Built in support for threading. An event loop is limited to the throughput of a single core and there'll always be a point at which you need either threading or multiple processes. Since I don't want to have a bunch of different copies of a particular thing in memory, it works better for me to have a bunch of threads with access to shared memory. Again, the docs say "While multiple threads are allowed, each thread must have its own event loop." but don't document how to do it (and particularly, do it while sharing a port). Plus, there's no suggestions on how to return a promise that's fulfilled by another thread (this was and still is a big problem for me, I still haven't figured it out). gRPC makes this dead simple by just using threads (or a queue system I haven't looked into much yet).
- No kj. I might be more okay with it if the documentation was better but at the moment it's near-nonexistent and wandering around the code base for ~30 mins just to figure out how to construct and return a promise was pretty painful.
Really, you can sum up most of my problems in one word: Documentation.
Oh and as you said, language support. The Python driver is really important to me yet seems unfinished.
Still wondering which one is better, although really, it probably doesn't matter nearly as much as the rest of the code I write. For now I just want it to be easy.
Totally agree that the grpc docs are much better. It took me quite a while to figure out how to just take a message in memory and get it out (all the examples I found are about reading from a file descriptor).
Also agree about the threading model suiting my purposes more, since I don't need any synchronization (I've got reader/writer locks for that).
(It is actually possible to write a multithreaded server, but it needs to be easier and with better documentation...)
We'd be really happy if you just provided many little code-examples that can be put together =))
I've not used Cap'nProto yet, but want to build a ML endpoint and wish to have Python support, so I'm also interested in that
What's the message size?
[1] https://performance-dot-grpc-testing.appspot.com/explore?das...
It looks like go's non-compacting/non-generational gc is at play here. Even csharp is faster than go on that benchmark.
The one you posted has no deserialization overhead since its just ping/pong.
[1] https://performance-dot-grpc-testing.appspot.com/explore?das...
edit: fix typo
When you say you used Cap'n'proto, did you mean the RPC and serializer, or just the serializer? I ask because gRPC's serialization is pluggable, and doesn't have a dependency on Protobuf.
https://github.com/theRealWardo/proto-bench
So yes, if your application has to send a lot of structured data around, you will end up paying a non-trivial serialization cost. On the other hand, if your application is sending really simple messages then I highly doubt proto serialization is the performance bottleneck you should focus on.
Will there be any progress on JS web, or can it just not be done at all with HTTP1? Even a subset of features for basic GET/POST would be fine...
EDIT: I enjoyed the GopherCon talk... I don't think the video is up yet but when it is I recommend watching...
Please tell browser devs to provide access to them! They have been part of the HTTP spec for years. Only recently have browsers entertained the idea:
https://bugzilla.mozilla.org/show_bug.cgi?id=1339096
https://bugs.webkit.org/show_bug.cgi?id=168232
https://bugs.chromium.org/p/chromium/issues/detail?id=691599
https://developer.microsoft.com/en-us/microsoft-edge/platfor...
Regarding the complaint about errors, there is already protocol support for structured error handling: https://godoc.org/google.golang.org/grpc/status. https://github.com/grpc/grpc-go/pull/1358 should make this easier to use. In practice, the provided error code set is very good, so give them a try before making things more complex.
Worst case, you can just stuff things in a header or trailer.
If that sounds interesting, I'm hiring engineers, shoot me a note at joholland at Expedia.com
It's designed to be generic for any kind of async operation, and to be "mixed in" with your APIs. There are utilities for waiting for an operation to finish (via polling); in theory it should be possible to use some type of server-push to avoid polling, but I'm not sure if anybody's doing this.
(I'm an Alphabet employee who's working on APIs delivered via gRPC / REST/JSON.)
gRPC was designed from the start for HTTP/2, which comes with some benefits: It's able to work wherever HTTP works (load balancers and proxies), can multiplex calls over a single stream (Thrift on the JVM, where it's most popular, uses a thread per socket), supports cancelation and streaming and so on. gRPC is arguably more opinionated than Thrift here; Thrift is both a serialization format and an RPC mechanism, and for Thrift RPC you can choose between different transports and framings, of which HTTP is just one. gRPC is HTTP-only, and is (at least nominally) serialization-format-agnostic; you could, in principle, use Thrift over gRPC instead of Protocol Buffers, for example. So in this sense, gRPC is more pure and generic than Thrift. (I'm sure Thrift fans might disagree here.)
I haven't actually used Thrift, so maybe someone else can chime in about other reasons gRPC is preferable.
It's not implementing the gRPC interface. I don't think gRPC was open-sourced [2015] early enough for Thrift [2007] to be built to its API.
Disclosure: Google employee who uses Stubby. [but is not on the team]
I disagree that Thrift failed, but GRPC has more momentum right now IMO.
Source (GRPC at current startup, multiple years of thrift)
I haven't tried to use Protobufs or gRPC in a serious way so I can't say if it's better or worse, but I would hope it supports fewer targets because it takes stability and QC more seriously.
REST APIs don't _have_ to be text-based, AFAIK. Why not just send/receive binary?
I can see where self-describing is better when it comes to side issues like debugging, exploration, or the hassle of configuring your build to generate code from the IDL files, though. But if I had to choose only one, those are lower priority to me.
That said, JSON is still slower and more bloated than Protobuf, and you get Protobuf for free when using gRPC.