gRPC-Go Engineering Practices
grpc.io
grpc.io
What is the advantage of gRPC - just more efficient?
That means that you are storing your schema, which also happens to contain interoperability features.
...and the serialization format is more efficient.
It's worth mentioning that gRPC is actually format-agnostic. While I can't say it does the best job of this (gRPC+protobuf works best), there's precedent for using any transport in gRPC. gRPC-Java includes examples of JSON serialization and Thrift serialization.
gRPC is just the transport protocol (its self built on HTTP/2). That said, I'd highly recommend protobufs, they're great. Define your schema in an agnostic way such that it can be easily browsed and compiled to multiple languages.
(For those curious, IIRC it is now been handed off to a subgroup of the Linux Foundation, called the Cloud Native Computing Foundation. The change does not appear to be well advertised on their website, which contains no reference to it.)
Note that this is an open-source project, and the website itself is on GitHub Pages, so feel free to send us pull requests. Sometimes the reason behind things not changing is not a huge conspiracy, but no one actually having spare time to do it. :)
I generally wouldn't file a PR on someone's copyright line or similar legalese, I don't know what the impact would be for any given organization or how they need to format it.
[1] https://github.com/grpc/grpc/blob/master/CONTRIBUTING.md
Now on the wire, you simply sent field numbers (e.g. the numbers that you have to manually specify in monotonically increasing way, and never reuse smaller number), and either simple type (int, float) or symbol reference to another.
But you never send the actual schema... Now some specific storage formats, might have a duplicate version of that schema, but this may be just part of their design.
So the nice thing, is that if you follow some simple rules, like: never reuse previous number, be careful when changing types of existing fields (there are only so and so ways it can go), then you can upgrade independently your services, and the data being pushed.
For example if you've added a new field (and new number), then your existing code (that's still compiled with the older schema), might just ignore it, since, even that it's there, there is no endpoint (e.g. method call to read it) for it to be accessed.
Off course, it's not so simple. After all, there are tricks, should I simply pass through fields that I do not understand to other services, or should I filter them? I certainly don't know.
I wish protos are used more and more, and ways to put them into various databases (like mysql, postgres, sqlite, etc.) is done. Then also formats that store such data in more optimal way, by rearranging such that compression can be gained, and later faster retrieval, etc.
Also "just more efficient" is a funny way to characterize the performance difference between just data bytes vs data + structure bytes (read: the gap is large). You gain in transmission and you gain during deserialization / parsing.
Here is an example.
{"My key":"my value"} has n=21 characters. When you parse you must scan the whole 21 characters O(n), every time, just to read the thing.
If you instead store this in fixed size bytes, where you have some fixed # of bytes that tell you "my value starts at address 0x43", then you can skip to just the values you care about. You don't need brackets or quotes. And you can use other nifty tricks to compress the binary representation further for savings on the wire.
Something I've read from others is that gRPC is not the easiest thing to use directly in SPA. If you have services exposed to a web front end and mobile, double exposing a REST-ish API + gRPC might not be worth it. This is a problem with which I'm currently struggling.
https://auth0.com/blog/beating-json-performance-with-protobu...
It will be huge when that's a reality and we start building all of our APIs around protobuf.
REST is the language the rest of the web speaks so if providing a consumer service that isn't latency sensitive you should feel at liberty to give a RESTful API.
If you're building internal services, consumed internally, gRPC is a decent way to go.
If say you're providing a logging service to a number of large customers you might want to provide a gRPC service for the performance characteristics.
gRPC also provides you the option to write all your services internally using gRPC and then you can expose a JSON service without having to write it yourself.
Hmm...can you quantify that? I found that while these performance difference you mention exist, they are not actually that large and completely dwarfed by actions later in the chain, particularly if you have generic intermediate representations.
Its never been problematic adding or extending JSON endpoints and its never been a problem using basic gzip compression on the fly either.
And JSON endpoints are a damn sight easier to debug and wireshark and all the rest.
I've spent a lot of time writing fast JSON serialization for various languages including Java etc; its staggering how inefficient most libraries are. But that's not really JSON's fault.
gRPC was born out of Google's Stubby rpc system[0], which is used heavily for communicating between different jobs. If you are going to stand up a lot of different services that are going to talk, it provides a lot of advantages that you don't get with JSON. For large companies that use multiple program languages, this is really nice, as proto3 and GRPC have code generators for a slew of different languages.
There are a lot of other niceties that are useful in gRPC that you don't get with JSON (like streaming data).
But if I were writing the next Uber or Facebook or whatever I'd probably get on gRPC or thrift or another RPC system as soon as we hit a non-trivial number of users.
history is completing full circle. Before JSON (HTTP REST) there was RPC. JSON was a new and shiny thing while the RPC was for the old-fashioned. The JSON took over exactly for all the great advantages over RPC that you listed (and a bunch of others) and despite all the advantages of RPC that people in this thread tout today as the gRPC advantages. I suspect that in 10-15 years some young guys at Google will come up with gJSON, and their arguments for it will look like your comment today.
Basically none of the "advantages" you describe exist when you're in an environment like Google's, and that kind of environment is the sort of environment one uses gRPC (or Finagle, or others) in. IOW, if you can use Wireshark successfully, you have an argument for using JSON. Just keep in mind many people are not in such an environment.
https://github.com/grpc/grpc-java/blob/master/examples/src/m...
message MyThing { oneof sum_type { TypeOne type_one = 1; TypeTwo type_two = 2; } }
The lack of inheritance seems awkward at first, but with oneof it isn't much of a blocker. The APIs for this aren't always great -- Go's in particular feel kind of awkward (IMO). Java's are nice -- it's a separate enum you can switch over.
An example from a Go project of mine:
switch req.StartAt.(type) {
case *pb.GetLogsRequest_Timestamp:
t, err := types.TimestampFromProto(req.GetTimestamp())
if err != nil {
return nil, errors.Errorf("Bad timestamp: %v", req.GetTimestamp())
}
filter.Timestamp = t
case *pb.GetLogsRequest_Offset:
filter.StartOffset = req.GetOffset()
case *pb.GetLogsRequest_Position_:
switch req.GetPosition() {
case pb.GetLogsRequest_LATEST:
filter.Position = LATEST
case pb.GetLogsRequest_EARLIEST:
filter.Position = EARLIEST
}
case nil:
default:
return nil, errors.Errorf("Unknown GetLogsRequest.StartAt type.")
}
}One of the biggest advantage imo is the contract between the client and the server, both are always in sync about what to send / receive.
I've seen many times things break because x,y,x added a field or change a type that the server / client couldn't understand.
Nothing about the REST architectural style prohibits a resource (or, rather, a particular resource representation) from being a stream.
Practical discussion: the biggest advantage of using REST is that you had a client readily available in pretty much every language. Most of those clients haven't been updated to deal with the HTTP/2 machinery necessary to drive streams the way gRPC does.
HTTP has supported streaming data for years, e.g. Server-Sent Events[0]. The short version is that you GET a URL, and the body keeps on arriving forever. I'm pretty sure that we had this sort of HTTP streaming back in the late 90s …
I'll grant that many JSON-oriented pseudo-REST clients probably don't support it, though.
Chunked encoding allows to do something like what gRPC does, but it's my understanding (I might be wrong) that the data needs to be base64 encoded, whereas HTTP/2 supports raw byte arrays to be sent.
If you mean binary data would need to be Base64-encoded with server-sent events, that's true. SSE uses text/event-stream as the content type, which is expected to be text. It would be interesting if someone defined a content type for streaming binary events.
Speaking from my own experience in Java, especially testing my gRPC/HTTP bridge[1], streaming is never overly easy with HTTP clients. You're usually just handed a byte stream to handle yourself. Streaming to the server from the client has been another really interesting exercise to get right.
That's part of what's nice about gRPC (aside from the RPC model and protobuf serialization). If you want unary semantics, done. If you want streaming, it's just a keyword away, and it works across all languages.
My use is not at all complex but I can't imagine that more involved streaming would be any tougher.
[1] https://github.com/chrissnell/gopherwx/blob/master/storage_g...
[2] https://github.com/chrissnell/grpc-weather-bar
Edit: here's a simple example protobuf definition for streaming to a client: https://github.com/chrissnell/gopherwx/blob/master/protobuf/...
The http2 and grpc focus was just to get the alpha release out there.
Good luck implementing a half-decent HTTP/2 client or server library if your language doesn't support it already. This is easily a 6 months job.
The other issue with HTTP/2 is that most client libraries require TLS to work. Development environments become harder to setup. It becomes harder to sniff the traffic during debugging. In client/server scenarios where both are on the same machine this also creates unnecessary overhead.
First challenge is, that you have to keep some kind of reference on what proto are you using within your project. What Google (and some others) do, is that they a) put all the proto files in a separate repo [0] and then generate them for each language separately (python [1], ...). This way you can use whichever proto file you need within your project, however you have to load more libraries than you need to. To be honest, it only makes slight difference when deploying, so not too bad.
The second challenge is, that you have to generate the result files every time you make some kind of change. If you have a lot of proto files, then it may take some time to generate them and there are very little tools available to help you. Google open-sourced Artman [2] although it's more focused on APIs than managing shared protos.
The massive advantage is that, because proto files are self-explanatory and if you put enough information in them can function as direct documentation of your API's interface, you don't need to fish out the requirements in the project or in the documentation but rather just directly read the proto file itself. But this does depend on the developers, to make it as consistent as possible, which is not always the case [3].
[0] https://github.com/googleapis/googleapis
[1] https://pypi.python.org/pypi/googleapis-common-protos
Only difference is we publish the resulting code for all languages into a single artifact repo, which other projects use as a dependency.
We also use a lot of inheritance of non-default types (timestamps, errors, ...) so it's important to make sure that we don't break anything for others.
NoSQL is an interesting parallel - Google publish map reduce and bigtable and amazon publish some influential papers and suddenly everyone is using NoSQL in order to be "web scale". Then it turns out that Google themselves were doing sql web scale and spanner and all that.
There's a risk that gRPC is the same? In chat yesterday ex-googlers said that Google was increasingly moving over to flatbuffer...
Personally I have an aversion for tools with generators. Harks back to the damage CORBA did me I guess... I also have a preference for plaintext eg JSON - so much easier to debug.
Oh well. Guess we're in the fashion business ... ;)
>seniorsassycat: I don't understand why AWS released Go support instead of binary support and I don't understand why they chose to rely on go's net/rpc [...] which encodes objects using Go's special [gobs] binary format
gRPC is a protocol and set of libraries for cross-language rpc based on protobuffs. Also doing a lot of codegen for you, like generating clients.
Seems to generate a REST proxy server side.
The only issue I have with full HTTPS + JSON is CRIME. I'm not sure how that works.
I've actually only used the gRPC JSON gateway mentioned in the other replies so I'm not sure how it compares, but it looks interesting.
I think they're pretty similar and you can't lose either way. Facebook's support of Thrift and Google's of gRPC make both decent options.
One thing I will say about gRPC is that it plays nice with Google's build system (Bazel) and some Google APIs now have first-class gRPC support. If you choose thrift in your stack you'll have to call APIs using JSON or support gRPC anyway if you want to use them for those API calls...so gRPC might be an attractive choice. Furthermore gRPC's go interop is also excellent if you happen to be a fan of golang.