A detailed comparison of REST and gRPC
kreya.app
kreya.app
The whole point of REST is that it's discoverable. Now web standards didn't quite manage to standardise HATEOAS, so sadly full machine discovery is unlikely via an API, but you can build quite adaptable clients that can respond to changing APIs well, going as far as optimising client usage by changing API responses without redeploying clients. That may or may not be something you want in an API, but it's worth considering, because it's not going to happen with gRPC.
GraphQL, like REST, is about the objects and relationships, and lends itself well to building highly capable client-side caching layers when you introduce the Relay patterns to it – particularly globally unique IDs. Given GraphQL's well defined schema it's relatively easy to build generic caching mechanisms. Again, this may or may not be something you want in an API.
gRPC doesn't really allow for any of this, if you want it you've got to invent it all yourself. But that might be ok! Server to server calls rarely need a big cache to work around poor networking. gRPC does however offer considerably more control over streaming behaviour, more standardised error handling than GraphQL, and more.
There are quite a few factual errors in this post, like REST not being schema based (that's up to the implementer), no streaming in REST (not necessarily true).
When deciding on an API technology the first questions must be "who is the consumer", and "what are their constraints". These will often lead to just one or two of the options here, then you can drill down into the details like tooling, API design, and so on.
Where did it say that? I got the opposite impression:
> Handling large data sizes with REST APIs, such as file uploads, is rather straight forward. The received file can be treated as a stream, using very little memory.
Re: streams, I've found that an opaque binary stream in REST is okay, but the moment I want to stream messages, my life is hell. I keep running into this, and it's such a pain. If I have to stick with REST, I end up inventing a bespoke RPC protocol to allow for sending errors as part of the stream, either as part of the message or as a trailer header (and trailers have their own issues like no browser support). And then the semantics of HTTP status codes have to change.
The reason I often stick with REST is that infrastructure supports HTTP/1 much better than h2/gRPC (such as L7 load balancers, caching, etc.). JSON RPC is not always a good answer either as caching POSTs is typically not well-supported and using GET with JSON RPC is discouraged and fraught with issues.
I would definitely prefer to just use gRPC but the poor middlebox support sometimes makes it cost-prohibitive.
What a mess.
You're right that files can be streamed and it does say that, but additionally there's no reason why messages can't be streamed down an open connection.
Out of interest, as I hadn't considered middlebox support, what sort of things do you find becomes problematic? Doesn't TLS mitigate that? Or are you thinking of niche environments like companies that require intercepting all traffic on their networks? I realise those use-cases do exist, but they're pretty uncommon now as more people realise they're terrible security practice.
Is the client always going to use a library to interact, is connecting without any code/schema published by the API provider a core requirement? Are clients going to be in languages with good protobuf/GraphQL support? Some don't have this. Is the code that talks to the API also owned by the API provider? Public vs private. Support lifecycles.
These factors all play into whether the type safety could even be utilised by clients. And it's not like REST doesn't have this – OpenAPI and Swagger can get a lot of the benefit with fairly minimal work. Both are very common.
They do differ fundamentally in that type safety is an afterthought with REST but is built in with some of the alternatives. That can have it's benefits too of course, it's easier to integrate with and more broadly supported in part because it doesn't concern itself with that.
OpenAPI can help but keeping your schema in sync with reality can be a pain depending on what libraries you have available. In the best case it really is minimal work, but if you don't have good tooling for whatever web framework you're using it can be a bit of a pain. In my experience it often requires more manual effort to maintain & more risk of mistakes causing the schema to be inaccurate.
* manual:
* schema is maintained manually, independent of the implementation
* usually an afterthought and used mainly for documentation
* implementation-first:
* schema is autogenerated from the code
* used as a reference, possibly to generate clients, and often to run tests
* probably the most common approach, supported by many frameworks
* schema-first:
* schema is maintained manually, with client and server code generated from it
* very rare, but the most correct approachUnless your models are very simple, the best approach is to use three separate layers of model definitions:
* API models for serde and conversion of external requests
* business logic models that carry the actual internal functionality
* (optionally) ORM models to convert the data for persistence to a RDBMSThe thing I’ve noticed when stepping into a codebase where this problem has been allowed to occur is the lack of layers of abstraction. Having those different models built up from the start allows for an application to shift along with the needs of the product. Having a single layer, with the endpoints talking literally directly to the ORM models, almost inevitably leads to calcification, spaghettification, and disastrous performance.
How is using, for example, JSON Schema to define your types any better or worse than a "proto" file that requires a compiler, a parser, and a client and server library?
There's many tradeoffs of course, it's not like dealing with proto files is painless either.
JSON has a limited set of types. There is object, array, string, number, boolean, and null. That's it.
JSON Schema adds to that by providing definitions of types that build on those basic JSON types.
Any good examples on how to do this for non-trivial API changes?
- Error codes are well defined (vs "should I return 200 OK and error message as a JSON or return 40x HTTP code?")
- no semantics ambiguity ("should I use POST for query with parameters that don't fit URL query param, or POST is only for modifying?")
- API upgrades compatibility out of the box (protobuf fields identified by numbers, not names)
Not to mention cross-platform support and autogenerating code.
I use it in multiple Flutter+Go apps, with gRPC Web transparently baked in, and it just works.
Once I had to implement chunked file upload in the app and, used to multiform upload madness, was scared even to start. But without all that legacy HTTP crap, implementing upload took like 10 mins. It was so easy and clean, I almost didn't believe that such a dreadful thing as "file upload" could be so easy. (years of fighting HTTP legacy).
Compared to the "traditional" workflow with REST/JSON the downside for me is, of course, the fact that now you can't help but care about API. With web frameworks the serialization/deserealization into app objects happens automagically, so you throw JSON objects left and right, which is nice until you realize how much CPU/Memory is being wasted for no reason.
Also, check out drop-in replacements for cases where you don't need full functionality of gRPC:
- Twirp (Twitch light version of gRPC, with optional JSON encoding, HTTP1 support and without streaming) - https://github.com/twitchtv/twirp
- Connect - "Better gRPC" https://connect.build/docs/introduction/
I'm truly curious to know how REST or JSON or HTTP is not cross platform
With REST/JSON mostly you write code twice – server implementation and client implementation. If you add new platform (say, iOS), you write new code again. So for each new platform you maintain separate code => not cross platform.
Imagine having API with 100+ endpoints. How do you add another one in gRPC? You edit .proto file and run codegenerator, which creates stub code for clients and server. You just need to connect that code to the rest of your app (say, talk to database here or display the response on the screen here). You don't write actual code for the network part.
With traditional HTTP/REST/JSON approach you would have to write handlers yourself for each codebase separately. And there is no way to make sure they're in sync – only on organizational level (by static checks, four-eyes policies and such). I know it doesn't sound too big of a deal because most people used to it, but gRPC gives a different experience. Hence my initial comment that it's so much easier to work with.
That's not just REST/JSON, of course. That's the general feeling I have from web-ecosystem. Recently I had to work with a simple HTTP form with one field and a single checkbox that is shown conditionally. Hours of debugging revealed that you can't just POST unchecked checkbox [1]. There is no difference between "unchecked checkbox" and "no checkbox". Instead you have to resort to hacks with hidden input field. It's just all feel hackish and you constantly question yourself – am I doing something wrong or it's just this whole stack is a set of hacks on top of hacks?
Same feeling with REST/JSON. Once your API grows past simple CRUD you start caring about optional values and error codes. Ambiguity seems fine until project grows and more people join and introduce inconsistency: one call returns 200 OK for error, other returns 5xx/4xx (cost of choice). Now you have to enforce rules with static checkers and other tools. You bring more tools just to keep the API sane, and it all, again, feels hackish and ill-suited. I don't have this feeling with gRPC - it feels like perfectly designed for APIs.
[1] https://stackoverflow.com/questions/1809494/post-unchecked-h...
What is the ubiquitous utility for interacting with gRPC? We have curl for REST. What is openAPI of gRPC?
This is technically true, but part of the "grpc philosophy", if you will, is to not make breaking changes, and many of the design decisions of protobufs nudge you toward this. If you follow this philosophy, change management of your API will be easier.
For example, all scalar values have a zero value which is not serialized over the wire. This means it is not possible to tell if a value was set to the zero value, or omitted entirely. On the surface this might seem weird or annoying, but it helps keep API changes backwards _and forwards_ compatible (sometimes this is called "wire compatible", meaning that the serialized message can work with old and new clients).
Of course you still can make wire-incompatible changes, or you can make wire compatible changes that still break your application somehow, but getting in the habit of making only wire-compatible changes is the first step toward long term non-breaking APIs.
GraphQL, by contrast, lets you be more expressive in your schema, like declaring (non)nullability, but this quickly leads to some pretty awkward situations... have you ever tried adding a required (non-nullable) field to an input?
The proto file. Grab that , use protoc to generate bindings for your language and off you go....
grpcurl[1] combined with gRPC server reflection[2]. The schema is compiled into the server as an encoded proto which is exposed via server reflection, which grpcurl reads to send correctly encoded requests.
[1] https://github.com/fullstorydev/grpcurl [2] https://github.com/grpc/grpc/blob/master/doc/server-reflecti...
Kreya, for example, haha (check the original link of this post).
There are many, actually, including curl-like tools. But I almost never use them. Perhaps it's because my typical workflow involves working with both server and app in a monorepo, so when I change proto file, I regenerate both client and server.
Just once I had to debug the actual content being sent via gRPC (actually I was interested in the message sizes) and Wireshark did job perfectly.
Anyone who returns a 200 OK on error is, of course, in a state of sin.
You’re welcome to return a 4xx + JSON, if that is congruent with the client’s Accept header.
Yes, I’m aware that GraphQL uses 200 OK for errors. That is only one of the reasons that it GraphQL is unfit for purpose. Frankly it is embarrassing that it has become popular. It’s like a clown car at a funeral.
Does JSONRPC have tooling for generating clients or API documentation?
You mean typed client with runtime assertions, cross language spec etc, right? That's what I mean, people should be discussing jsonrpc with joi vs zod vs jsonschema, their runtime overhead, cross language support, codegen support etc.
Short answer is whatever currently we have for json will work. There is no point in complicating things.
I don't like REST because it's basically become a specific url structure with semantic http methods (e.g. POST for create). Practically nobody bothers with the HATEOS stuff. Things start breaking down if you have a complex resource hierarchy and custom methods.
I've personally just settled on rpc using json. Any consumer of the api REST or otherwise is going to check the docs for the correct endpoint, at which point there's little to no difference between rpc and rest.
It’s not at all obviously wrong on the surface, there is some setup involved in the sense you’re bringing at a minimum one new tool into your build process but you are picking up a LOT in exchange for that trade off from prebuilt client libraries down to an extremely efficient wire format.
RPC using JSON sounds like the worst of both worlds to me. You ended up trading away a huge amounts of the benefits for things like code generation and general efficiency for a pretty reasonably one time setup cost.
Maybe that’s different for the language of your choice I don’t know…
The concept of resources maps well to URLs, which name things, ie a URL is a noun not a verb. The REST style is about two endpoints transferring their current state of a resource, using a chosen representation, whether JSON, XML, JPG, HTML etc.
gRPC is a remote procedure protocol that uses protobufs as a TLV binary encoding for serializing marshalled arguments and responses. It requires clients and servers to have compiled stubs created to implement the two endpoints. Like most RPC, it is brittle and its abstraction as a procedure call leaks when networks fail.
GraphQL is a query protocol for retrieving things, often with associated (ie foreign key) relations. The primary use is defined in the "QL" of the name.
The benefit of HTTP and the use of limited verbs and expressive nouns (via URLs) is that HTTP defines the operation, expected idempotency, and expected resource "state" after an HTTP request/response has occurred. It has explicit constructs for caching, allowing middleware to optimize responses.
There's nothing in HTTP that requires JSON, the choice of media type is negotiable between the client and server. The same server URL can serve JSON, XML, protobufs, or any other format.
gRPC is yet another attempt to extend the function call of imperative languages to the network. It is the latest in a long line of attempts, Java had RMI, there was SOAP and XML-RPC. Before that there was CORBA and before that there was ONC-RPC. They all suffer from the lack of discoverability, the tight binding to language implementations, and the limitations of the imperative languages that they are written in.
They all end up failing because of the brittle relationship between client and server, the underlying encoding (XDR, IDL, Java etc etc) of the marshalling of arguments and responses is essentially irrelevant.
For example https://jsonapi.org/format/ focuses on traversal of relations.
If you are doing RPC-on-HTTP then it is a bad idea.
I'm fine with the idea that URIs reference resources and HTTP actions are verbs.
I also take issue with the claim that gRPC and protocol buf is brittle; the protocol was explicitly designed to allow older servers to process messages from newer clients (and vice versa) to the best of their ability. More importantly: there are enormous production servers that see waves of server updates (and their associated clients are also getting updated) sending petabytes to exabytes to each other every day; in that sense, it's clearly not brittle or far more users would have a negative experience.
I just spent several weeks onboarding a new system that is based around REST and JSON schema. JSON schema... like most of the things with JSON and Javascript, feel like they were implemented in a hurry by non-experts who wanted to solve a problem, and made something simple enough that large numbers of users adopted it. Now we're stuck with the "core technology is based on less-than-awesome technology". most folks don't even use schema, and the document databases that receive the JSON blobs just sort of treat them as a dynamically created schema defined by the envelope of all extant messages (see, for example, dynamic mapping in elasticsearch).
(my experience includes: XDR and SunRPC, CORBA, protocol bufs, stubby/grpc, XML, WSDL, SOAP, and many more systems. I am not authoritative, and I have my own strong opinions based on experience. But I have to say, I'd rather work in an grpc/protobuf world than a REST/JSON one. It's much more robust.
What you're talking about is how can you publish a machine readable discoverable API. Just like RPC, there is no "well known" endpoint for getting the API specification.
The RPC IDL is effectively the same as "go read the docs". How is downloading a "formal published IDL" any different to a "formal published OpenAPI specification"?
gRPC just has an entire infrastructure of compilers, parsers and language libraries that generate stub code that you then have to go and "fill in".
OpenAPI is a pretty good standard for defining an HTTP based REST API.
In gRPC the definition is a requirement to use, so at the bare minimum you have the typed structure of requests and responses. There is no such requirement for REST
Not really. Grpc is just sending/receiving messages to/from an address. Other protocols like COM/CORBA tied the address to an object.
Also there is nothing to prevent you from writing a grpc service in a resource/entity oriented style while REST makes expressing non-resources like actions seem a bit awkward.
GraphQL especially is excellent with this, with Apollo federation [1] allowing you to interconnect all the APIs of your complete company into 1 single endpoint and reference each other's data, which is extremely powerful and allows development to move fast if done right.
Developer experience is also much nicer. Instead of repeatedly writing JSON serializers/deserializers on the backed + fronted you get them generated for free. I know there's various open API/json schemas code generators but for me i'd rather use Protobuf's at that point and get the benefits of a strongly typed schema, schema evolution, reduced payload size so faster experience for the user, cheaper bandwidth and cheaper storage if you currently store opaque JSON blobs in the db.
JSON API's are api's in debug mode.
As for the content negotiation for JSON vs protobuf, that's also defined in the HTTP standard.
https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/Ac...
[citation needed]
You're using it via your browser on this very website.
The reason for the performance difference is that a WebSocket is a single statefull TCP socket with a tiny extra binary frame header sending messages back and forth. HTTP is stateless so it creates a new socket for each and every message. HTTP also has a round trip, request/response. The round trip means twice the message frequency as compared to the fire and forget nature of WebSockets.
Or are you using something like trpc?[0]
Another thing I’ve also thought is that Content-Type plus the appropriate headers could easily open up using compressed/efficient binary serialization on the web as well and unlock even further gains.
One thing though is the issues with how websockets complicate scaling horizontally for simple request handling.
I can see how protobuf would be handy if you are just talking only to a database or directly transmitting a binary BLOB, but I rarely do any of that. Most of my messaging are microservices used by the application so I just use JSON.stringify for each message. I know there is overhead to that, but it works well for me. I should probably find a way to pull my application data directly into a binary format and back out without string handling, but I have not figured this out yet. Instead of 7x faster than HTTP my message handling would be 10-11x faster than HTTP.
Well, HTTP 1.1 supports persistent connections and pipelining[1].
Though I suppose if you're a JavaScript in the browser it's easier to just use WebSockets.
[1]: https://developer.mozilla.org/en-US/docs/Web/HTTP/Connection...
Imo, what could have been a decent piece of engineering has been killed by that "you should do like Google does" mentality, with some highly debatable design choices leaking in an unrelated standard (hello defaults).
HTTP is designed for transferring media types. It can transfer the application/octet-stream if you want to transfer "raw" binary.
If want you want to do is stream binary data, there are better protocols than TCP or UDP for the purpose. The use of Websockets is a way to get binary streams through the HTTP "firewall hole" on port 80/443, not because it is more efficient.
The internet and common protocols used to be designed with real world use cases in mind, but now all technology is designed with only one generic consumer user in mind, and everybody else gets left behind.