Using Protobuf instead of JSON to communicate with a front end
blog.wearewizards.io
blog.wearewizards.io
We rewrote our app eventually to send protos in JSON format to the app, while just letting our backends still pass around native protos, it worked a lot better.
And looking forward, JavaScript parsing of protobuf binary format has gotten a lot faster, thanks in large part to newer JavaScript technologies like TypedArray. Ideally JSPB would be deprecated as a wire format in favor of fast JavaScript parsing of binary protobufs, but this would of course be contingent on the performance being acceptable.
Finally, JSON is becoming a first-class citizen in proto3, so protobuf vs. JSON will no longer be an either/or, it can be a both/and. https://developers.google.com/protocol-buffers/docs/proto3#j...
JSON is just more popular as a serialization format. It doesn't matter what what programming language or OS I am on, there is almost always a built-in library that de/serialize JSONs at reasonable speed. To send the JSON objects around from one service to another, I can just gzip the string if it's big, or just plain UTF-8 string if it's not.
ProtoBuf has to provide more values for people like me to switch. I would rather try out Apache Avro first as a replacement for what I am doing right now.
With a plain JSON-based API, you copy and paste field names out of sample code or the documentation. If you spell a field name wrong, there will be no error on the client. If you're lucky, the server might error out because it didn't recognize the property name, but it also might not. If you send an integer when the server was expecting a string, the server might automatically convert or it might not.
With protobuf, the schema is explicit in a .proto file. That means that the client library can tell you, at the precise moment that you say msg.misspledFieldName, that the field name doesn't exist. Or if you try to put an integer in there instead of a string, it can tell you about that too. Basically it makes for a tighter feedback loop, which is almost always better.
In statically-typed languages like C++ or Java, the schema can be used to generate static types too, so it's actually a compile-time error when you misspell a field name.
> It doesn't matter what what programming language or OS I am on, there is almost always a built-in library that de/serialize JSONs at reasonable speed.
Yep, that's one reason that proto3 will support JSON as a first-class citizen: https://developers.google.com/protocol-buffers/docs/proto3#j...
:-)
Protocol Buffers were designed from the ground up for RPC, and as a result are far simpler and more convenient to use than XML. Seriously, nobody who uses Protocol Buffers compares them to XML, because it's not even a comparison.
https://developers.google.com/protocol-buffers/docs/overview...
Haberman also mentioned the schema benefits.
All that said, I'm using JSON for my current startup. I view them as optimizing for different parts of the product's lifecycle: JSON lets you quickly adapt the protocol and switch out different languages for different services when you're figuring out what product to build, while Protobuf saves you money when you're trying to scale it. I'm also pretty intrigued by Cap'n Proto as a high-performance serialization format, since it fixes a lot of the problems we faced using protobufs at scale at Google, but its language support just isn't up to protobuf/JSON yet, and the protocol is quite complicated.
Coming from a larger startup, I've also experienced the pains of trying to maintain JSON objects between different services. Protobufs have some quirks, but I think its a great solution to get behind at any stage.
The proof of concept would be to send an array of strings as bytes in a TypedArray, deserialize it to an array of JavaScript strings using JavaScript (not native code), and show that this is about as fast as doing the same thing using JSON.parse(). It seems likely that JSON.parse() will have an easier time creating all those JavaScript strings and other objects at once from native code.
https://developer.mozilla.org/en-US/docs/Web/API/TextDecoder
This will allow you to switch between JSON and protobuf binary on the wire easily, while using official protobuf client libraries. So you can choose easily whether you care more about size/speed efficiency or wire readability. Best of both worlds!
I work on the protobuf team at Google and would be happy to answer any questions.
http://hperadin.github.io/jvm-serializers-report/report.html
One unavoidable issue is that, unlike JSON, protobuf serializers have to do two passes over the message tree, because in protobuf binary format all submessages are prefixed by their length. The first pass just calculates lengths, while the second performs the actual serialization. This could potentially slow down serialization compared to JSON, especially for message trees with lots of nodes/depth.
There are a lot of practical challenges though -- to provide a useful API you'd have to reverse the string at the end, since socket APIs don't generally provide streaming "WriteReverse" functions, and for good reason, because it would force them to buffer arbitrary amounts of data.
So we'd have to reverse the entire string at the end. The question is whether this would be cheaper than doing a second pass over the message tree. And also keep in mind that you would need the first pass to decode everything -- including UTF-8 data -- in reverse. But since the UTF-8 APIs for Java strings probably don't support this, you'd probably have to encode it, then reverse it to put it in the encoding buffer. That way when it gets reversed again at the end, it would be proper UTF-8.
At the end of the day, this probably wouldn't end up faster than what we do now. But can't say for sure without trying!
Why is a mapping to camel case necessary? I imagine it creates the potential for collisions, no?
A benefit of this decision is that if you create JSON manually in JavaScript (ie. without a protobuf library) your JSON objects will match JavaScript conventions.
[1] https://microsoft.github.io/bond/manual/bond_cpp.html#simple...
[2] https://microsoft.github.io/bond/manual/bond_cs.html#json
On top of which, you've defined built-in default values for empty fields; this means that, without warning, an accidentally missing field will inject bad data into any consumer that doesn't carefully check for the existence of all required fields.
These are basically killer issues for us; we're not going to adopt an "update" that requires us to write JSON-style "hey, does this field exist?" code everywhere.
Also, how do you deal with the bytes type?
[1] https://code.google.com/p/protobuf-json/ https://github.com/benhodgson/protobuf-to-dict
edit: it's exactly what you would do if you wanted to pass any binary data as json over the wire, regardless of whether you're using protobufs. you'd just get it "for free" (meaning you wouldn't have to write the boilerplate, not that you don't have to en/decode).
I recently became the maintainer of the Dart protobuf library which supports both JSON and binary format [1], [2]. However, the JSON format isn't necessarily compatible with other protobuf libraries you've seen.
[1] https://github.com/dart-lang/dart-protobuf [2] https://github.com/dart-lang/dart-protoc-plugin
The most common method is to use an array indexed by the field number. I've seen protobufs with hundreds of fields so that's hundreds of nulls as the string "null".
The alternative is to have JSON objects with attributes named after the protobufs field name. This isn't without warts either and seems to be less prevelant in my experience.
Another problem is JavaScript doesn't support all the data types you can get in protobufs, most notably int64s.
Protobufs are relatively space efficient (eg variable width int types). JSON encoded protobufs much less so.
Perhaps the rise of browser support for raw binary data will make this less awful.
Many consider it a virtue to use the same code on the client and server. It explains things like this and GWT. Personally opi think this is horribly misguided and a fools errand. You want to decouple your client and server as much as possible (IMHO).
Disclaimer: I work for Google
What you are describing here is known as the "JSPB" wire format. This is a serialization that is only ever used for JavaScript, and only used there because, historically, parsing binary protobufs in JavaScript was too slow. With TypedArray and other JavaScript enhancements, this is changing. Ideally, JSPB wire format would be phased out completely.
> The alternative is to have JSON objects with attributes named after the protobufs field name. This isn't without warts either and seems to be less prevelant in my experience.
It's about to become a lot more prevalent with proto3, which features first-class JSON support. See: https://developers.google.com/protocol-buffers/docs/proto3#j...
Disclaimer: I work on the protobuf team.
It would be nice if we had a standardized JSPB wire format that used tag numbers, rather than the various unofficial implementations we have now.
if you are using a statically typed language, binary formats like Protobuf are a big win, but if you are going to have the dynamic language overheard that comes with JS, there isn't much gain to be had from binary formats.proto3 supports both binary protobuf encoding and JSON natively, so you can switch between them as desired. https://developers.google.com/protocol-buffers/docs/proto3
proto3 is currently in alpha, but we are working to bring it closer to release (I work on the protobuf team at Google).
proto3 is especially designed to be paired with gRPC, which is also in alpha but also going for wide language support: http://www.grpc.io/
I have it on a todo list to port an SBE parser to ScalaJS. ScalaJS already backs java ByteBuffers with javascript TypedArrays. That should be really fast, the same stuff that is being worked on for making asm.js fast will also make the Cap'n Proto / SBE approach fast, so I think this has the most promise of bringing really high-performance data transfer capabilities to the browser.
"Don't use protobufs if you don't have to".
Protobufs can be much faster, and provide a strict schema, but it comes at the price of higher maintenance costs. JSON is much simpler, easier to implement, and MUCH easier to debug. If your GPB looks like it's building properly, but fails to parse, it's a huge pain to try and decode/debug the binary. You'll wish you could just print the JSON string.
If you need the speed and schema, then GPBs are great. In our case, we got a huge speed boost just by avoiding string building/parsing inherent in JSON.
Are the maintenance costs related to debugging unparsable messages? We've almost never had an issue there, so maybe we've just been lucky?
Protobufs can be encoded as JSON and as text, so there are some ways to address the readability I guess.
As for Protobuf, I tried using it in a number of places, but found it to be very inflexible (schema!) and hard to debug in case of problems.
I'm looking forward to seeing protobufs in Rust as a macro. It should be possible; there's an entire regular expression compiler for Rust as a compile-time macro, which is a useful optimization.
No, it isn't true, but regardless of what format you use, there will always be someone who's not happy. Actually, I think that applies to everything in life.
1. Dictionary keys are repeated when you have an array of similar objects.
2. Non-text data. JSON can't natively represent binary data, forcing people to use things like base64 for binary and base10 for numbers.It transmits type and field names. Depending on how complex your data is those strings could be a large part of the data.
{ "person": { "age": 30, "shoesize": 10 } }
The above is what, 4-5 bytes of protobuf? I'm not sure what the gzipped-json data is but likely a lot more. If you were to send a list of 100 such person objects, the difference would be smaller.
Assuming the integer fields are regular varint types (and not the "fixed" integer encoding), and assuming the tag numbers were all under 16, then this would be a six-byte protobuf.
IMHO trying to make everything look like an object was a mistake, and newer RPC frameworks like gRPC, Thrift, and JSON-over-HTTP are much easier to use than the late-90s frameworks like RMI, CORBA, and DCOM. Sometimes you don't want abstraction, because it abstracts away details you absolutely need to think about.
842 words including code.
Average adult reading speed: 300 words/minute.
Does not compute.
Only reading the text itself takes indeed less than 5 minutes, not sure which approach people prefer.
However, I do feel there is a strange swaying back to binary (Protobuf/HTTP/2/etc). Developers are trying to wedge it in now in places it may cause more problems because it is more efficient in performance but not in use or implementation. Plus, like mentioned in this thread, you can compress JSON to be very small to send over the wire which makes the compactness of it a non-issue in non real-time cases. Going binary just to go binary is more trouble than it is worth in most cases.
- Binary over keyed plain text (JSON) is harder to generically parse objects i.e. dictionaries/lists for just a few fields/keys.
- Binary over JSON also seems to lock down messaging more, people have more work to change binary explicit messages because of offset issues and client/server tools must be in sync rather than just adding a new key that can be pulled as needed.
- Third party implementation and parsing of JSON/XML is more forgiving making version upgrades and changes easier to do. This is especially apparent on projects that are taken over by other developers.
- The language/platform on the backend leaks into the messaging. For instance Protobuf only runs on js/python currently and has various versions. The best messaging is independent of the platform and versioning is easier.
I would bet binary formats end up causing more bugs over keyed/plaintext (JSON/XML and possibly compressed) though I have nothing to back that up by except my own experience largely in game development where networking state is almost always binary, for server/data I wouldn't use it unless it needs to be real-time.
That being said Protobuf is awesome and I hope developers are using it where it is best suited and that developers don't start obfuscating messaging for performance where it doesn't really need to be, better to be simple unless you need to make it more complex at every level.
See https://groups.google.com/forum/#!topic/grpc-io/5Ic8MKgltwY
However, you still need the FAQ to figure out that browser transport isn't supported
https://groups.google.com/forum/#!topic/protobuf/eNAZlnPKVW4
My understanding of ASN.1 is that it has no affordance for forwards- and backwards-compatibility, which is critical in distributed systems where the components are constantly changing.
...
OK, I looked into this again (something I do once every few years when someone points it out).
ASN.1 _by default_ has no extensibility, but you can use tags, as I see you have done in your example. This should not be an option. Everything should be extensible by default, because people are very bad at predicting whether they will need to extend something later.
The bigger problem with ASN.1, though, is that it is way over-complicated. It has way too many primitive types. It has options that are not needed. The encoding, even though it is binary, is much larger than protocol buffers'. The definition syntax looks nothing like modern programming languages. And worse of all, it's very hard to find good ASN.1 documentation on the web.
It is also hard to draw a fair comparison without identifying a particular implementation of ASN.1 to compare against. Most implementations I've seen are rudimentary at best. They might generate some basic code, but they don't offer things like descriptors and reflection.
So yeah. Basically, Protocol Buffers is a simpler, cleaner, smaller, faster, more robust, and easier-to-understand ASN.1.
Ugh was that only 5 years ago?
Ah Ok whew, so the title was wrong or designed for click bait.