Reasons to use protocol buffers instead of JSON
blog.codeclimate.com
blog.codeclimate.com
No reason you can't implement schemas over JSON. In fact, you typically implicitly do - what your code is expecting to be present in the data structures deserialized from JSON.
> Backward Compatibility For Free
JSON is unversioned, so you can add and remove fields as you wish.
> Less Boilerplate Code
How much boilerplate is there in parsing JSON? I know in Python, it's:
structure = json.loads(json_string)
Now then, if you want to implement all kinds of type checking and field checking on the front end, you're always welcome to, but allowing "get attribute" exceptions to bubble up and signal a bad data structure have always appealed to me more. I'm writing in Python/Ruby/Javascript to avoid rigid datastructures and boilerplate in the first place most times.[EDIT] And for languages where type safety is in place, the JSON libraries frequently allow you to pre-define the data structure which the JSON will attempt to parse into, giving type safety & a well define schema for very little additional overhead as well.
> Validations and Extensibility
Same as previous comment about type checking, etc.
> Easy Language Interoperability
Even easier: JSON!
And you don't have to learn yet another DSL, and compile those down into lots of boilerplate!
I'm not trying to say that you shouldn't use Protocol Buffers if its a good fit for your software, but this list is a bit anemic on real reasons to use them, particularly for dynamically typed languages.
Right, that's the point -- since in normal use you impose a set of "schema" requirements over all data interchange formats, even schemaless ones, it's a strictly good thing to have that schema explicitly written out. It means the compiler can verify your types are correct and the runtime can verify your messages have all the fields they'll need.
> JSON is unversioned, so you can add and remove fields as you wish.
Sure, but if you do, you have to handle the version-management code at the application level, manually, where it's really easy to make mistakes.
> And for languages where type safety is in place, the JSON libraries frequently allow you to pre-define the data structure which the JSON will attempt to parse into, giving type safety & a well define schema for very little additional overhead as well.
Sure, and if you're going to do that, you might as well use protobufs which is going to be much faster/more lightweight.
If you're looking for fast, you're probably not looking at an interpreted language. Plus, there are faster protocols out there than protobufs, if speed is all you care about.
> It means the compiler can verify your types are correct and the runtime can verify your messages have all the fields they'll need.
Again, the OP was referring to using protobufs in Ruby. Type safety and the compiler are not high priorities when writing in Ruby.
> Sure, but if you do, you have to handle the version-management code at the application level, manually, where it's really easy to make mistakes.
That depends on your application's requirements - for most cases, Python/Ruby/Javascript make it very easy to add or deprecate fields in dictionaries/hashes without having to explicitly worry about versions. Just specifying default values when requesting a key. And since you can ignore hash pairs with even less effort, deprecating portions of an API are simple.
Protobuffers definitely has its place, I'm just not convinced that place is in dynamically typed languages like Python, Ruby or Javascript.
Not quite. Let's say your JSON data contains the following attribute:
"access" : [ "view", "edit", "admin" ],
This field should be represented (in the language) as a set of values from an "access-levels" enumeration.In C#, you'd have the following boilerplate:
[DataMember(Name = "access")]
public HashSet<AccessLevels> Access { get; private set; }
In OCaml, it would be: access : Access.Set.t ;
The simple "json.loads" solution would return a list of strings instead. What's the Python code for turning it into a set of enumeration values, and failing if one of the values does not match ? if 'edit' in structure['access']:
# can edit
In which case it really doesn't matter if there's junk in access. If I really felt the need to validate the structure, I could do so with: if not set(structure['access']).issubset(all_access_privs):
raise ValueError('invalid access types passed in')
but more realistically, I'd rely on the ORM object to validate against the authoritative source - the database - and key off the errors there.It's kind of like saying "language X sucks because it doesn't have an automatic build tool; use language Y instead". You just need a build tool for X, not a new language.
Free backwards compatibility ? No. Numbered fields are a good thing, but they only help in the narrow situation where your "breaking change" consists in adding a new, optional piece of data (a situation that JSON handles as well). New required fields ? New representation of old data ? You'll need to write code to handle these cases anyway.
As for the other points, they are a matter of libraries (things that the Protobuf gems support and the JSON gems don't) instead of protocol --- the OCaml-JSON parser I use certainly has benefits #1 (schemas), #3 (less boilerplate) and #4 (validation) from the article.
There is, of course, the matter of bandwidth. I personally believe there are few cases where it is worth sacrificing human-readability over, especially for HTTP-based APIs, and especially for those that are accessed from a browser.
I would recommend gzipped msgpack as an alternative to JSON if reducing the memory footprint is what you want: encoding JSON as msgpack is trivial by design.
1) Doesn't support Visual Studio 2013.
2) Doesn't support Mac OS X Mavericks.
3) No "nice" support C++11 (i.e. move constructors)
(These can be at least partly solved by running off svn head, but that doesn't seem like a good idea for a product one wants to be stable)With JSON I can be sure there will be many libraries which will work on whatever system I use.
Meanwhile, there are two branches of Protocol Buffers: internal and external. They contain mostly the same code, and there are scripts that automatically merge changes back and forth, but those scripts do require some amount of human supervision. And since the internal users are a lot more demanding than the external ones, most development occurs internally.
This is a pretty crappy way to run an ostensibly open source project, and that's my fault. I probably should have put in the effort to better unify the two branches so that changes could be integrated automatically. It would have been a pretty significant amount of work, though, and I was mostly a one-man team (with a few "20% time" helpers), and most of my time was spent trying to figure out how to convert millions of lines of existing internal application code over from proto1 to proto2 (where proto1 is the original version of protobufs which has never been released publicly).
What I did do was make sure to run the fiddly release process on a semi-regular basis, so that the open source release did not fall behind. Luckily the core code was pretty stable, so it wasn't necessary to do releases all that often. But pushing releases was something I pretty much had to do on my own initiative; management never cared.
So when I eventually moved off protobufs, the replacement team (which management didn't even realize was needed until about six months later) naturally deprioritized open source releases. They've done a couple over the years, but as you've noted, it's rare.
On the "bright" side, judging from the way things were going when I left the company in early 2013, it's unlikely that there has been any significant change in the internal code from which you'd benefit, so maybe it doesn't really matter.
That's a really tiny issue which just requires including another header.
- network bandwidth/latency: smaller RPC consume less
space, are received and responded to faster.
- memory usage: less data is read and processed while
encoding or decoding protobuf.
- time: haven't actually benchmarked this one, but I
assume CPU time spent decoding/encoding will be
smaller since you don't need to go from ASCII to
binary.
Which means, all performance improvements. They come, as usual, at the cost of simplicity and ease of debugging.I believe that's usually spelled "I am making this up".
This is one of those times. Parsing the text "1234" is obviously going to be slower than loading its binary value, and text obviously takes more space than binary.
That said, the internet is full of benchmarks backing this up if you care to Google it.
I'd also guess it would be more expensive during decoding. So let's compare that with some real data:
https://github.com/eishay/jvm-serializers/wiki (old wiki but with some graphs https://code.google.com/p/thrift-protobuf-compare/wiki/Bench...)
it turns out that state of the art JSON serialisers aren't so bad after all, especially the serialisation.
In a civilized discourse, if you disagree with my assumption because you know better, you should provide information as to why I'm wrong. And I'll be delighted to obtain this new knowledge.
Otherwise, what's the point of your intervention?
This seems to me like a key issue, you need to really know beforehand that this won't ever be the case, else you need to make your application polyglot afterwards. A risky bet for any business data service.
Maybe if it's strictly infrastructure glue type internal service. But even then, maybe someone will come along wanting to monitor this thing on the browser.
From Protocol buffer python doc: https://developers.google.com/protocol-buffers/docs/pythontu...
"Required Is Forever You should be very careful about marking fields as required. If at some point you wish to stop writing or sending a required field, it will be problematic to change the field to an optional field – old readers will consider messages without this field to be incomplete and may reject or drop them unintentionally. You should consider writing application-specific custom validation routines for your buffers instead. Some engineers at Google have come to the conclusion that using required does more harm than good; they prefer to use only optional and repeated. However, this view is not universal."
So basically I will be in trouble if I decide to get rid of some fields which are not necessary, but somehow were defined as "required" in the past.
This will potentially result in bloated protobuf definitions that have a bunch of legacy fields.
I will stick to the JSON, thanks.
" old readers will consider messages without this field to be incomplete and may reject or drop them unintentionally."
That means the old readers -- the ones that are expecting required fields, can't accept new messages. That's good! The readers don't know how to read the new messages! The readers need to be updated to a new version before they can correctly start reading new version of the schema.
I've written more about this here:
https://kentonv.github.io/capnproto/faq.html#how_do_i_make_a...
Of course, once something hits production, you really cannot reasonably update them all at the same time. That's why the proto API has:
https://developers.google.com/protocol-buffers/docs/referenc...
bool ParseFromString(const string & data)
Parse a protocol buffer contained in a string.
bool ParsePartialFromString(const string & data)
Like ParseFromString(), but accepts messages that are missing required fields.
Or simply use "optional" instead as others suggested.Also, even with these "bloated" definitions with legacy unused fields, it's probably still smaller than JSON.
[ 4738723874747487387273747838727347383827238734.00 ]
Parsing that universally is a shit.however if schemas scare you (shame on you if they do) then msgpack might be a better choice.
> * You need or want data to be human readable
When things "don't work" don't you always want this feature? Over a long lifetime, this could really reduce your debugging costs. Perhaps protocol buffers has a "human readable mode". If not, it seems like a risk to use it.
With Node, I'd have to see a very good argument for why I should give up all of JSON's great features for the vast majority of services. Unless the data itself needs to be binary, I see no reason why I shouldn't use the easy, standard, well-supported, nice-to-work-with JSON.
The thing it seems we've learned more recently is that the more features and complication you add to a protocol, the more it sucks. At the end of they day your data is composed of primitives, records, and lists, and if your protocol offers to structure things in any other way, it's just creating confusion. This is why JSON beats XML, and Protobufs beats many of its binary predecessors.
Anyone who has used Avro in Java knows that this is not true.
tl;dr Avro JSON != Standard JSON
The OP's assertions that the protocol buffer's are awesome and reduce boilerplate conflict with my experience.
(Answer: Trivial.)