You're Using JSON, Why not MessagePack?
blog.andrewvc.com
blog.andrewvc.com
* It does not support our languages. Is there a Ruby module with no C extension? A C# library? Lua? What about that really cool language coming out next week? What about C?
* Even if it does, why have to deal with someone else's poor API design? JSON has a million parsers for everything, and if by chance you don't like any of them, you can write another one in about two hours.
* It's not human-readable. Enough said.
* It's smaller than JSON, sure, but is it smaller than gzipped JSON? PB's aren't for our purposes. Neither are TNetstrings. This might be, I haven't checked, but I doubt it.
Why is the "with no C extension" part important to you?
This could be alleviated with Ruby FFI on top of a really good C implementation, but FFI is tricky.
This first part was already addressed in another comment, but ...
> A C# library? Lua? What about that really cool language coming out next week? What about C?
... yes.
> It's not human-readable. Enough said.
"Enough said" is glib, but it doesn't seem obvious to me. The computer is the one using the data the vast majority of the time, not a human. I'd rather optimize for that, improve the user experience and reduce parsing overhead, and use protobuf --decode for my own debugging.
> It's smaller that JSON, sure, but is it smaller than gzipped JSON?
Is GZIP'd JSON smaller than GZIP'd protobuf? What is the decoding time of a GZIP'd JSON file vs a non-GZIPd protobuf file?
Regardless, all of these reasons seem to be more about making your life mildly easier (or at least more closely matching your preferences), and less about optimizing for CPU and bandwidth utilization of the client interface.
We just switched to protobuf because it allowed us to provide the best user experience by decreasing both parse time and transmission cost, AND we can auto-generate the serialization code, including validating the messages for correctness.
If anything, I'd choose to move to an even more rigorous message specification format, as having the validation done for us automatically keeps our client-side code very simple compared to the data extraction and type validation we have to write manually with JSON.
Exactly. I love having JSON libraries at my disposal no matter whether I'm coding in Python, C, Haskell, Racket, JS, whatever. That's hard to beat for coders like me (generalists).
I'm working on a Protocol Buffer library that can serialize/deserialize to either JSON or Protocol Buffers. That way you can do all your development with JSON, but if you ever find you need the efficiency improvements of a binary format, you can just change your Serialize() call.
Having a .proto file gives you the benefits of something like JSON Schema: a place to document all your fields and what they mean, and a few very simple validation constraints like the expected types.
It's sad, really. The people who try to call bullshit on those sorts of claims are frequently drowned out by those who are blinded by the "4x faster!" etc.
It consists of serializing 3 integers and a string 200,000 times.
MessagePack defines 27 different types¹ (excluding reserved ones) with variable bit length for the type identifier, length somewhat correlated with frequency of use.
A benchmark should therefor test real life data and a lot of it.
Their inability to produce such benchmark makes me question the sanity of splitting up e.g. the type marker for “array” into 3 different types depending on the size of the array — this adds complexity, so it would be good to know what exactly the authors based this design choice on, hopefully not that it made it faster to serialize a 3 element array 200,000 times.
¹ http://wiki.msgpack.org/display/MSGPACK/Format+specification...
For large docs, I'd prefer XML too. I mean, I'd rather hunt down a missing closing tag than a missing closing bracket when the thing is pages upon pages long.
Another poster mentioned tnetstrings, those look interesting, however I am not sure how well those would handle binary data.
My stock “basic data type” set comes mainly from Ruby these days. I wince a little at the lack of interned-symbol type, but it's possible to live without that. But what of string encodings? In a recent piece of code which I wouldn't mind replacing with MessagePack, I prefix strings with fixnum IANA encoding numbers. I suppose that's too much to ask in this context, though… ? How do other people deal with this—just force everything to UTF-8?
* if needed: use snappy for blob storage (ie for key value stores)
* if you have an need for very high perfomance in your application that is mature to some degree, evaluate some binary protocols or compression and see how they perform.
the bad thing is that i already see full stack frameworks popping up "now with <binary-protocol-xy>" and everybody will scream "YEAH!". Most YEAH!-sayers will seriously be butt-hurt by the plain fact that it isn't human readable. Switching between JSON/binary won't matter, there will be situations that are not debugable.
Trivial performance optimizations won't fix a broken design.
As far as JSON being human readable, most JSON sent over the wire has no whitespace and generally needs to be run through a parser if you're going to read any significant amount if it.
You even say in your third point that if you need high performance you'd look at a binary protocol, this is one such protocol. Also, you completely ignore the space savings mentioned.
Your name is green, you're new to HN, and frankly comments like this drive away the kind of users we want here, so learn to talk to your peers respectfully, or keep your mouth shut.
I completely understand the need for JSON when you are communicating with the frontend/Javascript. But if you are doing backend messaging, would you not rather use ZeroMQ (which comes with its own protocol).
From what I understand (from previous HN posts), it is super fast and handles binary data very efficiently. I also understand that you can tune stuff at the OS level to wring the last bit of performance from ZeroMQ.
P.S: yup, it is written in C, but all bindings apparently work very well.
Once you've got your data expressed as a sequence of bytes, you can send that with ZeroMQ. Or with HTTP, or raw sockets, or whatever. The serialization format and the transport protocol are pretty much completely independent.
I've done several of my own benchmarks as well and can confirm that the current Node JSON implementation is much faster than MessagePack. Their benchmarks are most likely quite old.
There are benefits to MessagePack that have already been mentioned here, namely not having to base64 binary data first (smaller size), but that's true for any binary message format. I'd love to see some other binary formats thrown into the ring and see how they compare to MessagePack in both size efficiency and encode/decode performance. BSON seems like an interesting option, but I don't know enough about it to comment...
Honestly I'm not sure why you wouldn't just simply use gzipped/JSON. My test show nearly no difference (1-2%) in performance with MessagePack, yet you get to leverage all kinds of things that already understand JSON.
http://wiki.theory.org/BitTorrentSpecification#bencoding
Seriously, this is not an either or question. Just use the right tool for the right job. I wouldn't serialize my data to XML on a micro controller but so I wouldn't drive my JS frontend with binary serialized data.