MessagePack: like JSON, but fast and small
msgpack.org
msgpack.org
and complete.
Both are minimal self describing data serialization formats, but JSON is incomplete. It misses a type to represent one of the fundamental type: byte blobs. And also parts of floats.
Which means there are a lot inconsistent ways to hack that in e.g. base64 encoded strings. But this means it's loosing partially it's property of being self-describing (same if you allow numbers as strings e.g. "-12").
Just to be clear I'm not saying JSON should have bytes directly in it. But a "native" base64 string additional to string, number, null, list and map would help. E.g. `{ "bytes": b"YWJjZGU=" }`
Wrt. floats it's missing (+/-) Infinity, NaN is more like a error variant so it's okish to be missing, but then why not have it.
Also for completeness it would be better to differentiate between int and float as float is imprecise due to rounding errors.
----
PS: I don't like `null` but having it in this kind of data serialization format is still required. But the difference between a field not being there and it being `null` can be a mess and not all tools handle it well.
PPS: Both JSON and MsgPack can be used for non self describing serialization, e.g. by serializing the a record (fixed number of fields in known order) as a list of values instead of a mapping. But both are focused on enabling self describing serialization.
EDIT: PPPS: Yes, it's a very weak form of self-describing is use here, there are systems which are much more self describing e.g. XML+XML Schema linked from the XML document but also tend to be far more complex
Schemaless has become a horror-show.
That's webtech in a nutshell isn't it?
Nobody says "what a shame I can't open that jpeg image in a text editor".
Everyone understands that an image is an image, and there are tools for dealing with images, like photoshop.
The same could be said for data interchange formats like json.
JSON vs. Bencode, JSON can get you the data good enough a lot faster, while a bencode parser is easier to write. My feeling is that most formats with schemas usually fail somewhere in the details anyways. Honestly I think I want too much from them; e.g. the YANG Data model language is a kitchen sink but lacks a lot.
Nobody except me, of course. I'd be thrilled if I could open up a JPEG in Emacs and e.g. view/edit the EXIF metadata in a text buffer (and/or pop up a graphics buffer for editing the image itself). Similarly, I'd be thrilled if Emacs was able to parse MessagePack data and let me view/edit it in a buffer. Both of these things are theoretically possible, but to my knowledge nobody's actually done them yet.
Granted, calling Emacs "just" a text editor is pretty stretchy, but still.
We're not. We're gzipping the json when we transmit it.
The problem with JSON is that it's just human-readable enough that people think it's a config file format. It's not.
I call BS.
As a counterexample search for the word "comment" in https://www.w3.org/Protocols/rfc2616/rfc2616-sec14.html. Or are you saying that http is not a computer-to-computer interchange format?
It turns out these "comments" were a disaster and user agents are moving away from even providing them. You put "firefox" in the user-agent header, and web servers would read the comment and send you "your browser doesn't work" instead of the actual document. That did not make for a very compatible web... turns out comments in machine-to-machine communication are a bad idea.
But that doesn't change the fact that people put comments into machine protocols.
For another example, lots of machine to machine protocols specify XML. Such as SOAP. They therefore all support comments.
In accord to the general trend, eventually they get abused into having a semantic meaning. For example https://support.ptc.com/help/windchill/wc111_hc/whc_en/index... shows how one system uses comments in SOAP to create documentation and a WSDL that lets interfaces to your code to be automatically generated.
And so comments eventually become executable. But that doesn't change the fact that comments can exist in machine to machine formats.
How about ".comment" sections in ELF binaries, do they count?
You have to look at JSON in a historical context to understand it. It exists so that someone could get some data into their Javascript program with "eval". It was then standardized so that you didn't need a full Javascript engine to understand it, because Python and Perl and Ruby didn't have one, and it turned out that evalling random data from the Internet was a security disaster. Does that make it a good human-readable config file? Nope. Does that make it a good interface definition language? Nope.
It exists because it got popular early, and now we're stuck with it. Now it's too late to apply band-aids to make it a human-readable config language or a good computer-to-computer interchange format, because you will always have old parsers around and people will naturally want to target those. Trust me, 99% of developers will moan when you tell them that they have to use GRPC+Protos to access your API. They will be equally mad about your JSON extension that has comments and integers in it, because the dialect of Brainfuck they use for all their projects doesn't have a library that supports those extensions and their editor can't syntax-highlight it or autoformat it.
I would like to give you a solution to this problem, but there isn't one. JSON "won", but it's bad at everything except being well-understood. That seems to be all that people really care about.
That attitude is misguided. Many people do in fact use JSON as a config file format, so it is a config file format. Real world usage outweighs prescriptions about what is and isn’t correct.
Looking at the first standardized JSON spec, it’s not particularly prescriptive about usage, simply saying:
JSON is syntax of braces, brackets, colons, and commas that is useful in many contexts, profiles, and applications.
Of course, it would be a better config file format if it supported comments, but people use it even despite that. I think we should try to understand the reasons for that rather than just telling people they’re doing it wrong.
JSON is one of the worst config formats, even worse than the Apache or nginx config formats. Not being able to simply comment out a config line temporarily for testing purposes or to leave a explanatory comment for special setting is simply bad.
I don't know why people would prefer that over YAML or others but at least you can actually add comments in them and the reason for their popularity doesn't seem to be baked in parser support because these tools are adding their additional "JSON" parsing anyway.
Still, everyone forgets HOCON [1] is a thing, which solves many of the problems expressed here about JSON -for configuration-. It's easy to clearly specify things like time, reference other parts of the config, or if you want to change just one value, you can add that to the 'end' of the HOCON file and be GTG.
[1] - https://github.com/lightbend/config/blob/master/HOCON.md
Skimming through it, concatenation of unquoted values is where it goes off the rails for me. This is quite the gotcha: https://github.com/lightbend/config/blob/master/HOCON.md#not...
My wishlist for a friendlier JSON would be, in priority order:
- Allow trailing commas
- Comments
- Allow newlines instead of commas
- Allow unquoted dictionary keys
And I think that’s it. I’m not even sure the last one is worth the extra complexity.
...and those people need to face the consequences of their poor and missguided technical decisions.
In this day and age we have no excuse to repeat the "but everyone is using XML for that" mistake. Just pick the right tool for the job and stop complainig that the tool needs to change to compensate for your poor judgement.
If you want comments, pipe it through json5 first or something. VSCode again supports comments as a courtesy. More tools that use JSON are starting to as well.
It's just not a big deal.
That's just tooling compensating for the shortcomings of a format being shoehorned into a use-case that falls outside of it's scope.
It's the XML nonsense all over again.
> It's better done and more mainstream than any other solution I've seen compared to people saying "well technically you could build that for <pet format>."
That's the same short-sighted line of argument that was used to force the mistake of using XML everywhere.
There are right tools for the right job. JSON os the right tool for a lot of jobs, but config files is not one of them.
Even though it’s not ideal, I actively prefer JSON over every other random format I’ve come across. This is speaking as somebody who mostly has to tweak existing configs, rather than extending them or writing new ones.
YAML (along with I guess TOML etc) looks nice at first glance, but it has too many weird syntax shortcuts that make it hard to figure out what’s actually going on. And googling for syntax like “[ ]” is hard!
With JSON, the data model is super simple, and for a given set of data there’s only one way to write it down. Sometimes the data model is too simple, sure. But even for fiddly cases like dates and times, there’s often an obvious solution (in this case, ISO 8601 strings).
Do we use the same Json me and you?
> Json is an open standard file format, and data interchange format, that uses human-readable text to store and transmit data objects
- https://en.m.wikipedia.org/wiki/JSON
Json is text based, and meant for storage and transmission in a human readable format.
> Although Douglas Crockford originally asserted that JSON is a strict subset of JavaScript, his specification actually allows valid JSON documents that are not valid JavaScript; JSON allows the Unicode line terminators U+2028 LINE SEPARATOR and U+2029 PARAGRAPH SEPARATOR to appear unescaped in quoted strings. > JSON is a strict subset of ECMAScript as of the language's 2019 revision.
https://en.m.wikipedia.org/wiki/JSON#Data_portability_issues
But crockfird said it was done to keep the parsing simple (and thus secure), and between seeing how XML has it, and seeing how poorly many json parsing implementations are despite this simplicity, I have begrudgingly decided to accept it. Still wish there was a way without those problems though.
I think this was stupid. First, it happens anyway to some extent with people just using custom byte encodings for "strings". Second, using comments for metadata could be useful -- even if optional. For example, adding types after strings to specify things like the date format.
Not true for most numbers. A 64-bit IEEE float can handle any 53-bit integer with no error at all.
It's also non-extensible so these types it's missing cannot be added without assumptions or overhead from convention.
Transit is a good solution to this.
1. ISO UTC 2. ISO local 3. Epoch 4. 2018-04-20 5. Jan 01 2016 11:58am
Probably others that I can't recall. It's a damn mess.
That will probably cause a twitch to any dev that used Asp.NET up until a few years ago.
I never understood why simple integer unix timestamps are not more prevalent in API-s? Or some monotonic count from any reasonable epoch, depending on the context. How many APIs really have to ever return dates predating unix?
Current precision with a 32-bit float (JSON/JavaScript are 64-bit usually, though in non-JS it's common to use Bigdecimal in slight non-compliance) is much worse than one second.
That's the point, timestamps should be self-describing. If "look up the structure in documentation" suffices we might as well just use protobufs.
I dunno, how do you distinguish "April" (the month) vs "April" (someone's first name) or 18 (kg) vs 18 ($) vs 18 (-th of the month) vs 18 (page number)?
> It is lower fidelity (seconds vs milliseconds).
This isn't a problem if you have a conforming implementation (2^52 milliseconds is over 100'000 years), but as nitrogen pointed out, you do apparently have to worry about that.
This is a frequent gotcha with Typescript, where even if your type is declared with a Date field, Javascript won't care when it deserializes it, as that type information is all erased and not available at runtime.
(Also, Number in JS being floating point and all, it lacks the integer precision for high resolution timestamps - if you start serialising 64 bit timestamps and expect a JS-style runtime to do good things with them it doesn't end well.)
Ah, thanks, that makes sense; I'd forgotten that Javascript had a built-in Date type. (Strictly speaking, then, you ought to be able to write something like `{"time":new Date(...)}`, but that obviously doesn't work in practice.)
Types! Extensible types. Like I said above, Transit is a good solution to this. With a fixed set of types, arbitrary data will always have ambiguity.
You alarm will be an hour off when you convert to UTC beforehand.
A timestamp is a point in time, whatever the time is in Paris or Tokyo. It is an abstract value and it is way better this way.
A timezone is a filter through which you display your timestamp, and it tells you what your local time was when that timestamp occurred.
So yeah, always store and exchange time information as timestamps. Timezone is extra styling information, just like CSS.
> Timezone is extra styling information, just like CSS.
This is an incredibly reductive view of time data and the possible applications that use it.
Not all time data is timestamps, and UTC timestamps strip out information that is not "styling".
You might not have encountered such scenarios but they are very much out there.
Expressing time based on the standard reference timezones (i.e., UTC) is not stripping the time zone away.
Example: The EU will most likely get rid of DST around 2022. Any time you set beforehand in CEST or CET will be an hour off, depending on which the EU will get rid of.aybe the EU doesn't get rid of DST and keeps it, you don't know. So unless you have a time machine, you cannot convert the time to UTC.
There are lots of examples like this. Some time events you want to happen at absolute time points (in which case time zones are only for display purposes), but very often your events need to be ”time-zone aware”. There is no universal rule, and thinking that there is will lead to all sorts of trouble.
A timestamp is a point in time, and is the same everywhere (excluding relativistic effects).
What you describe is definitely a valid use case (and having developed software that programs hardware that is controlled by a calendar, I have painful experience in this). However, it's not a timestamp. I'd call it time-of-day or wallclock. Some systems like SQL refers to it as simply "date" and "time". Whatever you prefer to call it, it's not a timestamp.
How you represent it doesn't change the definition. It could be stored as the number of seconds since some arbitrary point in time, such as Unix timestamp, or a Julian timestamp.
Or, it could be a given time and date combination with some fixed point of reference, such as UTC. I'm sure we all have favourite ways of representing timestamps.
When I set an alarm for any time in CEST in 2022 and convert this to UTC before saving it, it will very likely ring at the wrong time, simply because CEST will probably not exist by then and be replaced by CET due to the EU getting rid of DST.
Dates should not be encoded with time zone info anyway, as timezones is context-dependent. Dates should be encoded in UTC and then let clients interpret them according to their context.
Maybe OP meant timestamp, but that's not what they wrote.
Read the spec. Basically all it says is that the numeric type is a float, but doesn't say anything about its precision.
I'm amazed that there haven't been any security vulnerabilities found yet that takes advantage of different number formats in different JSON implementations.
You'd think that a data serialisation format should get at least numbers right. All the ones that came before (and after) did, yet we're still stuck with JSON.
As it turns out, the only way you can reliably store an integer in JSON is to use a string field. This is incidentally also the correct way to store a username in JSON.
So I ended up using JSON. Yes the message sizes are larger in byte size with JSON but using the jsmn[1] parser I could avoid dynamic memory usage and code size was small. The jsmn parser outputs an array of tokens that point to the buffer holding the message (ie start and end of key name etc), so overhead is quite limited.
For JSON output I modified json-maker[2]. It already allowed for static memory usage and rather small code size, but I changed it to support a write-callback so I could send output directly over the data link, so I didn't have to buffer the whole message. This is nice when sending larger arrays of data for example.
Combined it took about 10kB of program (flash) memory, of which float to string support is about 50%. Memory usage is determined by how larger incoming messages I'd need, for now 1kB is plenty.
A nice advantage of using JSON is that it's very easy to debug over UART.
Though having compact messages would be nice for wireless stuff and similar, so does anyone know of a MessagePack C/C++ library that is microcontroller friendly?
I threw away NanoPB in favor of TinyCBOR and haven’t looked back.
https://github.com/ludocode/mpack
It can be built to a very small code size, especially when you disable libc, allocations, etc. There are some people using it on embedded devices like Arduino. There's someone working on a port to 8-bit microcontrollers that don't have a 64-bit float, so you may want to look into that as well; see the open issue for it on the GitHub link above.
If you want small code size, CBOR seems like a good bet:
> The Concise Binary Object Representation (CBOR) is a data format whose design goals include the possibility of extremely small code size, fairly small message size, and extensibility without the need for version negotiation. [1]
This [2] C-implementation fits in under 1KiB of ARM code.
[1]: https://cbor.io/
And Amazon has picked it up as a first class citizen in some of their IoT Core features. It’s definitely here to stay.
Anybody know if (and why) U2F uses CBOR?
They’re extremely similar.
Also I don't know how they got "code sizes appreciably under 1 KiB". On my STM32F1 release mode with -Os it adds about 12kB.
But yeah, maybe I should reevaluate CBOR.
You’re on your own for malloc, which for me is great because FreeRTOS Heap4 management is quite good. So I malloc an object I’m decoding into and parse away.
There are two options parsing arrays and strings/bytestrings and I just chose the option where I specify the pointer to use, vs them using normal malloc then free() later.
I really like this setup. I made a deinit(bad_message) that works anywhere it failed (parse, validate, eval, etc), goes through and looks for pointers that I previously would have malloc’ed.
There is another popular library but I forget what it’s called.
But libmpack is: https://github.com/libmpack/libmpack
- libmpack serialization/deserialization API is callback-based, making it simple to serialize/deserialize directly from/to application-specific objects
- libmpack does no allocation at all, and provides some helpers to simplify dynamic allocation by the user, if required.
- C89
I am not 100% sure why I would choose to use this - maybe for super tiny documents?
Typical gzip decompression speed is somewhere in the 200-250 MB/s region, compression is much slower. LZ4 for example tends to compress at ~600-700 MB/s, and decompress at several GB/s. zstd is tweakable over a very wide range of ratio-speed trade-offs.
LZMA(2) (xz) is a rather troubled format and should not be used any more. bzip2 has always been slower than gzip with usually marginally better compression. It has been irrelevant for a long time.
You mean xz. Not sure where you got the idea that it’s a troubled format, but if you’re talking about the infamous “Xz format inadequate for long-term archiving”, IMO that’s just bzip2 authors taking a dump on xz for no good reason, and fortunately for us it’s bzip2 that’s basically irrelevant today, not xz.
Funny you say that, my company uses bz2 for compressing pretty much everything.
Brotli is better, zstd is faster, lz4 is often good enough and a lot faster.
gzip is fast relative to other compression algorithms and relative to internet speeds.
gzip is slower than an ideal gigabit network, but faster than 100mb.
https://www.rootusers.com/gzip-vs-bzip2-vs-xz-performance-co...
> gzip is fast relative to other compression algorithms
gzip looks fast perhaps compared to "xz -9", but not to anything modern.
Gzip is quite close to the Pareto frontier, meaning it is a good trade off of time and space. See the charts at http://mattmahoney.net/dc/text.html (And read up on the Hutter Prize!)
With MessagePack the schema is embedded alongside the data, exactly like JSON. So it's just binary instead of text encoding, which saves some space, but as you pointed out, standard text compression algorithms are going to likely perform similar or superior.
B64 as originally used (with an extra >2.6% for line splitting) is handily over 35%. As we use it now it’s still a bit over 33%, because of the == endings to detect truncation... 66% of the time.
I guess the ability to have a schemalass and no type checking format (just like JSON), while enjoying some performance benefits? It just feels like a weird niche to me.
Why would it be your default?
https://github.com/thekvs/cpp-serializers#results
That said, currently I don't work with it in any projects. Apparently it hasn't really caught on.
At the end of the day, though, I'd seriously much rather just have a halfway well commented *.proto file. Work smarter, not harder. It's just that I won't recommend that for public-facing APIs because GRPC isn't universally well supported, while JSON over HTTP is.
Avro is pretty similar to MsgPack in concept. It seems less popular (at least in my circles), and has fewer implementations, which may or may not matter. As for performance and efficiency, the first benchmark I found [1] shows MsgPack is more space efficient, and faster to serialize, while Avro is faster to de-serialize. The second benchmark I found [2] found exactly the opposite. So it's not clear to me that either is better, and I'm sure it depends on your data, which library you're using, and how you're using it.
Being nearly as performant as the predefined-schema serializers, while being nearly a drop-in replacement for JSON, seems like a major and valuable use case to me.
[1]: https://medium.com/@nitinpaliwal87/compression-and-serializa... [2]: https://github.com/saint1991/serialization-benchmark
While I see the same numbers as you do, there's no way a format that includes plaintext keys alongside the values is going to be smaller than a format that only includes a packed field number+type and the corresponding values. There's something fishy going on here, and I can't seem to find a link to the source behind his benchmarks?
In force documentation:
https://www.itu.int/rec/dologin_pub.asp?lang=e&id=T-REC-X.68...
Blog on ASN.1 schema for JSON:
The big difference is that messagepack is schemaless.
>ASN.1 is a mature standard. As I already mentioned, it has been around since the 1980’s. Though stable, it is not stagnant; the most recent revision occurred in 2008.
It might be old but it is not widely used. To me it doesn't matter if something was created yesterday or 20 years ago as long as I can use it and using ASN.1 is significantly harder than necessary.
>ASN.1 has tool support. There are both commercial and open source tools that will generate code from your ASN.1 specification.
Ok, so where is it and why is it so difficult to find?
Even if we assume that we should just build it today then all the above claims suddenly become worthless. If someone builds ASN.1 tooling in 2020 then it is just as immature as e.g. GraphQL tooling that was built 2020. If there is renewed interest in ASN.1 then it might gain new features that will cause it to become less mature/stable again.
Having "mature" software is useless if it doesn't meet user demands.
I've spent a lot of hours poring over hex dumps of BER messages and CBOR messages (which are basically MsgPack) and I vastly prefer CBOR. But I prefer JSON way more.
How much time would it take you to either
a) turn on gzip support on your HTTP server and keep using JSON
b) rip out your JSON stuff from controllers and whatnot and replace it with a custom serializer/deserializer that's far from an industry standard?
1. msgpack + gz would be a more fair comparison to json + gz when comparing file size
2. have you run any time comparison? It would seem to me that msgpack gets pretty close to json + gz with far fewer resources, time
I havent had need to deal with such constraints personally, but maybe its something I should consider more. It doesnt need to be a 10x win to still be useful, and if its a simple import to use then i would likely be crazy not to consider it.
Without GZIP:
JSON 583 bytes
MessagePack 304 bytes
52 %
With GZIP:
JSON (GZIP) 289 bytes
MessagePack (GZIP) 248 bytes
85%
I just took a quick look and they do compress their queries, using Brotli, which is even more efficient than GZIP. They are behind Cloudflare, so it's probably from them.
I just tried with https://jsonplaceholder.typicode.com/todos and I get this:
Raw:
JSON: 24,311
MessagePack: 14,704
60 %
With GZIP:
JSON: 3,965
MessagePack: 4,063
102%
With Brotli:
JSON: 3,495
MessagePack: 3,704
106%
So essentially, MessagePack compressed is WORSE than JSON, and not compressed, is at least 3 times worse than compressed.
In the case where your data can't be cached but CloudFlare proxies and compresses for you, MessagePack wins because your uncompressed connection to CF is 40% more bandwidth efficient so egress out of your app server for any traffic is reduced. In the case where CF can cache responses you get the same back end bandwidth win for a negligible (if any actual) amount of extra egress out of CF.
Use what you want but if you need a schemaless serialization format, MessagePack works well and is compact for "free".
{"red": 127,"green": 28,"blue": 27,"x": -18,"y": 27,"z": 2008,"time": 1583862185,"dollars": 1298,"cents":12,"name":"tom",timeseries":[1,2,3,4,55,66,77,218,239,340,9009,80008,400004]}
gives 106 bytes with msgpack and 157 with gzip. Also, if you gzip the msgpack, for your message example, it's 248 bytes.
- https://docs.microsoft.com/en-us/security-updates/securitybu... - https://en.wikipedia.org/wiki/Billion_laughs_attack - https://en.wikipedia.org/wiki/Zip_bomb - ...
current_addr + message_len - start_addr < buffer_len
Or am I missing something?Similar to HeartBleed, where there wasn't validation on the heartbeat message, and the server would echo back buffer_len instead of just what was sent.
I can think of a very contrived situation where this can be a problem, but in most cases this will be perfectly safe.
ie just because MessagePack is a binary format it doesn't mean you can skip the same string checks that JSON requires; which means parsing MessagePack strings is unlikely to be any faster than JSON strings (contrary to the suggestions others have implied with the "just memcpy" comments). It's just with JSON that validation is done as part of the parser (remember JSON only technically supports a subset of ASCII and any extended characters or unicode is encoded via escape codes) where as with MessagePack you'd need to do that validation as an additional step.
Integers, on the hand, might differ since JSON would need additional validation (again, backed into the parser) which MessagePack would not because MessagePack encodes the integers as binary integers where as JSON encodes them as ASCII values that would need converting back to binary integers.
(hint: read the message I'm replying to).
However if you’re accepting MessagePack encoded data from insecure systems (such as end users) then you absolutely should be validating your input somewhere along the pipeline and it’s usually better to do that early on.
Also it’s not generally the distributed systems you worry about when it comes to this specific degree of micro-optimisation (which is basically what this is). It’s the monolithic ones. Distributed architecture is meant to solve various problems (for example but not limited to, high availability, reduced geographical latency, single site but running on cheaper commodity hardware, etc) but often at the cost of CPU cycles. Whereas your monolithic infrastructures where you have fewer servers (such as Stack Overflows set up) would be greatly more dependant on reducing computational overhead where corners could be cut. However they’d also be significantly less likely to need networked RPCs via MessagePack anyway (simply due to the monolithic design of their architecture).
Yes, so what's the problem?
Meanwhile keep in mind that the JSON document was already passed in a HTTP request body, or that its trivial to put string length checks on a parser, or even nesting limit checks.
That said, not something I want to prematurely optimize for.
I don't see your point at all. The state machine required to parse a JSON string essentially has only 2 states: current token is either a string character or a string delimiter. That's it. Adding a record size increases the number of states because now you not only need to track if a token is a string character but also if the character makes sense to be there within that state.
Moreover, when you parse a JSON string, once you hit the string's end delimiter you already are able to calculate the string size. Thus, if the string is short enough to fit the buffer then not only is the lexer simpler but it also requires the same memory allocations than strig formats with record sizes. If however a string doesn't fit the buffer then we are already in the territory of a ropes data structure in both cases, thus the number of memory allocations tend to be equivalent.
Edit: found my post https://stackoverflow.com/questions/9884080/fastest-packing-...
I no longer have any connection to this repo, but the benchmarks (and a schema validator gem I wrote) are here: https://github.com/deseretbook/classy_hash/blob/master/READM...
I'm not against messagepack or protobuf, however, much like all things, I'd rather start with simple http+gz+json (maybe websockets) and optimize as needed. Not everyone is at the scale of FAANG, and most don't really need this level of optimization.
Um... no? Epoch as binary, 4 bytes. As JSON, 10 bytes. 64bit big num in binary, 8 bytes, in JSON 19 bytes.
For a small example object I can think of
{ “x” : “y” }
In binary, 5 bytes. In JSON, 9 bytes.
Now... light ENOUGH because you’re using a PC and gigs of ram and a 100mbit internet, sure. Light in terms of a microcontroller? No.
I flat out could not use JSON in any of my projects and I’m not at FAANG level.
Sure, the HTTP overhead and TCP overhead under that are significant if you're transferring a single bignum. How about 10,000 of them?
Even if the best case of all strings where binary gets limited by ASCII's inefficiency, "":"", is "only" 5 bytes of wasted data, but what about a million item JSON? 5 million wasted bytes does seem like it outpaces NIC then TCP then HTTP layer overhead. But uea, IDK.
The claim was that JSON is lightweight, it is not. It's not as bad as XML, I'll give it that.
Then I tested it with test data streams generated by the reference C++ implementation and Python implementation.
It's actually kind of a pain, because the C++ serializer generates a variety of different data types depending on the values you are packing, not the type of the values you are packing. Let's say I encode a uint32_t field. The stream might get a uint8, a uint16, a uint32, or even an int32 (for reasons that completely elude me).
Also, C++ strings come out as the 'ext' type while Python strings come out as the 'string' type, so I have to accommodate both, even though they are both basically byte strings.
So, I want to tell the deserializer what kind of output field I'm expecting, for each struct member or array, and then look in the data stream to see if there is a data object there that came from the same type. But this is impossible, so the per-type decode functions have to be quite complicated to handle a variety of types.
So - it works but I can't do much on the deserialization side to verify that the data I'm unpacking really matches what was encoded. I can only detect very broken cases, like when I'm unpacking a uint32_t and in the stream there is an int32_t with a negative value.
I guess this is mostly done for optimization, but two things would make it a lot better:
- if the spec actually specified how data types in different languages were allowed to be encoded
- if the encoded data contained _two_ type fields, one indicating the original source data type and another indicating the type it was encoded into.
Basically the spec is just way too "loose" to make it usable for the use case I'm trying to use it for, which is to easily generate data that is sent to a micro and stored in EEPROM, then deserialized out of EEPROM later.
That's probably not very close to a use case the original designer had in mind. But I haven't found anything that works better (less decode logic).
You generally shouldn't worry about the low-level type that a value was encoded into in the MessagePack stream. It's dynamically typed, so you should just care about values. When encoding you should allow the encoder to use the most efficient representation, and when decoding you should be able to tell your MessagePack parser the integer width you want instead of caring about the original type or how it was encoded. It should then accept any packed integer type as long as the value is in range.
This is how my MessagePack implementation works as well. If you expect to receive an integer that fits in, say, `uint16_t`, you can call `mpack_expect_u16()` or `mpack_node_u16()`, and it will allow any integer representation as long as the value is in range.
It sounds like this is where you were going with your implementation as well, so this may not be comforting because it's not what you want, but it is at least the correct way to understand the format. I've talked about this pretty extensively and wrote up a protocol clarifications document that explains a bit more about how and why MessagePack libraries discard integer width and signedness:
https://github.com/ludocode/mpack/issues/35
https://github.com/ludocode/mpack/blob/develop/docs/protocol...
If you really want things like original integer width represented in the format, ultimately you're going to want to use a different format, probably one that is non-dynamic and uses schemas.
As far as the string vs ext, you may have meant string vs bin; there was a format change a while back that separated string and bin types and not all MessagePack libraries have adapted to that. Many libraries (including mine) support a compatibility mode so they will use only compatible string representations.
> Msgpack is an PECL extension, thus you can simply install it by:
Having something available in PECL is a good first step, but nobody will use it unless you either:
1. Get it into the standard library (which requires an RFC for PHP Internals), OR
2. Write a pure-PHP polyfill installable from Composer, OR
3. Do #2 then #1 (using the polyfill's popularity to argue for the importance of the RFC acceptance to make #1 a reality).
Reason: A lot of the places PHP is deployed, you can't compile C code or install binary dependencies (.so, .dll files). You can't access the OS package manager, either.
But Composer is a pure-PHP package manager that still operates in these environments.
So if anyone on HN ever wants your thing to be used by PHP developers, don't just stop at "PHP extension, written in C, available in PECL".
If you’re in the target audience for msgpack you probably aren’t relying on shared hosting and can build a pecl extension.
I have multiple racks of company owned bare metal that I deploy to. Still don't want to build a pecl extension. Much simpler and way less likely to not run into random build issues if I can just install via composer.
Docker? Build it as a system (RPM/DEB) package? Man, my life would be difficult if I just flat out refused to use packages for PHP/Python/Ruby requiring native extensions because it required some minimal effort on my part to deploy it.
Seemed like a free way to offer slightly better performance on the API (both serialize/parse times and bandwidth), since it can just serialize any data without specifying protocols or schemas, like json serializers can too. I didn't know about any alternatives that would also require no further configuration or infrastructure, so it seems like msgpack fills this 'free performance boost' niche quite nicely
Msgpack, unlike json, is not human readable on the wire, not a purely text based format, and I doubt it is smaller or faster than protobufs or flatbuffers.
Or CBOR, that has an RFC (which is important for some folks):
https://github.com/neuecc/MessagePack-CSharp#comparison-with...
FlatBuffers doesn't have as many client libraries. There's a MessagePack library for a ton of languages.
One thing I really like about MessagePack is that the Python client (and others too) supports reading from a stream. So you can write a bunch of msgpack messages to a file or TCP socket and it just works.
Protobuf can't do this out of the box because it doesn't include how long the message is. You can write a wrapper that specifies the message length, which isn't that hard and I've done before, but it is another thing to maintain. And other formats do that out of the box (ex. Cap'n Proto)
And as someone else mentioned, Protobuf doesn't have NULL which is useful in some cases. (I understand that Go has strong opinions about there being a useful default value, but that doesn't map well to a lot of languages)
One other thing with Protobuf is that the Python client is not very pythonic. I've been keeping my eye on this project [0] which makes Protobuf messages work just like dataclasses. They don't support OneOf types currently, which I happen to need for some of my use cases. But they're working on it [1]
So, it's do-able. But, I'm just saying that other formats/libraries support this natively out of the box.
1: https://www.tensorflow.org/tutorials/load_data/tfrecord#tfre...
Yeah, it's all about the tradeoffs with formats. Protobuf is a good choice a lot of the time and I always think about it as one of the first/preferred options. gRPC is great too.
It supports compression too. Even if the resulting serialised size isn't always smaller than compressed JSON, from a performance standpoint it's a lot better.
Especially if you use the MessagePack-CSharp library. That dude knows his high-perf .Net.
https://www.adobe.com/content/dam/acom/en/devnet/pdf/amf-fil...
It does support compression, but messages sizes are generally comparable to compressed JSON (YMMV, it will depend on your messages).
That would let it fully take on the role of json
But what's the text format? -- now there is endless happy bikeshedding
Maybe a library with json (or compatible superset, to get all messagepack features). Then the standard just serializes to messagepack, before gzip or whatever
https://github.com/ludocode/msgpack-tools
It works best if you use MessagePack like JSON where all map (object) keys are strings, so you can easily understand a message without context. If you want to optimize your MessagePack more, you would tend to use integers for map keys, but this makes the JSON-equivalent view not super clear because you just see a bunch of numbers in a tree structure.
1.) Decide upon a binary format.
2.) Use it.
No 3rd party libraries required; you could write writer or reader code for a Commodore 64, if you had to.
Keep it simple. I suspect any simple "type-length-value" type of binary file format would be written or read at least as quickly as this, without third party code.
I am opposed to the blind use of libraries like this. Developers need to understand what they're doing and how what they're doing is being done at a reasonably low level if they ever hope to become better developers. Masking it all away behind a third party library is not how you understand the code that's running on your systems.
Also proto bufs gave us those nice schemas, the code gen, and easy forwards/backwards compat.
Was it slower? Yeah. Was it crap ton less work? Yup.
In JSON a number-type value has to be parsed like in javascript, so IEEE754 double.
for example:
{"number": 0.1}
in reality it is parsed as: {"number": 0.100000000000000005551115123126}
Solely deserialization followed by serialization might introduce a change in serialized value.Because of problems with IEEE754 in all of our api we use only floats as strings, like {"number": "0.1"}. For us enforced IEEE754/double format for floating point numbers renders number type near useless.
I've worked on a production app where we had to swap out the JSON parsing libraries in the server (Rails) and all clients (Android, iOS, Rails again) for ones that preserved numbers as BigDecimals. This was a huge pain, it made everything slower, and even then it wasn't ideal because different BigDecimal libraries aren't even necessarily compatible. Foundation's NSDecimalNumber has different limits than Android's BigDecimal for example, and these are the types used by the parsers so you can't just get the raw data to do it yourself.
If I had to do it over again I would never rely on JSON's decimal support. I'd rather stuff my decimals in strings and parse them myself.
We prematurely utilized it and paid the price of "nothing's human readable without first unpacking it" without actually benefiting much given we weren't shipping much data that often."
MessagePack is great, but libraries generally only support deserializing the whole thing. I have an application where these structured documents can be very large, and scanning code sometimes only needs a very small subset of keys.
I believe Cap'n Proto has this feature, but unlike MessagePack it's not schemaless.
For example, given a struct like this:
{
"id": "123",
"name": "Developers",
"members": [{
"id": "567",
"permissions": [
{"type": "read"},
{"type": "write"}
}
}]
}
Let's say I only want the name, the ID of each member, and whether permissions.read is set. I may want to do something like (Go): StreamingUnmarshal(b,
func(keypath string, parse func() interface{}) {
switch keypath {
case "id", "members.id":
value := parse()
// ... use value ...
case "members.permissions.type":
if parse().(string) == "read" {
// ...
}
}
})
Random access could be also work, as long as it didn't need to sequentially parse from the beginning of the data to get to the right value each time. Something like: id := GetKey(b, "id")
memberIDs := GetKey(b, "members.id")
permissionTypes := GetKey(b, "roles.permissions.type")
The trickiest bit is treating arrays of nested structs (as in "roles.permissions.read") correctly, although efficient scanning of keys becomes an important optimization point, too.Probably the best method would be to store the keys pre-sorted at the beginning of the data, so that they'd better fit in the CPU cache, and have pointers to the offsets of the values:
KEY1,KEY2,KEY3,VALUE1,VALUE2,VALUE3
Arrays of structs are tricky here, again, but this is solvable.No, seriously: It doesn't require an external schema, supports every kind of indexing and fast random access you could want, is supported in like every language, OS, and architecture, has copious tooling, documentation, and community support, and is battle-tested across literally billions of installations worldwide.
Also, I am talking about individual documents that already live in a database such as PostgreSQL. I can't store an entire SQLite database in a single column.
Sure you can, it's just a file. :) sqlite scales down nicely to data sets of just a few kilobytes -- if you're worried about parse time of your documents then I assume they are larger than that.
That said, if you're already loading the whole blob from a single row in Postgres anyway, then is random access such a big win? Or is the idea that you would selectively read byte ranges out of Postgres? If you're already pulling the bytes into RAM then avoiding the parse isn't that huge of a win.
(I say this as the author of Cap'n Proto which is all about zero-copy random access... it's only a big win in certain use cases, like mmap() or shared memory IPC.)
I can't imagine that the overhead of initializing an SQLite database from a small byte array in memory is that small, not to mention the overhead of maintaining the table schema.
Out of pure curiosity, I glanced at the Go bindings for SQLite, and there's no provision for initializing a database from a byte array, or accessing the raw underlying byte data of a live database. The C API supports implementing your own VFS for custom storage, but that's not supported by the Go bindings, and seems like a lot of work.
You're right about loading whole blobs; I was misremembering a little bit. The application in question already pares down the document keys in its queries to avoid sending everything. I'm in the middle of a research project into an alternative backend where the documents are stored as binary data, not JSON, and given a set of keys/keypaths, I want to do a little better than deserializing the whole blob.
"But, maybe some of these binary JSON formats are smarter!" you think. Well, almost all of them aren't. Only two are: BSON, and Amazon Ion. Unfortunately BSON limits the message size to 2 GB.
Amazon Ion is also the only format that actually deduplicates object keys. Definitely the most capable and well-designed of these formats. Unfortunately it is also the most complicated.
Sadly a lot of these formats make questionable choices, like storing numbers in big endian format (why?), using explicit `uint8`/`uint16`/`uint32` sizes rather than something like Protobuf's varint, etc. And there are also a load of them that are nearly identical. You really have to dig deep to find the critical flaws.
https://github.com/liteserver/binn
The reason most formats don't length-prefix everything is because it makes it costly to encode in both time and space. You have to basically encode a message inside-out to calculate the nested sizes of everything. This is going to be hugely slow and memory-intensive if you're encoding a 10 GB file, and it's useless for messages on the scale of kilobytes so there isn't any point. MessagePack on the other hand can be encoded in one pass from start to finish as long as you know the element counts of your maps and arrays beforehand.
> storing numbers in big endian format (why?)
Embedded processors tended to be big-endian, like older PowerPC and older ARM. These formats are designed for embedded so it (probably?) improved performance on those processors. This is less true now since virtually all modern ARM processors and probably most other embedded processors now run in little-endian mode.
Ultimately what it comes down to is that these formats are designed for the opposite of your use case. I don't know what you're using a 10 GB JSON file for but there must be a better storage solution for you than a schemaless serialization format.
Yeah this is true, except for BSON because it uses fixed-size length prefixes, so you can just go back and fill them in later. Presumably that's why they used fixed-size lengths. The downsides are it is less space efficient and limited to 2GB.
In any case Amazon make the very good point that formats are read more often than they are written. It makes sense to optimise for the read case.
> Embedded processors tended to be big-endian, like older PowerPC and older ARM.
Nobody uses PowerPC anymore, and ARM hasn't been big endian for ages. Also MessagePack isn't designed for embedded systems and it still uses big endian. I don't think that's the reason. I suspect it's from a misguided belief that "network byte order" still matters.
And I totally agree, a schema-based format makes way more sense for my use case - changing is difficult though.
Anyone have any insight?
- CBOR has two ways of encoding maps and arrays: fixed length and variable-length. This complicates decoders, especially those that would pre-allocate arrays and maps to the proper sizes, which significantly reduces decoding performance. The CBOR spec has nothing useful to say about this; it just requires you to allocate indefinitely.
- CBOR defines a canonical representation, including a key sorting order based on binary representation which is just awful. It requires multi-pass encoding which is slow, complex, error prone, and completely non-intuitive: [1,2,3] comes before 100000 which comes before [1,2,3,4].
- CBOR has more types in the core spec, ones that are extremely specific to certain applications or programming languages. It has a 16-bit float, and it has both null and undefined as separate types.
- CBOR defined a system of "tags" with a huge number of extension types. These are supposed to be optional, but of course they only work if both ends support them. Some features like BigNum are well-supported in some programming languages but not others, so CBOR implementations tend to diverge in supported message types.
CBOR as a standard is far worse than the "non-standard" MessagePack it purports to replace. Here's a great HN comment on it from another user (and another MessagePack library implementer) a few years back: https://news.ycombinator.com/item?id=14072598
Your parser or encoder does not need to support indefinite arrays, that feature is clearly designed to be used with some practical limitations like “I don’t know how many but let’s assume less than x, and I’ll send a STOP when I’m done”. Canonical ordering is optional. Yes, a typed system that has more types, IDK what to say about that other than you don’t have to use them. And yes, tags need to be supported on both ends, just like ANY DATA that is being transferred, compare to a strict schema’ed system and I no difference except that you’re only partially required to adhere to the plan.
Maybe msgpack is just objectively better because it has less features. IDK. Doesn’t matter because CBOR got an RFC and is actually popping up in places. If there was a competition, cbor won, right or wrong.
A common way to mess with JSON parsers is e.g. to nest a lot arrays and objects. So you needs max nesting depth. There are a bunch of other ways to mess with JSON parsers, too.
(In case of MsgPack lists and maps: A parser should only per-allocate memory for given size if it has done proper sanity checks, e.g. compares the given length with the remaining byte length of the message. Also you can simply not preallocate and instead, like in JSON parsers, grow your list on demand and just use the length to know when the list ends).
But yes you have to make sure the parser works for your use case.
An annoyance of TCP was that I never knew whether I read all the data. I either read less and leave data unread, or I read more and end up blocking for a long time (or implement timeouts and get the worst of both worlds).
What's a good alternative? Maybe 0mq? All I need is to send some MsgPack bytes to another client over the wire, hopefully without having to guess whether there's more to read in the socket or not.
For some applications, though, the simplest solution is to shut down one half of the TCP connection once you've finished sending your data. That's how rsh, finger, and HTTP/0.9 responses work, and it's a supported option in HTTP/1.0 and HTTP/1.1. Failing that, preceding each message with a byte count, a la netstrings, is fairly simple; or you can use SLIP-like or COBS framing.
So, to take the canonical concrete example, a chat channel might number the messages on it in a monotonically increasing order, and you might tell the server the channel name and the number of the last message you saw, at which point it sends you the messages since that point, if any, then closes the connection. As I understand it, this is how Kafka works, except for the connection-closing part.
In all probability, your life will be easier and your performance will be better with ØMQ, but these hacks are things that work reasonably well and are extremely easy to implement with off-the-shelf tech.
SCTP in many cases suffers from the fact that it doesn't run on top of TCP, so NATs don't know what to do with it. If you have enough control over your network that that isn't a concern for you, UDP with IP multicast is another plausible solution, the one TIBCO used originally IIRC; you can allocate a multicast IP address per pubsub channel or multiplex them. With IP multicast, recovery from lost messages is a concern, especially if 802.11 is part of your network (since 802.11 uses hop-by-hop ACKs for unicast packets) but there are a variety of reliable multicast protocols like SRM to handle that.
Feel free to hit me up for more info, I've been hacking around with different ways of doing pubsub since the previous millennium.
https://gitlab.com/stavros/itsalive
Clients can connect to the server and get updates for the commands that are currently running, which is not high throughput or complex from a networking perspective. I was wondering if there was something lightweight that will do the same, and 0mq seems like the best choice, but a simple loop over the connections seems to work well as well.
I played around with 0mq for this and it works great, but in this instance I might not want to add the extra dependency (especially since I've already implemented it, minus a bug where it'll block if a packet is exactly 4k).
I think adding an "end of message" character (eg a newline) would be the simplest thing to do in this instance.
Uncaught RangeError: Maximum call stack size exceeded
I'm curious to see the savings difference and hoped to with "Try!" but it'll have to wait.
1) Identify medium-scale similarity boundaries in the data structures. E.g.: a sequence of messages in a protocol, such as a C "struct" with a bunch of fields.
2) Compute the binary difference between these structures so that most of the subsequent bytes after the first message are either zeroes or small numbers. Both the sender and receiver have to keep the previous message in a buffer to allow this.
3) Use a high-performance compression algorithm that supports "user provided dictionaries", such as Zstandard. Train it with sample data.
This above is surprisingly straightforward because it doesn't require complex changes to the underlying data structures. You don't even necessarily need to be able to parse it at all, as long as it has large-scale repeating structures that you can identify.
> It’s that is not enough to just know some new cool technology, nod along and go about your day with your assumptions unchallenged. You need to find out more, test it out, have a grasp before committing to it, and, if you’re lucky, learn a thing or two in the process.
It was very clearly just a page on the vendor's site where you could manually upload a zip of documents, which they had simply declared to be an API.
Parsing code has to be branchy to handle many different possible structures.
If you want extreme performance, variable-length strings are a problem -- the old mainframes that had fixed-length "HOLLERITH" strings had a good idea. I like just about everything about Apache Arrow except that it ignores the problem of fast/portable string handling.
Other formats might have been faster, but MessagePack was very easy to use.
(You mentioned strings, so I'll use that to give flavor: a string is represented as a one-byte prefix, followed by a byte length, followed by that many bytes of UTF-8. I'm not sure whether you'd categorize this as fixed or variable length, but that's how it's represented on the wire.)
It seems msgpack still needs a deserialize step to partially read data right ?
Netflix uses flatbuffers and works wonders in low powered devices.
Schema-driven no-compromise fast compact binary formats with no cross-version compatibility: Cap'n Proto, FlatBuffers, SBE, ASN.1 PER, XDR, OMG CDR.
Schema-driven binary formats which allow some cross-version compatibility: Protocol Buffers, Thrift.
Self-describing binary formats: MessagePack, CBOR, BJSON, Bencode, ASN.1 BER, Avro (?), Fast Infoset, AMF3.
Self-describing textual formats: JSON, XML, YAML, TOML.
I'm using "self-describing" here to mean simply that you can recover the structure of the encoded data without a separate schema, rather than that you can attach any semantic meaning to it.
This is incorrect: Cap'n Proto absolutely allows cross-version compatibility, using roughly the same semantics as Protobuf. I believe FlatBuffers does too. (I'm unsure about the rest, haven't studied them in a while.)
> I'm using "self-describing" here to mean simply that you can recover the structure of the encoded data without a separate schema, rather than that you can attach any semantic meaning to it.
Protobuf, Cap'n Proto, and probably several of the other binary formats can parse data into a message tree without the help of a schema, but all the fields will be labeled numerically. MessagePack is only considered "self-describing" in comparison because in encodes human-readable field names on the wire.
- Avro
- CBOR
- SMILE
And BSON anyone? I think not much people besides MongoDB using it though.
Yes, compressing JSON with gzip-style compressor usually yields 0.5-1% better results then equally compressed binary format (in my limited testing). Still the serialization speed and savings on compression are great to have.
I agree- that's why json is honestly great. I avoided it for the longest time, but now I totally see the appeal.
You can easily make a function wrapper that uses JSON in a dev environment and MessagePack in Production, for example.
(Except if you want to use JSON for configs, which a strongly recommend against for a bunch of reasons including missing support for comments).
I mean by now nearly all times you send thinks over the wire they are ecrypted or at least compressed. So inspecting on-wire messages without "proper" more complex tooling doesn't really work. But if you have already more complex tooling involved there is no reason why you need to be able to read the raw message, it could be just converted on-the-fly into a readable format.
Same applies for application development when e.g. logging, you always do some formatting/conversion when logging some data, so not a problem to convert msg-pack to a human readable format. But then today logging of e.g. servers should always go to some form of log server which does add features like search ability by indexing the log and similar. So no problem to have a msg-pack viewer there, too.
Honestly I believe the only reason human readable formats made/(still do make sometimes) sense was due to limited tooling sometimes caused by server limits in computation power on the developers system. And the fact that many binary formats are totally over engineered making handling of them painful. Especially debugging of slightly corrupted data. Which isn't the case for MsgPack.
This has been my experience with JSON as well. Almost all JSON in the wild especially in web services and RPC is at least minified, so you need to pass it through something like `python -m json.tool` to reformat it for viewing. So you might as well use MessagePack and pass it through `msgpack2json -d` to view it instead. It makes no difference whether the underlying format is human readable.
Now we need an IDL that will let you define a structure and have it produce <language> marshalling and unmarshalling routines.
https://github.com/ludocode/schemaless-benchmarks
I haven't compared them to schema formats like Protobuf or FlatBuffers yet because the use cases are pretty different. I like MessagePack for small projects or rapid prototyping because you don't need to integrate any big libraries or set up code generation as part of your buildsystem. (Mostly I got sick of integrating the C++ Protobuf library into embedded projects.)
The MessagePack format is a lot simpler than Protobuf and the best implementations are nowhere near as allocation-prone as the reference implementation so I expect they would beat it flat out on performance, though the messages may be slightly larger. They would probably beat FlatBuffers for encoding speed as well, but I don't expect any schemaless format could beat FlatBuffers for decoding speed.
The somewhat silly video on that page shows the actual difference in performance our users felt after the change. It was a _huge_ benefit, both in terms of loading time, but also in terms of memory, vastly increasing the number of CS:GO rounds that could be analyzed simultaneously.
In a previous life I inherited a service that shipped tons of data (billions of requests a week) as base64 encoded protobuf strings over HTTP. It was a bad solution in so many ways, but there were historical reasons why it had gotten there. The system required a number of servers and I decided to do some profiling to see if there were some quick gains that could be made. As it turns out, about 60% of CPU time was spent decoding base64 using Python's standard library. I was shocked.
I still don't see why you'd compare uncompressed MesasgePack to compressed JSON.
In other words, compressed MesasgePack is probably only tiny amount smaller than compressed JSON.
I interpreted your original message as requesting a comparison between compressed JSON vs. uncompressed MessagePack, which didn't make sense (but which I see people ask a lot, including elsewhere in this thread). Sorry if I misunderstood.
You interpreted it correctly. And it makes sense. If MessagePack is alternative to JSON so in JSON.zstd if I need compactness.
I've been building a new ad-hoc data format to replace JSON for a couple of years [1], and am nearing completion of the reference implementation in go. It natively supports the following types:
* Nil : No data (NULL)
* Boolean : True or false
* Integer : Positive or negative, arbitrary size
* Float : Binary or decimal floating point, arbitrary size
* Time : Date, time, or timestamp, arbitrary size
* URI : RFC-3986 URI
* String : UTF-8 string, arbitrary length
* Bytes : Array of octets, arbitrary length
* List : List of objects
* Map : Mapping keyable objects to other objects
* Markup : Presentation data, similar to XML
* Reference : Points to previously defined objects or other documents
* Metadata : Data about data
* Comment : Arbitrary comments about anything, nesting supported
But the most important feature is that it is a paired format: a binary format [2] and a text format [3], which are 1:1 compatible. This allows you to transmit in the binary format, and only convert to text when a human is involved.
I've put together a quick comparison here: https://github.com/kstenerud/concise-encoding#comparison-to-...
Currently, I'm finishing off the go implementation [4], which so far I've managed to get running 30% faster than the json codec, using less than half the memory. I'll be pushing the binary codec to master soon, and the text codec shouldn't take much longer since the code is pretty modular.
[1] https://github.com/kstenerud/concise-encoding#concise-encodi...
[2] https://github.com/kstenerud/concise-encoding/blob/master/cb...
[3] https://github.com/kstenerud/concise-encoding/blob/master/ct...
[4] https://github.com/kstenerud/go-cbe/tree/new-implementation
The idea is to make a format for the 80% case, so things like ISBN are definitely out.