CBOR – Concise Binary Object Representation
cbor.io
cbor.io
I wrote [1] a pretty comprehensive (and admittedly biased) critique of the CBOR standard years ago.
[1] https://news.ycombinator.com/item?id=14072598
Disclaimer: I wrote and maintain a MessagePack implementation.
CBOR was inspired by MessagePack. MessagePack was developed and
promoted by Sadayuki Furuhashi ("frsyuki"). This reference to
MessagePack is solely for attribution; CBOR is not intended as a
version of or replacement for MessagePack, as it has different design
goals and requirements.
I fail to see what is wrong with having both CBOR and MessagePack, or with trying to bring a MessagePack-like thing to the IETF, if MessagePack's design (lack of extension points, etc.) was problematic for enabling future applications.Also: CBOR is not one person. The CBOR working group, and the IETF in general, put forward the spec. If you object to how the process went down, your quarrel likely lies with IETF, not any individual.
As for the rest of the process, that HN post I linked to has a basket of links at the end which are a good summary of what happened. I still (as I wrote before) won't characterize events, I think people should make up their own mind about it. But I think it's not biasing to reiterate:
> At least cbor-ruby started with the MessagePack sources. The story is that Carsten took MessagePack, wrote a standard and added some things he wanted, and called it something else.
- Msgpack serialization and deserialization is very fast in many languages - often 100x faster then JSON
- Msgpack natively supports encoding binary data
- Msgpack has type extensions, making it trivial to represent common types in an efficient way (eg. IPv4 address, timestamps)
- Msgpack has good libraries available in many languages
If you do not care about those things (no binary data, no need for extended types, not performance critical) then JSON is just fine.
FlatBuffers, Protobuf, Cap'n Proto, etc., all require an external schema configuration that you compile into a code chunk that you include into your program. Without this it is impossible to make sense of the data. In our case, the data is semi-structured and changes frequently. The prospect of maintaining a schema registry for all the data users and keeping everyone up to date and backwards compatible is enough of a burden that it was excluded.
Avro also uses schemas, but since the schema is embedded in the file it is self-describing so the reader does not need to do anything special to interpret the data. But Avro's C library is buggy and the python deserialization performance was terrible, so Avro was not selected.
MessagePack is marketed as "It's like JSON, but fast and small", but I don't view it as a JSON replacement at all. I think JSON is great for most projects, especially anything talking to a JS front end, and as soon as you move into something requiring more performance, you probably need schema/versioning guarantees that something like FlatBuffers or Cap'n Proto would give you.
I think MessagePack is a great on-disk binary format, and a pretty good network binary format if you don't have complicated application architecture (i.e., a video game client/server, a chat client, etc.) It's way, way faster than JSON which can be important for mobile, IoT, and embedded work). It's more compact than JSON if you want to avoid the overhead of gzipping/gunzipping everything. It compresses similarly as JSON+gzip if that's important to you. It's not confounded by things like canonicalization or attempts to add a schema. It's also amenable to streaming, something you... mostly can't do with JSON without going to a lot of trouble. It's also easy to implement -- which is important in some cases. Try implementing a JSON encoder/decoder!
Edit: to be a little more constructive, I read this a while ago and it scared me off implementing anything JSON related forever: http://seriot.ch/parsing_json.php#4.
And on the topic of cute names, back when I was learning Elixir, I wrote a wrapper for an Erlang YAML parser. I called it Mark Yamill :)
The format should use contemporary machine representations of data
...and then see in section 1.2,
All multi-byte values are encoded in network byte order (that is, most significant byte first, also known as "big-endian")
I know it's largely a choice of tradition, but it seems almost anachronistic to specify any new protocols as BE when LE is the overwhelming majority of machines today, and probably has been for at least the past two decades.
Head desk.
E.g. for a switch connected to networks A: 0xa, B: 0x1, C: 0x3, a only the first byte 0x1 of the packet destined to 0x1234 needs to be processed before forwarding, saving some time compared to LE, where the entire address 0x4321 would have to be processed to find out that it's at the 0x1 network.
The gate savings on doing endiannes specific circuits are close to zero in comparison to many many other things a typical logic block comes with today.
Today, you have to process way way more bytes in a single clock cycle anyways.
Internally, for a chip designer, almost all modern high speed serial busses look way wider than a single byte. And all of its serialness is kept inside the transmitter/serdes/interface blocks without any external exposure.
So if you use big-endian byte orders you're actually sending your bits all jumbled up (bits from the most significant byte first, then the least significant bit from within each byte first).
2015: https://news.ycombinator.com/item?id=9597198
2013: https://news.ycombinator.com/item?id=6932089
2013: https://news.ycombinator.com/item?id=6632576 (the largest)
Anyone have any thoughts on using canonical CBOR for object signing? Currently, I'm building a system with a content-addressable data store, and I'm particularly interested in data formats with a canonical form for this use-case.
As a result, round-tripping through a CBOR implementation still may result in data structure changes. Depending on the type of change and any exploits in say the hashing algorithm, this could be a security issue.
On the flip side, you can just tag a byte array as CBOR data, and sign it. Unlike JSON, you don't need to perform an encoding/escaping to make one document safe to embed into another document.
* Cap'n Proto
* ASN.1
* gzip-compressed JSON
in various ways. (I don't know much about progress in serialization methods.)
[1]: https://en.wikipedia.org/wiki/Bencode
[2]: http://web.archive.org/web/20140701085126/http://tnetstrings...
Raw json is 90mb, cbor was 80mb. json+gzip takes it to 30mb and cbor+gzip was 31mb.
That being said, the schema has a lot of repeated keys, so that's why gzip helps a lot.
CBOR is not about compression to make it smaller, but for machine readability. For example for integers CBOR uses binary representation so that machines can read it directly without converting from string to integer in JSON.
You can have a look to results of performance testing involving these binary message formats vs json.
http://zderadicka.eu/comparison-of-json-like-serializations-...
https://github.com/ludocode/schemaless-benchmarks#speed---de...
1. https://github.com/boonproject/boon/wiki/Boon-JSON-in-five-m...
For me it does not really matter which binary message format you are using, I think there is not a huge difference between the different libraries. There are some feature differences for sure, library quality maybe. If you are going for the most performance that is possible to achiave you could look at SBE[2] developed by performance freaks.
because you _seem_ to imply that json parsing is faster/better than binary parsing...
Do I?
Also, how does this compare to JSON binary serialization, such as BSON?
Because it is pretty much optimized for machine readability and lightness, one of usage is microcontrollers for IoT, with CBOR + CoAP(kinda HTTP for IoT), although I wouldn't say it is common yet, since CoAP needs IPv6.
[1]: https://wicg.github.io/webpackage/draft-yasskin-http-origin-...
I didn't find a CBOR/MessagePack nor UBJSON implementation that was microcontroller-friendly and easy to use.
In the end I ended up just using plain JSON. Easy to debug as you can easily see what's on the wire, easy to implement, relatively small code size.
I almost had a better floating point compression format, but it turned out to be too complicated, and only works well for decimal floating point, so I'll probably not use it in CBE.
[1] https://github.com/kstenerud/concise-encoding/blob/master/cb...
This would make it easier for CBEv2 to add types without changing how a byte is interpreted.